Keeping Up with the Code My Agents Write
9 October 2026

How I work with coding agents today, from running several sessions to reviewing their changes and keeping the tests useful.
In August 2025 I wrote about a year of working with AI assistants, with a small disclaimer in the first paragraph: no relevant experience with agents yet. Back then I was skeptical, a little worried whether my skills would still matter in a few years, and I struggled to find flow with an assistant nudging me at every small decision. That changed quickly, and I've changed my mind on more things than I expected.
This post is a snapshot of how I work with agents today. It will be outdated in a few months, and that's fine. Part of the point is that it keeps changing.
Juggling Five Sessions
On a typical day I have five or six agent sessions running in parallel, each on its own task. I start them, steer them, unblock them and review what comes back. One agent reviews, another researches, another writes code.
Funnily enough, this is what brought my flow back. Last year the constant back and forth with an assistant kept breaking my concentration. Now I'm slowly finding my way back into the same deep focus I used to have when I wrote all the code myself, just one level higher.
My role has shifted from writing most of the code to scheduling and reviewing it. The bottleneck isn't how fast the agents type anymore; it's my attention, and a lot of what follows is really about that.
Who Reviews All This?
The great thing is that you simply get so much more done now. No matter how good you were at coding, it took time, and there was always something left to do. Now you can finally finish the work and build the solution you actually had in mind. You can also codify your experience in skills and guidelines and lift the quality of the whole codebase.
But it brings new problems. Do you really need all that extra code? Who is going to review it? I don't think the tools are there yet. We'll need better ways to see what changed and where the risk is, such as visualizations or risk-based pull request reviews.
The Rule of Three, Again
When I notice I'm explaining the same multi-step procedure to an agent for the third time, I write it down as a reusable workflow and let it run automatically from then on. It feels a lot like the rule of three I used to apply before extracting duplicated code.
I keep the skills, prompts and checks I reuse in my agentic-engineering repository. It includes lookout for code hotspots and review risk, probe for replayable application checks, and recall for keeping up with my own codebase, along with prompts for refactoring agent instructions and skills.
Plans Are Disposable
Before a bigger change I think in writing: what are we changing, what are the trade-offs, what could go wrong. Often an agent grills me on it until the gaps show. Then the agents build, and once it's merged I delete the plan.
The better the models get, the less I need this. Sometimes I just write the core of a feature in code or pseudocode myself and point an agent at it with a short prompt.
The code and the tests are the record. A plan that outlives its change only drifts away from reality, and sooner or later someone, or some agent, trusts it. If your repository is full of 500-line specification files that were obviously written by an AI, I doubt anyone is reading them, and I'm not even sure your agents do.
Specs Won't Replace Code
This is where I disagree with a lot of what's being sold right now. The idea that we'll write detailed specs and let agents generate the code from them sounds appealing, and I don't think it will work. Nobody wants to maintain walls of text. Specs drift from the code the day after they're written, and an agent following an outdated spec is worse than one with no spec at all. To me it mostly looks like a new way to sell tooling.
That doesn't mean I write nothing down. But the specs I trust are the ones that can fail: tests, architecture rules checked in CI, types. The only spec that doesn't rot is one that can fail a build. Everything else is either a short, pruned instruction file or a plan I'm about to throw away.
Pruning Is Part of the Job
Models improve fast. Instructions I wrote for last year's model, such as workarounds, warnings and step-by-step hand holding, often get in the way of this year's. So every few weeks I go through my instruction files, workflows and skills and delete ruthlessly.
The same goes for the code. I regularly do cleanup and refactoring sessions where I simplify the codebase and clean up abstractions. Static analysis tools help, and so do metrics like cyclomatic and cognitive complexity, which point me to the right places. I also generate diagrams and small live apps of my codebase to see which parts change most often.
Staying Close to the Code
We are not moving away from code. Code is the only description of a system that is guaranteed to be accurate, because it's the thing that runs.
Will most developers need to understand it at the same level of detail as before agents? Probably not. But it will be essential to know where things happen and why they work: the fundamental building blocks of your system. Yes, in a couple of months anyone will probably be able to fix your production incident by following an agent's instructions, Chinese Room style. But finding new ways to apply agents, extracting workflows, and coming up with useful abstractions for your system all require the fundamentals. Great ideas are usually closer to the fundamentals than you'd think.
So I still use an IDE. I read the diffs, navigate the code, run it and debug it. When something looks off, I want to understand it myself rather than ask another agent to explain it to me. That's how I notice when an agent elegantly solved the wrong problem, and it's what lets me review fast enough to keep several sessions moving.
And honestly, I became a software engineer because I love writing code. I love navigating text with Vim motions. I love understanding and learning new things. Writing a specification and hitting enter on repeat is not going to cut it for me. I don't see software engineering disappearing any time soon, so I'd rather full-ass my imperfect approach than check out. Learn in the most chaotic way possible that is still fun enough to keep you going.
Testing Has Never Mattered More
If agents produce more changes, faster, then verification becomes the bottleneck. That's good news for anyone who cares about testing, because suddenly everyone else has to care too. Having preached about tests on this blog for years, I'm not going to complain.
In the loop, that means guardrails that reject bad output and fast feedback: compile, unit tests, integration tests with real dependencies. I've been writing about integration testing for years; with agents, those tests become the agent's eyes.
After the merge, it means monitoring. More changes reach production, so we need to notice regressions quickly: good observability, feature flags, gradual rollouts.
And because agents add entropy fast, I'm experimenting with visualizing my codebases: tracking changes per file to assess risk, and making those visualizations data-driven, verifiable and testable themselves.
Just Do the Thing
I can't help feeling that a lot of developers are picking up pennies in front of a steamroller right now. The amazing thing is already here, so use it. Don't spend your days building elaborate scaffolding that tells the model how to think; it mostly hamstrings it.
All the gains are downstream from that. Finally finishing your side project because you can power through the annoying parts. Learning with these tools instead of spending countless hours on Stack Overflow and forums. There never was a better time to learn and be curious.
What's Next
I'll write more often and in smaller pieces: notes on what works, what doesn't, and the occasional engineering deep dive. I also put together an agentic engineering page that maps how I approach this: guardrails, feedback loops and evals, plus the tools and skills I use.
If you work differently, especially if you think spec-driven development will work, I'd like to hear why.