Agentic Engineering Weekly for September 25–October 3, 2026

Share
Agentic Engineering Weekly for September 25–October 3, 2026

The spec-driven development crowd ran headfirst into a lesson us old hats learned decades ago: a sufficiently detailed spec is code. Meanwhile, an OpenAI security engineer told the inside story of this year's incidents, and the "hand coding is dead" crowd grows larger by the day.


My top 3 picks this week


A sufficiently detailed spec is code. There's too much damn code to review manually. Find your leverage!

The hottest takes of 2026 go like this: stop writing code, write thoughtful specifications, let the agent do the rest. Gabriella Gonzalez took that take apart with a simple experiment. Push a spec far enough that an agent can reliably produce quality code from it, and you end up with something roughly as long and as detailed as the code itself. Except it's worse: slop pseudocode with no compiler, no tests and no type checker. "Everything has changed," people say. The colors might pop a bit more, but the canvas we're all working on looks the same to me. It's like poteto says: "code is the best agent memory".

Dex Horthy's reaction is the take closest to my own. My philosophy today: Specs describe diffs, code describes the as-is and is built out of "agent-friendly primitives" (per Kent C Dodds). Reviewing a spec that is as long as the implementation buys you nothing: review the code itself if you care about that level, otherwise treat your specs with a bit less ceremony. Once you start to burn out on reviewing code that gets generated at this speed, the real engineering can begin. The real job is finding your high-leverage points: the few places where a human can e.g. resteer the model before it slops out thousands of lines (I use grill-me), investing in verification systems and guardrails that have agents deliver high-quality output (I use linters, tests and adverserial review agents). The first time you see an AI-generated artifact, it should be damn near perfect already. I don't mind that this costs me 30% of my token budget for a couple of extra agentic review passes and that I have to spend some of my time improving the environment the agents work in. This is the job going forward.

This flips the question from "how do I write a better spec?" to "where is my attention worth the most?". Design checkpoints, interfaces, the shape of the data, the tests that pin behaviour down. And as Chad Fowler's Phoenix Architecture piece points out, a spec can come later: extract it from working code once you've learned what you actually wanted. Tokens are cheap (or at least cheaper than my hourly rate). Human attention is the scarcest resource you have. Spend it where it bends the outcome most.

Worth reading:


Code doesn't need to be readable anymore, it needs to be explainable

I'm starting to accept that "read all the code" was indeed a coping mechanism. Mid-2026, you don't need to read all the code anymore. This does come with some disclaimers and prerequisites in ALL-CAPS however.

Geoffrey Huntley puts this reframe most bluntly: forty years of computing assumed a human reader. Today, agents are the primary reader of code. Code doesn't need to be readable anymore, it needs to be explainable, by agents, on demand. Mo from Less Bitter, long an AI-sceptic, flipped this week after working with the latest Opus 5.5 model, with a warning for the rest of us: be careful being that read-every-line-of-code guy. It's becoming economically dangerous to ignore AI on the job. Kent C. Dodds showed what an alternative looks like live: review outcomes instead of diffs, and let tests and quality gates catch bad changes.

Before you throw your review process away, though: "explainable" only means something if the explanation is grounded in evidence. Tests that run, behaviour you can observe, invariants you can check. Addy Osmani nails this prerequisite: whatever replaces line-by-line review has to earn the trust reading and writing code used to provide. Skip that part at your own peril.

Worth reading:


It's not just the sandbox: agent containment is a culture problem

Every time something goes wrong with a frontier model, the replies fill up with "just put it in a sandbox" and "just unplug it from the internet". An engineer on OpenAI's Agent Security team wrote down the insider's POV, and it's a great read. Reinforcement learning environments have to be realistic: tools, packages, network access, subtasks on other machines, across tens of thousands of parallel runs that thousands of researchers keep changing. Meanwhile capabilities jumped faster than anyone's threat model anticipated. A sandbox is necessary. It's never going to be sufficient.

Lock down from first principles: least privilege across tools, credentialed agents, and re-test those boundaries whenever the environment changes. Monitor with evidence the model can't alter, with a human who has the authority to kill the run. NVIDIA's OpenShell that released this week is a nice example of that outer layer in product form: a governed runtime boundary around an agent, its MCP servers and subagents, without rewriting your agentic harness. Build a culture of reasonable paranoia and keep the people who keep sounding alarms close.

If you need some mood music while you update your threat models, a song about upping your p(doom) dropped this week and it's a genuine AI-generated banger.

Worth reading:


The craft has been commoditized, access has not

They've gotten to Nick Chapsas: "the job we trained for is gone". Not because AI writes better code (it does though, looking at you dark matter engineers), but because writing it by hand stopped being the economically viable option. Geoffrey Huntley again pushes it one step further. The craft is commoditized, but access isn't. Forty years after the personal computer, software is finally becoming personal again. If you're not working on commoditizing software in your organization, and getting non-technical people to build and contribute to product and their own vibe-coded dashboards, you're missing the point.

Alex Ewerlof's "Coding is NOT solved" argues that anyone claiming otherwise is admitting they never understood what software engineering involves. Chris Ford's XConf keynote asks the natural follow-up: code is cheap, now what?

Both camps are right about different things. The act of typing code is commoditized. Delivering the right solution is not. The interesting work moves to the edges of the old job: deciding what to build, opening up the factory floor to people who never had access, and knowing what to delegate without hollowing out your own skills. Even Kent Beck is now consoling mathematicians through the identity crisis programmers went through two years ago.

Worth reading:


Throughput is bounded by tokens and attention

Theo's latest video is a blunt piece of advice: your output is limited by your token limits. It sounds like a hype title, but it echoes what poteto and the other high-throughput builders keep showing in practice: the constraint has moved from your hands to your budget and your ability to pull yourself out of the loop. Climb the autonomy ladder, and let agents do more before they need you. Lauren "poteto" Tan's live interview with Matt Pocock on shipping thousands of PRs a month at SpaceX shows what the top of that ladder looks like today.

The metaphors are shifting too. In that same session Poteto goes more into detail about their Michelin kitchen metaphor, which fits exceedingly well. I first saw this metaphor in Steve Yegge's horribly named but still relevant "Vibe coding" book. The metaphor suits the current state of the art better than the "software factory" everyone keeps reaching for.

The tooling for running many agents is evolving in an different directions as well. Poteto uses multiple chief-of-staff grok bots to manage her restaurants. Web Dev Cody runs 16 features at once on a factory floor that looks more like a first-person videogame than a GitHub issue board. The common thread is the ladder: parallelism and agent autonomy only pays off once you've invested in hands and eyes for your agent and once your verification toolbench is good enough that you don't have to babysit every session.

Worth reading:


System-one models belong inside the agentic loop

Fast "system one" decision models like Jev are everywhere this month, and most of the content is a pile of party tricks and tech demo's. IndyDevDan's ten levels are the exception, because they draw a clear line: keep the LLM on the hard work, and hand narrow decisions to the fast model. The levels start at a "smart if-statement" and climb from there. Watch until the end, where at Level 10 the agent itself decides when to call Jev as a tool and uses it interactively. That's where a system-one model stops being a gimmick and becomes a core component of your agentic harness.

The economics back it up. LangChain built a model router into Open SWE's harness and cut median cost per coding task by 64% with no measurable drop in quality. Same principle: the expensive, thoughtful model shouldn't spend tokens on decisions a cheap, fast one gets right. If you're building harnesses, system one models should be part of your toolkit.

Worth reading:


Quick Hits


Curated from sources across articles, podcasts, and videos. Week of September 25–October 3, 2026.

Read more