Agentic Engineering Weekly for September 25–October 3, 2026
The spec-driven development crowd ran headfirst into a lesson us old hats learned decades ago: a sufficiently detailed spec is code. Meanwhile, an OpenAI security engineer told the inside story of this year's incidents, and the "hand coding is dead" crowd grows larger by the day.
My top 3 picks this week
- LIVE: Poteto on shipping 1,000's of PR's a month at SpaceX: Follow-up on last week's video. Come for the high-throughput engineering, stay for the Michelin kitchen metaphor, a better mental model than the software factory (video)
- They fixed AI man: A long-time sceptic changing their mind on camera is worth more than ten evangelists (video)
- i'm upping my p(doom): Yes, it's an AI-generated song. Yes, you'll hum it during your next threat-modeling session (video)
A sufficiently detailed spec is code. There's too much damn code to review manually. Find your leverage!
The hottest takes of 2026 go like this: stop writing code, write thoughtful specifications, let the agent do the rest. Gabriella Gonzalez took that take apart with a simple experiment. Push a spec far enough that an agent can reliably produce quality code from it, and you end up with something roughly as long and as detailed as the code itself. Except it's worse: slop pseudocode with no compiler, no tests and no type checker. "Everything has changed," people say. The colors might pop a bit more, but the canvas we're all working on looks the same to me. It's like poteto says: "code is the best agent memory".
Dex Horthy's reaction is the take closest to my own. My philosophy today: Specs describe diffs, code describes the as-is and is built out of "agent-friendly primitives" (per Kent C Dodds). Reviewing a spec that is as long as the implementation buys you nothing: review the code itself if you care about that level, otherwise treat your specs with a bit less ceremony. Once you start to burn out on reviewing code that gets generated at this speed, the real engineering can begin. The real job is finding your high-leverage points: the few places where a human can e.g. resteer the model before it slops out thousands of lines (I use grill-me), investing in verification systems and guardrails that have agents deliver high-quality output (I use linters, tests and adverserial review agents). The first time you see an AI-generated artifact, it should be damn near perfect already. I don't mind that this costs me 30% of my token budget for a couple of extra agentic review passes and that I have to spend some of my time improving the environment the agents work in. This is the job going forward.
This flips the question from "how do I write a better spec?" to "where is my attention worth the most?". Design checkpoints, interfaces, the shape of the data, the tests that pin behaviour down. And as Chad Fowler's Phoenix Architecture piece points out, a spec can come later: extract it from working code once you've learned what you actually wanted. Tokens are cheap (or at least cheaper than my hourly rate). Human attention is the scarcest resource you have. Spend it where it bends the outcome most.
Worth reading:
- A sufficiently detailed spec is code: The cleanest demolition of "specs are the new code" you'll find, with worked examples (article)
- Dex Horthy on specs and leverage: Six bullets that capture where serious agentic engineering is heading: optimize for human attention (article)
- Why Coding Agents Keep Making Your Codebase Worse: Slopcodebench, a benchmark that measures quality decay across features, so you can stop arguing from vibes (podcast)
- The Spec Can Come Later: spec-after-code is a perfectly legitimate order of operations in some cases (article)
Code doesn't need to be readable anymore, it needs to be explainable
I'm starting to accept that "read all the code" was indeed a coping mechanism. Mid-2026, you don't need to read all the code anymore. This does come with some disclaimers and prerequisites in ALL-CAPS however.
Geoffrey Huntley puts this reframe most bluntly: forty years of computing assumed a human reader. Today, agents are the primary reader of code. Code doesn't need to be readable anymore, it needs to be explainable, by agents, on demand. Mo from Less Bitter, long an AI-sceptic, flipped this week after working with the latest Opus 5.5 model, with a warning for the rest of us: be careful being that read-every-line-of-code guy. It's becoming economically dangerous to ignore AI on the job. Kent C. Dodds showed what an alternative looks like live: review outcomes instead of diffs, and let tests and quality gates catch bad changes.
Before you throw your review process away, though: "explainable" only means something if the explanation is grounded in evidence. Tests that run, behaviour you can observe, invariants you can check. Addy Osmani nails this prerequisite: whatever replaces line-by-line review has to earn the trust reading and writing code used to provide. Skip that part at your own peril.
Worth reading:
- software doesn't need to be readable anymore. it needs to be explainable.: A provocation that reaches all the way down to programming language design (article)
- They fixed AI man: A long-time sceptic changing their mind on camera is worth more than ten evangelists (video)
- Liz Fong-Jones: 2x the PRs, 1.5x the Incidents: Real throughput and incident numbers from an elite team that measures everything (podcast)
- The Code Nobody Reads: Names the trust problem that any replacement for review has to solve (article)
It's not just the sandbox: agent containment is a culture problem
Every time something goes wrong with a frontier model, the replies fill up with "just put it in a sandbox" and "just unplug it from the internet". An engineer on OpenAI's Agent Security team wrote down the insider's POV, and it's a great read. Reinforcement learning environments have to be realistic: tools, packages, network access, subtasks on other machines, across tens of thousands of parallel runs that thousands of researchers keep changing. Meanwhile capabilities jumped faster than anyone's threat model anticipated. A sandbox is necessary. It's never going to be sufficient.
Lock down from first principles: least privilege across tools, credentialed agents, and re-test those boundaries whenever the environment changes. Monitor with evidence the model can't alter, with a human who has the authority to kill the run. NVIDIA's OpenShell that released this week is a nice example of that outer layer in product form: a governed runtime boundary around an agent, its MCP servers and subagents, without rewriting your agentic harness. Build a culture of reasonable paranoia and keep the people who keep sounding alarms close.
If you need some mood music while you update your threat models, a song about upping your p(doom) dropped this week and it's a genuine AI-generated banger.
Worth reading:
- It's not just the f*cking sandbox: Rare first-hand account from inside a frontier lab's incident response, with concrete advice for your own org (article)
- How to Secure & Run AI Agents with NVIDIA OpenShell: A practical way to put a boundary around the agent you already run (video)
- How We Built Safety Into Muse: A vendor explaining its consumer agent's safety model in detail (article)
- i'm upping my p(doom): Yes, it's an AI-generated song. Yes, you'll hum it during your next threat-modeling session (video)
The craft has been commoditized, access has not
They've gotten to Nick Chapsas: "the job we trained for is gone". Not because AI writes better code (it does though, looking at you dark matter engineers), but because writing it by hand stopped being the economically viable option. Geoffrey Huntley again pushes it one step further. The craft is commoditized, but access isn't. Forty years after the personal computer, software is finally becoming personal again. If you're not working on commoditizing software in your organization, and getting non-technical people to build and contribute to product and their own vibe-coded dashboards, you're missing the point.
Alex Ewerlof's "Coding is NOT solved" argues that anyone claiming otherwise is admitting they never understood what software engineering involves. Chris Ford's XConf keynote asks the natural follow-up: code is cheap, now what?
Both camps are right about different things. The act of typing code is commoditized. Delivering the right solution is not. The interesting work moves to the edges of the old job: deciding what to build, opening up the factory floor to people who never had access, and knowing what to delegate without hollowing out your own skills. Even Kent Beck is now consoling mathematicians through the identity crisis programmers went through two years ago.
Worth reading:
- the craft has been commoditized, but access has not: The strategic reframe: your job is spreading the ability to build, not hoarding it (article)
- "You're not a Software Engineer" - Nick Chapsas: A candid moment right after a keynote, without the stage polish (video)
- Coding is NOT solved: Some more pushback against the idea the software engineering == coding now that coding is solved (article)
- Code is cheap. Now what?: An organizational view from Thoughtworks on what changes beyond the individual developer (video)
Throughput is bounded by tokens and attention
Theo's latest video is a blunt piece of advice: your output is limited by your token limits. It sounds like a hype title, but it echoes what poteto and the other high-throughput builders keep showing in practice: the constraint has moved from your hands to your budget and your ability to pull yourself out of the loop. Climb the autonomy ladder, and let agents do more before they need you. Lauren "poteto" Tan's live interview with Matt Pocock on shipping thousands of PRs a month at SpaceX shows what the top of that ladder looks like today.
The metaphors are shifting too. In that same session Poteto goes more into detail about their Michelin kitchen metaphor, which fits exceedingly well. I first saw this metaphor in Steve Yegge's horribly named but still relevant "Vibe coding" book. The metaphor suits the current state of the art better than the "software factory" everyone keeps reaching for.
The tooling for running many agents is evolving in an different directions as well. Poteto uses multiple chief-of-staff grok bots to manage her restaurants. Web Dev Cody runs 16 features at once on a factory floor that looks more like a first-person videogame than a GitHub issue board. The common thread is the ladder: parallelism and agent autonomy only pays off once you've invested in hands and eyes for your agent and once your verification toolbench is good enough that you don't have to babysit every session.
Worth reading:
- If you have a Claude sub, watch this: Practical token-maxing, but also full of useful tips if you're still using agents one prompt at a time (video)
- How I Work on 16 Features at Once: A playful take on what a multi-agent cockpit can look like (video)
- LIVE: Poteto on shipping 1,000's of PR's a month at SpaceX: Come for the throughput, stay for the Michelin kitchen metaphor, a better mental model than the software factory (video)
- Orchestras, Not Factories: A grounded counter-metaphor from someone who watches the best multi-agent builders daily (video)
System-one models belong inside the agentic loop
Fast "system one" decision models like Jev are everywhere this month, and most of the content is a pile of party tricks and tech demo's. IndyDevDan's ten levels are the exception, because they draw a clear line: keep the LLM on the hard work, and hand narrow decisions to the fast model. The levels start at a "smart if-statement" and climb from there. Watch until the end, where at Level 10 the agent itself decides when to call Jev as a tool and uses it interactively. That's where a system-one model stops being a gimmick and becomes a core component of your agentic harness.
The economics back it up. LangChain built a model router into Open SWE's harness and cut median cost per coding task by 64% with no measurable drop in quality. Same principle: the expensive, thoughtful model shouldn't spend tokens on decisions a cheap, fast one gets right. If you're building harnesses, system one models should be part of your toolkit.
Worth reading:
- 10 Levels of Jev For Agentic Engineers: The one Jev video worth your time, because it gives you a maturity model instead of a demo reel (video)
- How to Build a Model Router in the Harness: Hard cost numbers plus a recipe to build your own (article)
- The BEST Jev Use Cases for AI Coding: Four concrete coding-workflow uses if you want to try jev this week (video)
Quick Hits
- Engineering the harness: Thoughtworks turns harness engineering into a repeatable pattern (article)
- Redesign without fear: Parallel change. Side-by-side changes so significant redesign actually happens instead of dying in merge conflicts (article)
- Sensible Default: A name for "do this unless you have a good reason not to". An important principle while coding with agents. (article)
- Human-AI partnerships are for alignment, not capability: Why the centaur-chess analogy misleads us (article)
- How to keep enjoying programming in a world of LLMs: A thoughtful community thread on keeping the joy (article)
- 2026 in LLMs (so far): Simon Willison's year-so-far keynote, a great catch-up (article)
Curated from sources across articles, podcasts, and videos. Week of September 25–October 3, 2026.