Insights
AdaptationPerformCapability Isn't Judgement
25 September 2026
A while back I wrote about an experience building a design for a client using AI. The result was fast, comprehensive and remarkably convincing. It passed the obvious checks and looked ready to go. Something still didn’t feel right, so I went back over it once more before it went out the door.
What we found changed how we’ve approached every AI-assisted design since. At the time I framed that as a lesson about trusting AI’s output. Looking back, I don’t think that was quite it. The AI hadn’t hallucinated some critical fact or produced something obviously broken. In many ways it had done exactly what it was asked to do. The problem was that we’d done a better job testing the answer than challenging the question.
That distinction has been sitting with me, and this week two people I’ve been reading brought it back from completely different directions.
Capability isn’t authority
I’ve followed Greg Woolley’s work for years. His latest venture as a serial entrepreneur is Causant, an authority control plane built specifically for this problem: AI agents that can inspect a repository, implement a change, run the tests, diagnose failures and repeat the cycle across live codebases faster than any human engineering team could review it. That’s genuine capability, and Greg’s building the infrastructure that decides what it’s actually allowed to do with it. His recent piece on the problem names what’s missing precisely: “An agent may be capable of changing three hundred files, but that does not mean the current task authorises it to do so.” And: “A test can tell you that a change behaves as expected. It cannot tell you whether you should have changed that part of the system.”
Evidence theater
A few days later I read Steve Blank’s account of what happened to his Lean LaunchPad class at Stanford, a course that’s run the same way for 15 years: form a hypothesis, get out of the building, talk to customers, get proven wrong, adjust, repeat. This year, all eight teams arrived with finished, AI-built products before a single customer conversation. Blank calls it “evidence theater”: a polished result that looks like progress while actually skipping the only process that ever taught anyone what a market actually wants. Some students got attached to what they’d built and started interpreting real customer feedback through it, rather than the other way round.
Steve and Greg have never met, as far as I know. They’ve arrived at the same underlying concern from about as far apart as two smart people can get, one from enterprise software engineering, one from teaching first-time founders.
The friction was doing something
A client design, an enterprise codebase, a Stanford classroom. Three completely unrelated contexts, one identical failure: something that works gets mistaken for something that should exist, because AI has quietly removed the friction that used to force a harder question in between.
That friction wasn’t an inefficiency. It was the thing doing the work: exposing hidden assumptions, buying time for contradictory evidence to surface before you’re too attached to an answer, and, not incidentally, building the experience that later becomes judgement. A good architect doesn’t just know more patterns than everyone else, they’ve watched those patterns succeed and fail in different environments. None of that comes from knowing the correct methodology. It comes from having done the slow version enough times to recognise when something’s off.
An answer before you deserve it
AI can give you an answer before you deserve it. Not because there’s some initiation ritual you need to earn before using the tool, but because getting to an answer the slow way used to expose you to everything hidden inside it. The person who built a model by hand probably understood where the numbers came from, because they typed them in. The founder who couldn’t afford to build immediately had no choice but to understand the customer first.
Where the stakes are real, that can’t just be left to someone noticing in time the way I got lucky on my own design. That’s not a scalable control, it’s a hope. Greg’s answer is the one that actually holds up at speed: engineer the authority in as a structural constraint, rather than waiting for a human to catch it. Aviation didn’t solve “the pilot will notice if something goes wrong” by asking pilots to concentrate harder. It built operating envelopes, independent checks and fault containment that hold regardless of who’s watching. Judgement still decides where those boundaries should sit. It just can’t be the only thing standing between capability and consequence anymore.
AI will keep making us more capable, and we should use every bit of it. What I’m less sure of is whether it’s making anyone better at knowing what’s actually worth doing with that capability, or just faster at doing more of whatever we’d have done anyway.
How do you see it?