Insights

Adaptation

The AI Design Looked Convincing. That Was the Problem.

7 August 2026

We had a problem to solve for a client.

It was a real operational problem, involving physical assets, people, safety and technology. We needed to understand what was happening in an industrial environment and design a practical way of monitoring it.

We put our AI tooling to work. The result was impressive.

It developed two alternative architectures. One used relatively inexpensive consumer-grade cameras and computer vision. The other was a more industrial approach using infrared sensors.

It identified specific components, down to manufacturers and part numbers. It worked through how they would be installed and integrated. It considered the data flows. It designed dashboards. It described how operators would interact with the solution and how exceptions would be handled.

We pushed the model’s reasoning level up because this wasn’t casual research. We wanted it to work hard on the problem. And it did.

The resulting design looked like something that could have taken a conventional consulting and engineering team days, perhaps weeks, to assemble. It was detailed. It was coherent. It was technically literate. Most importantly, it was convincing.

We were getting ready to put it in front of the client.

And something bothered me.

I couldn’t tell you exactly what was wrong

This wasn’t a case of spotting an obviously invented product or some ridiculous technical assertion. I’d read the work. The architecture made sense. The components existed. The arguments were logical.

But experience was telling me to have another look.

So rather than asking the model that produced the design to review its own work, we did something different. We gave the completed design to a completely separate, heavyweight model. We turned its reasoning up as far as we could and gave it a very different job.

Don’t design the solution.

Attack it.

Look for unsupported assumptions. Check the specifications. Challenge the proposed components. Ask whether the sensors could really do what we were claiming. Look at the physical environment. Examine the integrations. Challenge the safety assumptions. Find the things that would cause this apparently excellent design to fail in the real world.

What came back was salutary. There were assumptions buried inside the design that were much less robust than the confident final document suggested. Technical capabilities had been extrapolated. Some conclusions were plausible rather than proven. Elements that looked perfectly reasonable when assembled into the overall design became considerably less comfortable when independently challenged.

It didn’t mean the whole design was worthless. Far from it: the AI had done an extraordinary amount of useful work.

But it had also produced something potentially more dangerous than a bad answer: a very good-looking answer containing weaknesses that weren’t obvious.

That changed our process

Our response wasn’t to stop using AI for serious design work. Quite the opposite: we added more AI. But we gave it a completely different job.

We built a heavyweight validation stage into our design process.

Our design tooling is encouraged to explore. It can consider alternatives, find components, develop architectures, calculate, prototype and move quickly. I don’t particularly want to constrain that process by making it obsess about every possible objection while it is trying to create.

But it no longer gets the final word. Before substantive design work leaves the building, it goes through a separate validation engine: different model, separate context, maximum reasoning, different objective.

Its purpose isn’t to make the original work sound better. Its purpose is to find reasons not to believe it. What assumptions have been made? What claims aren’t adequately supported? Do the nominated components actually meet the claimed specification? Has a manufacturer’s statement been stretched beyond what it really says?

What happens at the edges of the operating envelope? What has the design ignored? What would have to go wrong for the recommendation to fail? Where should we be telling the client that we don’t yet know?

Only then do we make the human judgement about what survives, what needs to change and what we’re prepared to stand behind.

AI has moved the bottleneck

There’s a much bigger lesson here than our particular solution design. For years we’ve talked about AI primarily in terms of productivity. How much faster can we write the report? How quickly can we analyse the data? Can we create the application in a day instead of a month?

Those questions still matter. The productivity gains are becoming extraordinary. But I think we’re approaching another stage: production is becoming cheap, and assurance isn’t.

AI can produce code, analysis, technical designs, strategies, contracts, presentations and business cases at a speed that would have seemed implausible only a few years ago. The challenge increasingly becomes determining whether they’re right, and that’s particularly important because the models are getting better.

A poor AI answer is relatively easy to deal with. You recognise that it’s poor and investigate it. A beautifully structured, technically detailed and entirely plausible answer is different.

Confidence is infectious. The better the output looks, the easier it is for us to stop asking difficult questions.

Don’t ask the designer to certify the design

There’s nothing particularly revolutionary about separating creation from assurance. We’ve been doing it in other professions for a very long time. Code gets peer reviewed. Financial statements are independently audited. Engineering designs are checked. Safety-critical systems go through verification and validation. Academic work is subjected to peer review.

We generally recognise that the person who created something isn’t necessarily the best person to identify its weaknesses. AI shouldn’t be different.

In fact, AI gives us an interesting new opportunity. Historically, putting a second highly capable team onto a piece of consulting or design work purely to find faults with the first team’s work would have been expensive. Now the incremental cost of heavyweight independent review can be tiny compared with the value of the work, and microscopic compared with the potential cost of getting it wrong.

So use it.

What happens to expertise?

This also changes my view of the argument about AI replacing experts.

AI is undoubtedly taking over parts of expert work. It can research faster than I can. It can explore more alternatives. It can produce extraordinarily detailed technical material. It can cross disciplines in ways that would traditionally require several different people.

But expertise hasn’t disappeared from the process. It’s moving.

The first model produced the design. The second model challenged it. But neither model decided that the first answer made me uncomfortable.

Experience did. And experience decided what to do about it.

That may be one of the more important changes AI brings to professional work. The value of expertise moves away from simply being able to produce the work and towards knowing what problem to solve, what questions to ask, what evidence to demand, when something doesn’t smell quite right, and ultimately what you’re prepared to put your name behind.

The lesson from our experience wasn’t that we should use less AI. We did the opposite: we added more. But we gave it a different job.

This is what operator, not spectator actually looks like in practice.