Insights

AdaptationPerform

Does Your Agent Need a Performance Review?

14 August 2026

Thirty years ago, in a consulting business I founded, we automated the production of an individual performance review pack for each consultant. The pack wasn’t the performance conversation itself; it was the factual backdrop to it.

It included the basic hygiene and financial measures: revenue generated, billable utilisation, project performance, contribution across the projects the consultant had worked on, and even whether they’d completed their timesheets properly. None of those measures told us everything about the person, but they gave us a reasonably objective base from which to have a better conversation.

We then layered role-specific expectations over the top. A project manager wasn’t expected to know how to tune a SQL statement to the same level as an experienced database administrator, and a senior technical specialist might be expected to perform at a much higher level in a particular competency than somebody relatively new to the role. The important point was that performance was assessed in context: we didn’t simply ask whether somebody was “good”, we asked whether they were performing well against the expectations of the job they’d been given.

I think we’re rapidly going to need the same discipline for AI.

We keep talking about AI agents as workers

The language is already everywhere: AI teammate, digital worker, copilot, agentic workforce, autonomous employee. Some of those descriptions are overdone, but the direction is real. AI systems are moving beyond answering questions and increasingly:

  • analyse information
  • prepare recommendations
  • update systems
  • coordinate workflows
  • communicate with customers
  • create technical artefacts
  • and initiate actions

Yet we often manage those systems nothing like we manage people. The pattern is frequently: build it, test it, deploy it, assume it continues to work. That seems increasingly inadequate.

If an AI agent is performing real work, then perhaps we should ask the same basic question we ask of any other operational capability.

How is it performing?

Start with a job description

Before measuring an AI agent, we need to be explicit about its job. Imagine an operational AI agent whose role is to identify maintenance exceptions and recommend appropriate action to a maintenance planner. That agent shouldn’t be measured against the same criteria as a sales-qualification agent, or one reviewing API specifications. Its performance measures should reflect the work it has actually been assigned, which means establishing both basic operating measures and role-specific competencies.

The principle is exactly the same as our old consultant performance framework. Don’t expect the project manager to be the DBA. Don’t expect the customer-service agent to make an engineering safety decision. And don’t let an AI agent become responsible for a task simply because it’s technically capable of producing an answer.

What goes into an AI performance review?

The first layer is basic operating hygiene. For an AI agent, that might include:

  • availability
  • response time
  • cost per task
  • task completion rate
  • retries
  • tool failures
  • escalations
  • and human intervention required

Useful information, certainly. But these are the AI equivalent of timesheets and utilisation: they tell us something about operational discipline, not whether the agent is doing a good job.

The next layer is quality. Depending on the role, that might include:

  • accuracy
  • false positives
  • false negatives
  • completeness
  • evidence quality
  • compliance with business rules
  • human acceptance rate
  • rework required
  • and downstream business outcome

I particularly like human acceptance rate as a measure. An agent that produces 10,000 outputs for almost no cost isn’t particularly productive if people substantially rewrite half of them. That’s why “cost per generated output” can be misleading. A much better measure is cost per accepted output: the production cost, the review cost and the remediation required to make the work usable, all counted together.

Then come the competencies

This is where things get more interesting. An AI agent should have a competency profile appropriate to its role. A maintenance-planning agent, for example, might be expected to demonstrate:

  • Asset history interpretation: Expert
  • Manufacturer documentation: Expert
  • Work-order analysis: Expert
  • Cost forecasting: Working knowledge
  • Safety-critical recommendation: Recommend only
  • Autonomous equipment control: Not authorised

Now we’re no longer asking the vague question “is the AI accurate?” We’re asking whether this system is performing at the level we require for this particular job. That’s a much more useful conversation.

Performance should affect authority

There’s another dimension that’s particularly important with AI: authority. An agent may initially be authorised only to observe, then perhaps to analyse, then to recommend, then to propose an action, then to execute that action with human approval. Eventually, after sufficient evidence, it may be allowed to act autonomously within tightly defined limits.

That means AI agents can effectively be promoted and demoted. If an agent performs well over time, its authority might increase. If performance deteriorates, permissions can be reduced. If a new model release suddenly increases false positives, the agent might move back into supervised mode. If a policy violation occurs, one capability might be withdrawn entirely.

This isn’t about treating AI like a person. It’s simply good operational management. Capability and authority are not the same thing: a system being able to perform a task doesn’t mean it should automatically be permitted to do it.

We can review AI more rigorously than people

There’s also a significant advantage here. Human performance reviews are often subjective and relatively infrequent. AI performance can potentially be measured continuously: every material interaction can leave evidence of what information the system saw, what recommendation it made, how confident it was, whether a human accepted it, what action followed, what happened afterwards, and whether the result was ultimately successful.

That gives us the possibility of a rolling performance review based on thousands of real decisions, rather than somebody trying to remember what happened six months ago. A quarterly agent review might say:

  • Tasks completed: 8,430
  • Accepted unchanged: 91%
  • Modified by human: 6%
  • Rejected: 3%
  • False-positive rate: 4.1%
  • False-negative rate: 1.2%
  • Evidence completeness: 97%
  • Material incidents: 2
  • Policy breaches: 1
  • Current authority: Recommend only

And then: recommendation, retain current authority, consider progression to supervised work-order creation after successful completion of the next validation cycle.

That looks remarkably like a performance review. Because functionally, that’s what it is.

And then there is drift

There’s one complication that makes AI performance management different from human performance management: the agent itself can change underneath us. The model provider releases an update. A prompt changes. A tool changes. A source system changes. The data changes. A business rule changes. Suddenly the system we trusted last month isn’t quite the same system anymore.

That means performance management must include regression testing. Before moving a changed agent into production, replay historical cases: give it the situations it has already encountered, ask how it would handle them now, then compare the result with the current production system and with known outcomes.

In effect, we can ask the new version to sit the previous version’s exam before it gets the job. That’s an extraordinary capability.

This is really an operating-model question

The point isn’t that AI needs an annual appraisal form. The point is that organisations need a disciplined way of answering four questions:

  • What job have we given this AI?
  • How well is it actually performing that job?
  • What authority has its performance earned?
  • What evidence would make us increase or reduce that authority?

Those questions move AI governance away from abstract policy and into day-to-day management, and they become increasingly important as agents are connected to real business systems and allowed to do real work.

If we’re serious about describing AI agents as part of a digital workforce, then perhaps we should start managing their performance like one.

Not because the AI is a person.

But because anything exercising meaningful business authority should have a defined role, measurable expectations, appropriate supervision, regular review, and consequences when performance changes.

Don’t give an AI agent a job without giving it a performance review.


Postscript, 21 August 2026: I’m clearly not the only one thinking about this. Five days after this piece went up, NVIDIA published Evaluating AI Agent Skill Performance with NVIDIA SkillEvaluator, a benchmarking framework that does in tooling exactly what this argues for in principle: don’t let a capability into production until you’ve measured whether it actually lifts performance. They ran it across 300+ skills and found the gains varied far more by the specific capability than by which agent harness was running it. Same conclusion, arrived at from the infrastructure side rather than the management side.

Second postscript, 29 August 2026: McKinsey published Your AI Agents Need Performance Management, Too on 26 August. The title suggested the same argument as this piece; the content doesn’t quite make it. It’s a genuinely good conversation about organisational talent strategy, how leaders model the learning they ask of others, and why teaching your domain experts AI usually beats teaching AI specialists your domain. But it’s about the humans working alongside the agents, not about reviewing the agents themselves. Different, complementary question, not the same one with a shared headline.


More writing on growth, performance and adaptation