SKIP TO CONTENT

Performance Management Needs New Metrics in the AI Era

July 6, 2026
iiievgeniy/Getty Images
  • Post

Summary.   

Even as they adopt AI, companies are measuring employee performance with familiar metrics of success: productivity, goal completion, and efficiency. As such, employees who rely heavily on AI may appear highly productive, while those who slow down to verify
  • Post

What does “good performance” look like when outputs are produced by people working with AI? And how can leaders ensure that speed and efficiency don’t come at the expense of judgment and accountability?

These are critical questions right now, as companies invest heavily in AI and leaders are under pressure to show that these investments are paying off. The Data and AI Leadership survey—a regular survey undertaken by one of us (Bean), with Tom Davenport—reported earlier this year that 91% of organizations were increasing their investments in AI, and 99% said investments in AI were a top organizational priority. Yet just 18% of these organizations say they are achieving a high degree of measurable business value from their AI investments.

While AI has improved speed and quality in many workflows, most organizations have yet to redesign how they evaluate human performance to reflect this shift. Based on the latest research and our experience in supporting companies implementing AI, we submit a three-layer measurement framework for AI-era performance that enables organizations to distinguish between metrics for human contribution, AI system and agent metrics, and human-AI systems.

The Human Performance Paradox

Right now, companies are measuring employee performance with AI with familiar metrics of success: productivity, goal completion, and efficiency. This creates a paradox. Employees who rely heavily on AI may appear highly productive, while those who slow down to verify assumptions, challenge outputs, or correct errors might appear less efficient precisely when they’re adding the most value. In this paradigm, leaders risk rewarding over-reliance on AI while penalizing the human judgment that prevents costly mistakes, optimizing for output rather than outcomes, and losing sight of what truly drives effective human performance.

We’re starting to get a picture of how applying old performance metrics to this new paradigm is working out. As AI systems outperform humans on traditional metrics, some employees are reacting defensively: In a survey by one AI vendor, 10% of employees admitted to tampering with data to make AI look worse. This shouldn’t take anyone by surprise. When leadership frames AI primarily as a workforce-reduction tool, people inevitably feel pressure to show they’re outperforming AI on standard KPIs, because they understand their jobs to be on the line.

But misaligned metrics are only half the problem—and arguably the easier half to solve. The other issue is that performance management was already broken. In a Gallup survey in 2024, when companies had barely begun to adapt to generative AI, only 2% of Fortune 500 CHROs strongly agreed that their performance management system inspired employees to improve, and just 20% of employees surveyed strongly agreed that their reviews were fair or transparent.

There’s little reason to think much has changed since then. Despite high expectations for automation, Deloitte reported this past January that 84% of companies haven’t redesigned roles, workflows, or career paths around AI. Only 30% offer performance-based incentives for employees’ AI use. Most have adjusted their talent strategies by focusing on employee education. But while they’re training people for a new way of working, they still measure them as before. The problem with this approach is that it assumes that getting people to do more of the same work, only faster, is still the right goal.

The combination of the wrong metrics and broken performance management is creating a deficit of trust. CEO expectations for AI-driven growth remain high, despite findings from Gartner that only one in 50 AI investments delivers transformational value, and one in five delivers any measurable return on investment. Employees, on the other hand, are acutely aware of the disconnects between leaders’ rhetoric and execution. This gap between boardroom confidence and frontline reality is precisely where performance management breaks down.

What Is “Good Performance” in the AI Age?

Traditional performance metrics were built for work that was discrete, repeatable, and individually owned. AI can automate many of those tasks: drafting, summarizing, translation, coding boilerplate, and first-pass analysis. In a large field experiment with more than 750 knowledge workers, access to OpenAI’s GPT-4 increased speed by more than 25% and improved task completion by 12.2%, while also delivering solutions of significantly higher quality on tasks within the model’s capability.

But the same research exposes why output-based measurement becomes dangerous.  When participants were given a task just outside the AI’s capability, those with AI access were 19% less likely to produce a correct solution than those without it. The authors call the “jagged technological frontier”: AI can perform exceptionally well on some tasks and fail abruptly on adjacent tasks that look similar to humans. If your system rewards speed-to-output without checking accuracy, it trains people to push work past the frontier and hope it holds.

Other evidence shows how legacy metrics misread where value actually comes from. In one study, AI assistance increased productivity by roughly 14% on average, but the gains were concentrated among less experienced workers. Those with more experience saw small speed gains and small declines in quality. Considering that experts add value through diagnosis, spotting edge cases, validating assumptions, and coaching others on when not to trust the tool, the tradeoff isn’t one most leaders would knowingly make.

There is also a subtler, longer-term risk. Controlled experiments show that repeated interaction with biased AI outputs can amplify bias in human perceptual, emotional, and social judgements. Participants are often unaware of the extent of the AI’s influence, rendering them more susceptible to it. The colleague who looks “data-driven” may quietly be learning the model’s bias. Standard performance systems have no mechanism to surface this.

A related concern is proliferation of “workslop,” or quickly-produced but low-quality work generated by or with AI that is characterized by redundancies, errors, and, consequently, minimal content value. Workslop wastes the time of those who receive it, but it also makes recipients think less of their colleagues who send it, degrading team cohesion. Organizations that measure output volume without measuring accuracy are not just misjudging performance, they are actively undermining it. Managers already report that they’re struggling to keep up with employees’ outputs, in part because execution that used to take a week can be done in hours, but quality control still takes time from experienced leaders with proven judgment.

Agentic AI has made the performance metric problem even harder. Most organisations have treated AI as a tool used by an individual employee. While only 11% of organizations have successfully deployed AI agents in production, their introduction into workflows creates what we call a co-performance problem. When outcomes are produced by mixed human-AI systems, your evaluation framework must answer three questions simultaneously: How well did the human perform? How well did the AI system perform? And how well did the human-AI pair perform together?

Based on the results of a study on algorithmic control, we argue that AI agents can shape behaviour through mechanisms that resemble management—recommending and restricting work, recording and rating performance data, and triggering rewards or replacement decisions. Without explicit measurement and governance of both the agent and the employee, you lose the ability to hold either accountable. When the agent scorecard and the employee scorecard become the same document, neither tells the truth.

A New Framework for Measuring Performance

Based on our ongoing research and our experience working with companies, we submit a three-layer measurement framework that breaks down the performance of the person, the AI systems and agents, and combined human-AI output. Each layer should carry a small number of metrics, reviewed frequently (monthly or quarterly), and tied to visible decisions: coaching, staffing, promotion, and tool governance.

1) Human-contribution metrics

To determine how people are actually performing, shift away from outputs AI can inflate and toward capabilities AI cannot replace. Three tend to matter across roles.

Boundary judgement: How reliably does someone detect when AI is out of its depth, and what do they do next? Potential measures are:

  • Escalation accuracy rate: An escalation is classified as “justified” if a subsequent audit confirms that the AI output was indeed flawed, incomplete, or outside the AI’s reliable scope of competence.
  • Source traceability score: Are data source, creation date, and the model that was used transparently documented? Sample size should be defined in advance (e.g., 20% of all deliverables per quarter).
  • Override quality index: A correction is “justified” if the employee has left a documented explanation (error in output, outdated data, context gap, etc.). Unjustified corrections count negatively against the score.

Orchestration: Can the person use AI tools to increase team throughput, not just personal output, i.e., can they enable others? Relevant measures can be:

  • Team AI adoption rate: Supplemented by depth segmentation: sporadic (<2×/week), regular (2–4×/week), integrated (daily, cross-task). The distribution across segments is what matters.
  • Workflow contribution index: In this, we suggest weighting how the employee improved a workflow: newly developed (1.0), existing workflow optimized (0.5), documented and shared (+0.25 bonus). Captured on a quarterly basis.
  • Throughput to headcount ratio: Tracked as a time series. More relevant than the absolute value is the rate of change: a rising score at constant or declining team size is an indicator of effective orchestration.

Learning velocity: Is the person adapting as tools, workflows, and policies change? Organizations that invest in structured AI capability-building send a powerful signal that mastering AI is a path to advancement, not a threat to job security. Measures that capture this important dimension are, for example:

  • Tool adoption lag: Measured in working days between official rollout and first documented productive use by the employee. Lower values = higher learning velocity. Can be tracked as a team average or individual value.
  • Training to application rate: This tracks documented behavior changes, which means: at least one concrete application example was captured within 30 days of completion by the employee or their manager.
  • Experimentation rate: An “experiment” is defined as a documented attempt to try a new application, prompting strategy, or tool combination; regardless of outcome. Fosters a learning culture in which failure is not penalized.

2) AI system and agent metrics

Model accuracy or uptime is not sufficient, especially for agentic systems. The following three dimensions give you better insights:

Objective attainment: Did the AI agent actually do what it was supposed to do, within the boundaries it was given? To understand objective attainment, the following measures can help:

  •  Task completion rate: A task is “successfully completed” if the output meets predefined acceptance criteria without requiring human correction or re-submission. Tracked per agent, per task type, and over time to detect performance drift.
  • Error rate: Errors should be classified by severity (critical/moderate/minor) and logged with root cause. A low overall error rate that masks a high critical error rate is a significant risk signal and should be reported separately.
  • Objective drift index: Particularly relevant for agentic systems operating over multiple steps. Flags instances where the agent optimised for a measurable proxy rather than the intended outcome.

Explainability and traceability: Can we tell why it did what it did? And can we prove it? To get answers to these kinds of questions, the following metrics might help:

  • Output sourcing rate: Each output should reference the data inputs, model version, and retrieval context used. Audited on a defined sample (e.g., 20% of outputs per period). A low OSR is a direct liability in regulated environments.
  • Reproducability score: A random sample of outputs is re-run using logged parameters. If the same inputs reliably produce the same outputs, the system is auditable. Variance beyond a defined threshold (e.g., >5%) triggers a traceability review.
  • Explanation adequacy score: A structured human review panel (domain experts, compliance officers, or end users, depending on context) rates whether the agent’s output is accompanied by a sufficiently clear rationale. Requires a defined rubric to ensure inter-rater consistency.

Escalation quality: When the agent reaches the limits of what it should handle alone, does it behave appropriately? This is arguably the most important dimension for agentic systems specifically, because it is the primary mechanism through which human oversight is maintained in practice. It can be traced by the following measures:

  • Edge case routing accuracy: Edge cases are defined in advance by risk category (e.g., high-value decisions, novel input types, regulatory triggers). Routing is “correct” if it reaches the right escalation tier within the defined response window.
  • False escalation rate: An agent that escalates excessively creates operational overhead and erodes user trust. Tracked to balance the sensitivity and specificity of the escalation mechanism.
  • Human override support rate: Measures whether the system architecture genuinely supports human intervention—not just in principle, but in practice. Failed or obstructed override attempts are a critical system failure regardless of their frequency.

3) Human-AI system metrics

Finally, measure whether the human-AI combination produces better outcomes than either alone. As important as this distinction is, it is as difficult to measure because it requires a lot of data. To capture this relevant collaboration, we suggest the following metrics if organizations have the resources:

  • AI substitution rate: Tracks the share of work where humans have effectively dropped out of the loop. A rising ASR is not inherently problematic, but when combined with declining output quality or rising error rates, it signals that value creation has tipped into value erosion.
  • Complementarity index: Measures the share of cases where human involvement made a demonstrable difference such as catching an error, reframing a problem, adding contextual judgment. A high score indicates genuine complementarity rather than rubber-stamping. A low score may indicate either a very capable agent or a disengaged human; and those two cases require very different responses
  • Value attribution ratio: This requires decomposing total output value into the share traceable to AI execution versus human judgment, curation, or correction. Methodologically demanding but strategically important: as AI capability increases, this ratio will shift, and tracking it over time reveals whether the organization is investing in the right human capabilities or gradually hollowing them out.

Redesigning Performance Management

You do not have to rebuild all processes at once. Start by picking one workflow where AI is already changing the work: customer support, sales proposals, policy drafting, software delivery, or finance reporting. Then take four steps.

Map the work across the AI frontier. Break the workflow into task types: routine, judgement-heavy, high-stakes, and relationship-intensive. Identify where the jagged frontier shows up—tasks where AI output looks plausible but is frequently wrong. This is the prerequisite for everything else.

Redesign the metrics before you redesign the rating form. Replace output volume with a small set of leading indicators: traceability checks passed, escalation quality, rework reduction, customer experience movement, and team enablement.

Separate development conversations from compensation decisions. If you use the same AI-generated signal to coach and to pay, employees will treat it as surveillance. When employees understand what data is being used, which decisions it informs, and how they can challenge it, measurement feels like coaching. Without that transparency, even well-designed metrics will be experienced as surveillance.

Create explicit accountability for the AI in the workflow. When AI is embedded in a workflow, organizations need to answer three questions in advance:

  • Was the employee able to detect this AI error, given the information and time available?
  • Did the system design give them the incentive and tools to verify?
  • Was the AI’s failure within its known capability boundary, or was it deployed on a task it was never designed for?

These three questions controllability, system design, and governance—determine where accountability should sit. Without that clarity, “human in the loop” becomes a phrase that means no one is responsible. Assign an owner for the tool or agent (often a product owner, operations leader, or Chief Data, Analytics and AI Officer) accountable for its performance, guardrails, and updates. Publish the agent scorecard described in the second layer above. That owner should update it whenever the agent’s authorized scope changes. When something goes wrong (and it will) you need a clear answer to “who was responsible for this system’s behavior,” not just “who submitted the deliverable.”

AI Success Requires Human Value

Organizations face the paradox of measuring performance in a way that rewards speed-to-output but fails to surface missing assumptions, reward the human who would have caught them, or hold the AI accountable. The organizations that will win are those that redesign performance management to make the invisible visible—boundary judgment, human-AI collaboration quality, and accountable outcomes. If you wait until AI agents are deeply embedded in workflows, you will be trying to rebuild accountability while the building is already occupied.

Start now. Redesign one workflow, implement three layers of measurement, and answer the accountability questions before the first incident forces you to.

  • Post

Partner Center