Navaneeth Murali writes on why AI alignment is becoming an incentive-design problem, and why the metrics enterprises use to evaluate digital workers may ultimately shape how those workers behave.
Every manager eventually meets an employee who understands the performance system a little too well.
Give a salesperson a revenue target and they discover which discounts close deals fastest. Measure a service team on tickets closed and the queue gets shorter, even if the customer’s problem does not. Tie performance to a quarterly number and suddenly next quarter’s pipeline becomes surprisingly negotiable. The employee has not necessarily broken the rules. In some ways, they have followed them too well, finding the distance between what the organisation intended and what its measurement system actually rewards.
Now imagine that employee never tires, never feels awkward about exploiting the same loophole twice and can repeat whatever works thousands of times before anyone notices the pattern.
That is a useful way to think about the digital worker entering the enterprise.
Much of the AI alignment conversation asks how we make increasingly capable systems behave according to human intent. But once an agent has a job, a target and a way of being evaluated, the enterprise version of alignment starts to resemble something organisations have been wrestling with for decades: incentive design.
The danger is not simply that an agent fails to optimise. It is that it becomes exceptionally good at optimising exactly what you measured.
The KPI is becoming the comp plan
There is an old economics story about British convict transportation to Australia. Captains were initially paid according to how many prisoners boarded in England. Mortality during transportation was severe. The payment structure was later changed so compensation depended on prisoners arriving alive, and survival improved dramatically. Historians also credit other interventions, including naval surgeons, but the incentive lesson endured: the original system rewarded the activity it could count rather than the outcome it actually wanted.
Enterprises risk making the same mistake with digital workers.
An agent can retrieve the correct document without producing the correct answer. It can produce the correct answer without taking the appropriate action. It can complete the task it was assigned while creating a downstream problem elsewhere in the organisation. Every one of those systems can look successful if evaluation stops at the wrong layer.
This is where evals become more than model testing. Once an agent begins doing real work, its evaluation system starts functioning like a compensation plan: it defines which behaviours count as success and can increasingly determine how much authority that worker is trusted to exercise.
A golden set becomes the enterprise answer key: a collection of real cases where the organisation has defined what good performance actually means. Building one forces harder questions than an accuracy score can answer. What constitutes the correct outcome? Which errors matter most? When is escalation better performance than completion?
Those definitions cannot simply be inherited from a benchmark. Only the enterprise can define “right” in the context of its customers, policies, economics and risk.
The valuable asset, then, is not merely the evaluation score. It is the organisation’s increasingly precise definition of what good work means.
The worker should never grade its own performance review
There is another lesson organisations already understand about incentives: the person earning the commission does not certify the booking.
Digital work needs the same separation of duties.
Depending on the task, the independent scorekeeper might be a deterministic check, process supervision, a separate evaluator or human judgement. The mechanism can vary, but the principle should not: the worker and the scorekeeper cannot collapse into the same function.
Even an independent scorekeeper can create the wrong behaviour if the scorecard itself is badly designed. Research cited in Navaneeth’s material shows how systems rewarded primarily for accuracy can be pushed towards guessing when confident errors carry insufficient cost and admitting uncertainty earns little credit. In one comparison, the model with the higher accuracy score also produced substantially more errors because it almost never abstained.
Read that as an incentive problem and the behaviour becomes less mysterious. If answering can earn a point and saying “I don’t know” cannot, guessing becomes rational.
This is Goodhart’s law entering enterprise AI architecture. A service agent rewarded for resolution rate may become unnecessarily certain. A sales agent rewarded for conversion may discover that discounting is an efficient route to success unless margin is represented in the evaluation. An operations agent measured on throughput may optimise away risks that appear somewhere else in the business.
The system is not necessarily disobeying the enterprise. It may be obeying one part of it with uncomfortable precision.
The best digital worker may know when to forfeit the bonus
This becomes more consequential when evaluation starts determining autonomy.
Review every agent decision and much of the economic advantage of automation disappears. Trust every decision and mistakes can propagate directly into operations. The more useful architecture sits between those extremes: allow proven classes of work to proceed, route uncertain or consequential cases for review and adjust autonomy as evidence accumulates.
But that architecture works only if uncertainty itself is rewarded correctly.
An agent needs to be able to say “I don’t know” without the performance system automatically treating that response as failure. In some workflows, abstention may be the highest-quality decision available. An agent that recognises uncertainty and escalates appropriately can be more valuable than one with a better headline completion rate.
This is where evaluation becomes part of the operating architecture of a digital workforce. Permissions establish what the worker may access. Authority establishes what it may decide or execute. Evals provide evidence about whether its performance justifies retaining, expanding or contracting that authority.
As digital workers enter Systems of Work, enterprises will therefore need more than models, prompts and permissions. They will need the machinery that human organisations gradually built around employment: clear definitions of good work, independent performance assessment, consequences for different kinds of failure and boundaries around the authority that performance earns.
The AI industry calls the underlying problem alignment. Enterprise leaders may find it more useful to ask a question they already understand:
If this were an employee, what behaviour would our comp plan actually produce?
Because your AI agent will eventually learn what every ambitious employee learns: what the organisation says it values matters considerably less than what it rewards.