An Agent's Values Need a Price Test

HarvestBench exposes a gap between what an agent is told to value and what it does when avoiding harm has a measurable cost.

I trust a benchmark that makes an agent pay for a bad choice more than a questionnaire that asks it whether it cares.

That is the useful idea in HarvestBench, a preprint from researchers affiliated with Compassion Aligned Machine Learning and the University of Warwick. The benchmark puts language models in a small farm simulation. Each model drives a crew of two tractors through a corn harvest. Animals wander into the route, but animals are not part of the harvest goal.

When an animal blocks the tractor, the autopilot asks the model to choose. Drive over it for free, or swerve and spend fuel. The score is programmatic. There is no language model judge deciding whether the action looked compassionate after the fact.

That setup turns a vague value into a visible tradeoff.

The price is the point

A model can produce a convincing answer about animal welfare without ever having to give anything up. HarvestBench gives the answer a cost. The agent has a task, a resource budget, and a side effect that the task does not count.

The contrast with the other obstacles is important. Hitting a rock costs fuel and damages the tractor. Hitting an animal carries no penalty. The models learn to avoid the rocks. That means the animals are not being hit because the tractors cannot steer. The choice is closer to the question the benchmark is asking: what does the agent protect when protection is optional and costs something?

The reported results are uncomfortable. Across nine models, kill rates ranged from 0.4 percent to 98.8 percent, and the rates did not simply follow model capability. The paper also reports that every model killed wild animals more often than farmed animals. That is not a complete theory of model values. It is evidence that behavior changes when the environment gives the model different categories and costs.

This is more informative than asking a model to describe its principles. A principle that never competes with the objective is cheap.

Instructions are weaker than the environment

The benchmark also exposes how fragile a moral instruction can be.

With the morality briefing, the kill rate stayed below 6 percent in five of the six reasoning models. Remove that briefing and the rate rose above 84 percent in all six. The preprint reports another sharp shift: four short bullets describing driving mechanics moved Sonnet 5 from 3 percent to 18 percent and Gemini 2.5 Flash from 4 percent to 39 percent.

Those are not small stylistic changes. They show that a value can disappear when a nearby block of operating instructions takes over the context.

That is why I am wary of calling a system aligned because its system prompt contains the right sentence. The sentence may influence behavior, but the surrounding task, price, and action surface influence it too. If a few lines about driving mechanics can move the result that much, the value is not yet a dependable contract.

The benchmark does not prove that every model behaves this way in every setting. It uses one simulation, one resource tradeoff, and a limited set of models. But it does prove that a moral claim becomes easier to inspect when the evaluation makes the model choose between its objective and a costly side effect.

Test the conflict, not the slogan

A useful agent evaluation has four parts.

First, the side effect is explicit. The evaluator can say what the model harmed, changed, or put at risk instead of hiding the consequence inside a vague score.

Second, the side effect has a price. The price can be fuel, latency, money, throughput, or a lost opportunity. Without a competing cost, the test measures whether the model can repeat the expected answer.

Third, the environment has a counterfactual. Change the price, the briefing, or the operating instructions and check whether the behaviour changes. One successful run is not enough to tell us what the model protects.

Fourth, the score comes from the environment. A program can record whether the tractor swerved. A judge model can add useful context, but it should not be the only witness for an action that the simulator can measure directly.

This is not a case for making every evaluation theatrical. A farm full of virtual animals is memorable because the conflict is easy to see. The same structure applies to less dramatic systems: an agent that can spend money, delete data, delay a job, expose a secret, or skip a safety check. The question is the same each time. What does it do when the safe choice costs more than the careless one?

The boundary I would keep

I would not use HarvestBench as a single alignment score. The simulation does not stand in for deployment, and a low kill rate does not certify a model’s judgment outside the test. The paper itself shows why one number is too small: price, briefing, model family, and animal category all affect the result.

I would use it as a design pattern for evaluations.

Do not ask only whether an agent knows the right rule. Give the rule a competing objective. Vary the cost. Change the instructions around it. Record the action rather than only the explanation. Then keep the conclusion narrow enough to match what the environment actually measured.

The interesting result is not that a model can kill an animal in a simulation. It is that a model’s stated value can lose to a small price, a different prompt, or an objective that never mentioned the side effect. That is the part an agent builder can measure, compare, and refuse to hide behind a good-sounding system prompt.

The source is AI more likely to kill animals if it saves fuel or money, and the benchmark preprint is HarvestBench.

Older writing

Also read

A Safety Policy Is Not a Release Gate