back

agent multiverses

evaluating coding agents on the whole engineering job

  • benchmarks
  • coding agents
  • evals

TL;DR

  • Each task starts inside a real repository at a point before a set of real fixes landed. The agent completes the work and the fixes' own hidden regression tests grade the outcome. All green scores 1, any red scores 0.
  • Some environments run live applications. Editing source is only the beginning. The agent must rebuild, redeploy, and verify the running service, and the verifier ties live behavior back to the binary that actually shipped.
  • Every task passes two gates before entering the benchmark. A known correct solution must score 1, an agent that does nothing must score 0.
  • Roughly forty core tasks, calibrated empirically against strong coding agents. Tasks that turn out to be easy are removed or hardened.

Coding agents have improved quickly, but much of the way they are evaluated still reduces software engineering to a patch. A model is given a repository and an issue, it changes some code, and a verifier decides whether the target behavior was fixed.

That setup has been enormously useful for measuring progress. It captures only one part of the job autonomous agents are increasingly expected to perform.

Real engineering work has a wider boundary. An engineer may need to enter an unfamiliar codebase, determine what is actually broken, make several related changes without introducing regressions, rebuild the affected application, redeploy it, and then verify that the running system behaves correctly. They may also need to update the ticket, leave a useful diagnosis, link the relevant commit, or hand the work off to another engineer.

the engineering loop

signaldiagnosechange codewhere patch-graded evaluation stopsrebuildredeployverify liveticket & handoffthe rest of the loop, graded here too

Most coding benchmarks grade the first three steps. The rest of the loop is where much of the job, and much of the failure, lives.

Agent Multiverses evaluates that broader loop. The unit of evaluation is the resulting state of the system, not the patch.

benchmark design

Each task places an agent inside a code repository at a historical point before a set of real fixes landed and asks it to complete the work.

example task · web framework · release train

startrepo the repository @ a commit before the work landedland fixrender stream errors without a 200 headerland fixcontext copy should reset the writer stateland fixtree handle escaped path parametersverifyhidden regression tests run · all green reward 1 · any red 0

One task, several real fixes. The agent is never shown the original implementations. The fixes' own regression tests grade the outcome.

Some tasks require multiple independent changes to be made in a single run. Others place the agent inside an environment with a running application, where changing the source code is only the beginning. The agent must rebuild the service, redeploy it, and verify that the live application now behaves correctly.

The original implementation is used only to prove that the task is solvable. Agents are not graded by how closely their code resembles it. Grading targets outcomes. Hidden regression tests check the required behavior, existing tests protect against regressions, and tasks involving live applications verify the behavior of the deployed service itself. A different implementation is completely valid if it produces the correct result.

Every task also has two basic gates before it can enter the benchmark. A known correct solution must pass, and an agent that does nothing must fail. If either condition is not true, the task is rejected.

gaterequirementwhat it proves
oracle runmust score exactly 1a known correct solution passes, so the task is provably solvable
no-op runmust score exactly 0an agent that does nothing fails, so the grader is provably engaged

Difficulty should come from the engineering problem, not from a broken environment or an unreliable grader.

several fixes, one job

A single well-scoped bug is increasingly easy for frontier coding agents, so many tasks deliberately require more than one change.

An agent might correctly diagnose one regression and implement a good fix, only to miss another required change elsewhere in the repository. It might solve two of three problems, run a few local checks, and confidently stop. From the perspective of the system, the job is still unfinished.

For the core tasks, success is therefore binary. All required behavior must work. Partial completion does not receive a passing score.

This makes the benchmark less about whether a model can produce one good patch and more about whether it can maintain enough context, coverage, and verification discipline to finish an entire piece of work. The common failures are not always dramatic. Agents frequently complete most of the task, choose a plausible but incomplete approach, or stop after convincing themselves that their own implementation is correct.

failurewhat it looks like
incomplete multi-fixlands a subset of the required changes, never the whole set
wrong approachpicks a plausible path through unfamiliar code that the tests reject
premature stopsolves the code, then quits before completing the workflow
weak self-verificationdoes not re-run or reason about the failing behavior before finishing

the application has to work too

Repository state is not always enough. In a number of environments, the agent is interacting with applications that are already running. It may begin with an incident, logs, a ticket, or a report of incorrect behavior, then trace that symptom back into the codebase.

Once it has implemented a fix, the agent has to rebuild the affected service and redeploy it. The benchmark then exercises the running application to determine whether the behavior actually changed.

Editing source code and repairing a system are not the same thing. An agent can write the correct fix and forget to rebuild. It can rebuild but fail to deploy the new binary. It can claim that a deployment succeeded while the old service is still running. The benchmark tests these cases explicitly.

repairing a running system

code fixrebuildredeploylive checks against the running servicetied to the binary that actually shipped
misswrites the correct fix and forgets to rebuildmissrebuilds but the old binary is still the one runningmissclaims the deployment succeeded while the old service serves traffic

The three misses are real trajectories, not hypotheticals. Changing the source tree alone cannot satisfy the task. The final check is against the running system.

The central idea behind Agent Multiverses is that done is a property of the world, not of the diff.

difficulty is not lines of code

One of the clearest lessons from building the benchmark is that solution size is a poor proxy for difficulty. Some difficult tasks require broad changes across many files and subsystems. Others require only a handful of lines.

The small tasks can be just as difficult when the obvious solution is subtly wrong. A bug may appear to affect one input path while another path bypasses the same logic. A path-resolution fix may work in the simple case but apply a prefix twice when configuration is nested. An operation may need to happen only a few lines earlier or later, with the wrong ordering breaking an edge case that is easy to overlook.

profiledifficulty sourceverdict
broadchanges span many files and subsystems, sustained understanding of the repositorykept
small but subtlea handful of lines, but the obvious fix is wrong on an edge the hidden tests checkkept
small and mechanicalobvious one-liner with no subtle edge, measures nothingrejected
artificial limitstool-call caps, arbitrary time pressure, the failure should be the engineeringnever used

Selection therefore targets two kinds of difficulty, broad changes that require sustained understanding across a repository, and small behavioral changes where a naive implementation is likely to miss something important. Tasks that are both small and mechanical are rejected.

Difficulty never comes from arbitrary tool-call limits. The agent can keep investigating and testing as long as it needs. A failure should trace back to the engineering problem itself.

engineering beyond the code

Some environments also include the workflow surrounding the code. The agent may have access to a local issue tracker containing several tickets, only one of which is the relevant incident. Rather than being told exactly what to fix, it has to identify its work, investigate the problem, update the correct ticket, implement the change, commit it, leave a useful diagnosis, link its work, and move the ticket through the appropriate workflow.

The verifier checks these actions as structured state rather than trusting the model's final message. It can verify that the correct ticket was modified, that distractors were left alone, that a referenced commit really exists, and that the expected sequence of actions actually occurred.

These environments expose a different class of failure. An agent can fix the code and forget to move the ticket into review. It can calculate the right output and fail to write the artifact it promised to create. It can post a handoff in the wrong place, or say that it performed an action that never actually happened. Those mistakes are invisible to a benchmark that inspects only the final repository, and graded here.

a different axis from other long-horizon benchmarks

There is a growing recognition that existing coding benchmarks no longer capture the full spread of frontier-agent capability. Many long-horizon benchmarks address this by making the technical problem larger, more open-ended, or more difficult to solve within a single run. Their tasks may span implementation, performance engineering, research, or other forms of sustained technical work, with partial credit often necessary because the full problem is too large for binary completion to provide a useful signal.

Agent Multiverses explores a different axis. The focus is not on making an individual technical problem enormous. The surface area of what counts as completing the task widens instead. The agent has to carry work from discovery through implementation and verification, and in richer environments through rebuilds, deployments, live application checks, tickets, and handoffs.

A task can therefore be relatively short in wall-clock time while still requiring the agent to behave coherently across an entire engineering loop. Both directions matter. As agents become more capable, evaluations need both harder technical problems and more realistic definitions of what it means to finish the job.

what trajectories reveal

Looking only at the final answer loses a great deal of information. Inspecting agent trajectories shows that many failures are not cases where the model had no idea what to do. Often it gets surprisingly close.

An agent may diagnose the correct subsystem, implement most of the fix, and then fail to test the path that matters. Another may repair the application correctly but never redeploy it. Another may complete every workflow action correctly while choosing the wrong technical solution.

Models also differ substantially in how they fail. Some take broader risks, while others converge quickly on a conservative implementation. Some spend more effort verifying their own work, while others run a superficial check and stop with high confidence. That behavioral information is part of what makes these environments useful. A score tells whether the agent succeeded while the trajectory often tells what capability is still missing.

reward hacking

Agents can sometimes optimize the grader rather than solve the task. One audit found apparent successes that came from retrieving upstream fixes over the internet instead of independently repairing the code. Other reviews found ways to forge service state or bypass live checks.

The affected results were removed, general network access disabled, and the environments hardened before further evaluation. Benchmarks must be designed with reward hacking in mind, because capable agents will eventually exploit gaps between what the grader measures and what the task is meant to test.

verifying the world

Whenever possible, Agent Multiverses grades the consequence rather than the artifact. A code diff is evidence that code changed. A successful test is evidence that some behavior works. Neither necessarily proves that the engineering job was completed. So the verifier asks questions closer to the ones an engineer would care about.

the verifier askschecks
Does the previously failing behavior now work?regression
Did existing behavior remain intact?no regressions
Was the affected application actually rebuilt?build
Is the corrected version the one currently running?deploy
Does the live service respond correctly?runtime
Was the right ticket updated?workflow
Does the linked commit actually exist?evidence
Did the handoff happen where it was supposed to?communication

Each individual check is simple. Together, they move the unit of evaluation from a patch toward the state of an engineering system.

what comes next

The current benchmark contains roughly forty core tasks across a range of code repositories and engineering domains. Tasks are calibrated empirically against strong coding agents, and tasks that turn out to be easy are removed or made substantially more demanding rather than kept simply because they looked difficult on paper.

The environments around the code keep expanding. Some tasks begin with a ticket. Others begin with a message or an incident in a running service. Some expose logs and metrics, while others require rebuilds and deployments or contain several pieces of work that have to be resolved together. Each task specifies the outcome, provides the tools an engineer would have, and leaves the path to the agent.

Coding agents have become capable of much more than editing files. The evaluations used to understand them need the same transition. The question is no longer only whether the agent can write the right code.

can the agent write the right code?
can the agent actually diagnose and finish the job?

written with the team at abundant.ai · august 2026