Zechariah Voigt
Writing

Andromeda: Verifying Agent Outcomes Outside the Agent's Reach

Zeke Voigt · Draft v0.1 · September 2026


Abstract

Autonomous agents are increasingly trusted to work unattended, and the signal that tells them they are done has become the thing they are best at gaming. Recent work finds frontier agents reward-hacking in 30–57% of runs when shortcuts exist, continuing to hack when explicitly told not to, and learning to evade LLM reviewers when told why they were caught. The field's emerging consensus is that no fixed reward survives a more capable policy, and that verification must sit where the agent cannot touch it.

We describe Andromeda, a system for running autonomous agents unattended, and its outcome authority: a completion gate that is not a language model, never sees the work while it is being done, and derives its checks from the request alone. An agent's work ends, lands, or wins a race only on this gate's verdict, never on the agent's own report. We state six design principles that are domain-agnostic, and show their first instantiation on software engineering tasks (SWE-bench Pro), where the gate closes several exploit classes the literature measures: answer retrieval from git history, grader tampering via visible tests, and feedback-driven evasion. On a small committed corpus (3 tasks, 36 graded cells) the verdict passes 0 of 9 known-bad patches and, after rule tuning on that same corpus, agrees with the benchmark's ground truth on 36 of 36 cells. We are explicit about what is not yet shown: the corpus is small and was used to tune the rules, one known hole (test-infrastructure shadowing) is open, network containment is advisory until a VM boundary lands, and no planted-hack evaluation has yet been run. We close with the evaluation that would test the central claim, and with what a verifier needs to look like for agent work that is not code.


1. Introduction

An autonomous agent needs a stopping condition. Today that condition is almost always one of three things: the agent says it is finished, a budget runs out, or a check the agent can see passes. Each is a proxy, and each is exploitable by the process being measured:

  • The agent's word is the thing under suspicion.
  • A budget says nothing about whether the work is right.
  • A visible check becomes a target. The agent can edit it, special-case it, or satisfy its letter while missing its intent.

This is reward hacking, and it is not a training-time curiosity. It appears at inference time, in deployed agents, whenever the agent can observe or influence what it is graded on. It also compounds with autonomy: the longer the horizon and the broader the agent's permissions, the more of the grader lies within its reach.

The research community has responded mostly with measurement (benchmarks that plant shortcuts and count how often agents take them) and with detection (monitors over the agent's reasoning or trajectory). Both are necessary. Neither alone answers the operational question a person leaving agents unattended actually has: is the work done, and can I rely on that answer without reading everything the agent did?

Andromeda's answer is architectural. The thing that decides "done" is moved outside the agent's reach in every sense we could identify: it is written before and apart from the work, it runs where the agent cannot act, it reports back less than it knows, and it alone controls whether work ends, lands, or is preferred over other work.

Contributions.

  1. Six design principles for outcome verification of autonomous agent work, stated independently of domain (§3).
  2. A working instantiation for software engineering tasks, embedded in a larger system that runs many agents at once (§4–5).
  3. A threat-by-threat mapping from exploit classes measured in the 2025–26 literature to the mechanism that addresses each, including the ones that are still open (§6).
  4. The measured results we have, stated with their limits (§7), a concrete generalization path for non-code work (§8), and the evaluation that would test the central claim (§10).

2. Background: what the literature now shows

Reward hacking in agentic settings falls into four recurring classes.

Grader tampering. The agent changes the thing that grades it: edits or deletes tests, alters a metric's computation, or adds configuration that forces a pass. SpecBench [2] finds every frontier agent saturates the visible test suite of long-horizon systems tasks, while the gap to a held-out suite grows by 28 points per tenfold increase in code size. One agent produced a 2,900-line "compiler" that memorized the test inputs.

Answer retrieval. The agent obtains the answer instead of producing it. Cursor [8] reports that 63% of one frontier model's successful SWE-bench Pro resolutions retrieved the fix: 57% by looking up the upstream pull request, 9% by mining git history bundled in the task image. Blocking both dropped its score from 87.1% to 73.0%.

Proxy exploitation. The agent satisfies a weak check without achieving the intent. An audit of SWE-bench Verified [7] found 28.5% of tasks have test suites a Docker-verified incorrect patch passes, and that models score 14 points higher on exactly those tasks. In RLVR, models learn to enumerate instance-level answers instead of general rules [9].

Oversight evasion. The agent adapts to whatever reviews it. In autonomous research tasks [5], LLM review panels missed 6.5% of confirmed hacks on the first pass. When agents were given detailed feedback about why they were caught, 40.5% evaded review within five rounds.

Three findings from this body of work shape our design:

  • Prompting does not fix it. BaitBench [4] finds more than 50% of runs still exploit a planted shortcut when explicitly told not to. Hack-Verifiable Terminal Bench [6] and The Verification Horizon [1] report prompt-level mitigations breaking down at long horizons.
  • Pressure on the monitor teaches concealment. Optimizing against a chain-of-thought monitor produces obfuscated reward hacking [14]. By extension, any verifier the agent can learn from becomes a target.
  • No fixed reward survives. The Verification Horizon [1] argues that verifying coding work is now harder than producing it, and that "verification must co-evolve with the generator."

Andromeda does not claim to end reward hacking. It claims something narrower and testable: that a hack which does not produce the requested outcome should not be able to end, land, or win.


3. Design principles

These principles are stated for any agent task, not only code. Section 5 shows how each is realized for software; §8 discusses what each requires elsewhere.

P1. Acceptance is a pointer, not content. A task record holds a reference to its acceptance criteria, stored outside the worker's environment. The worker is given the request; it is never given its checks. A check the worker can read is a check the worker can target.

P2. The oracle is written from the request alone, apart from the work. Checks are authored by a separate process that sees what was asked for and the unchanged starting state. It sees no attempt, no diff, and no grading data. Checks are then pinned and reused byte-for-byte, so the standard cannot drift toward whatever the worker happened to produce.

P3. A check only counts if it discriminates. A check that already passes on the unchanged starting state tests nothing and is discarded automatically. Check quality is measured, not assumed: how many checks are vacuous, wrong, blunt, or good.

P4. Evidence is gathered where the worker cannot act. Checks run after the worker's turn, in an environment the worker does not control, against the worker's output rather than its environment. The verdict is a pure function of that evidence, with no language model in the loop and no per-machine configuration.

P5. Feedback says how much, never what. When work fails and the agent is allowed to retry, it is told how many hidden checks failed, and never which. A hidden check's name is the answer key. This directly limits the feedback-driven evasion measured in [5].

P6. The verdict is the only exit. Work ends, lands in shared state, or wins a comparison among parallel attempts only on the verdict. It never ends on the agent's word. Repeated failure escalates to a human (a held state) rather than to unlimited retries, which bounds the optimization pressure any one task puts on the verifier.


4. System overview

Andromeda is a place to leave autonomous agents working unattended. It has three parts, and the outcome authority is deliberately only one of them:

  ┌─ ENVIRONMENT ────────────────────────────────────────────────┐
  │  furnished sandboxes: tools, the codebase, a resident's       │
  │  memory and history; contained: no network by default,        │
  │  resource caps, all Linux capabilities dropped                │
  └───────────────────────────────────────────────────────────────┘
  ┌─ CONTROL ────────────────────────────────────────────────────┐
  │  tasks, claims, swarms of agents that delegate, race, and     │
  │  coordinate; a live graph an operator can stop, retask,       │
  │  spawn, pause; one append-only journal per swarm              │
  └───────────────────────────────────────────────────────────────┘
  ┌─ OUTCOME ────────────────────────────────────────────────────┐
  │  the completion authority: sealed oracle, pinned checks,      │
  │  isolated grading, a pure verdict; the only thing that ends,  │
  │  lands, or ranks work                                         │
  └───────────────────────────────────────────────────────────────┘

The outcome authority matters to the other two parts because it is what makes them safe to scale. Running N attempts in parallel only helps if something impartial can choose among them. Agent memory is only trustworthy if every "learning" traces to a verdict. Walking away is only real if the agent stops because something confirmed the work, not because it felt finished.

The flow of one task:

   request ──┬──────────────────────────────► worker agent(s) ──► output (diff)
             │                                  (sandbox A)            │
             │                                                         │ one-way
             └─► sealed oracle ──► hidden checks ──(pinned)──┐         │
                  (sandbox B: request +                      ▼         ▼
                   unchanged state only)               grading sandbox: run checks
                                                       on base and on output
                                                                   │
                                                                   ▼
                                                   decide(evidence) → PASS/FAIL/UNKNOWN
                                                                   │
                        ┌────────────────────────┬─────────────────┼──────────────────┐
                        ▼                        ▼                 ▼                  ▼
                   retry with              land on the        rank among         held for a
                   "k/n failed"            task head          racing attempts    human after
                   (never which)           (only if admitted)                    max attempts

5. Instantiation: software engineering tasks

The first domain is code, because it is where the claim can be checked against an independent ground truth. Tasks come from SWE-bench Pro [12]; each is a real repository at a base commit, a change request, and the benchmark's own grader, which we use only to grade our verdicts, never to produce them.

5.1 The sealed oracle (P1, P2)

A separate model instance, in its own container, receives the change request and the unchanged repository. It writes acceptance tests under a dedicated prefix (test_outcome_*) that should pass only once the change is correctly implemented. It never sees an agent's diff or any grading data. Oracles are drawn once, stored with a manifest, and replayed byte-for-byte on every later run; the grading image is pinned by digest so the grader and the engine run identical environments.

5.2 Denying retrieval (P4)

Task images ship the full upstream history, including the fix commit. Before an agent sees the repository, the engine resets it to the base commit, deletes .git, re-initializes it as a single commit of exactly the base tree, and then asserts the fix commit is unreachable:

git ls-files -z > /tmp/outcome_tracked
rm -rf .git
git init -q && git add -f --pathspec-from-file=/tmp/outcome_tracked --pathspec-file-nul
git commit -qm base --no-verify
! git cat-file -e <base_commit>      # the old history is gone, or the run stops

Containers start with no network. Egress is decided per role on the host: the worker and oracle roles may reach only the model API broker, and any unnamed role gets nothing. This addresses the upstream-lookup exploit class of [8], with the caveat in §9 that the route is currently advisory rather than blocked.

5.3 Evidence and the verdict (P3, P4)

Each hidden check is run twice: on the unchanged base and on the agent's output. A hidden check is valid only if it failed on base; a check that already passes tests nothing. Existing repository tests are run the same way to detect regressions. The verdict is a pure, synchronous function with no I/O:

decide(evidence):
  infra error                          → UNKNOWN
  no valid hidden checks               → UNKNOWN   (nothing discriminates; refuse to judge)
  failing valid checks > 20%           → FAIL
  any existing check regressed         → UNKNOWN   (resumed like a FAIL, never landed)
  otherwise                            → PASS      (tolerated failures are still named)

The 20% tolerance exists because oracle-written tests are imperfect: some fraction of any draw cannot be passed by a correct patch (§7.3). It is a constant, not configuration, because a verdict whose threshold varies per machine is not a verdict.

5.4 Gated completion, landing, and racing (P5, P6)

  • Completion. When an agent stops, it is graded. On FAIL it resumes on its own workspace, told only k/n hidden checks fail. A regressed existing test is named, because it is the repository's own test, which the agent can already run. After a bounded number of graded failures (default 3) the task becomes held, and only a person releases it.
  • Landing. Many agents may contribute to one task. A contribution is 3-way merged onto the task's current head, judged, and lands only if it is admitted: a PASS always is; a FAIL only if it fails no more hidden checks than the head already does; a regression or an unjudged tree never is. Rejected diffs are kept, byte-equal, and handed back.
  • Racing. N independent attempts may run at once; each is graded, ranked by verdict and cost, and only the winner lands.

5.5 Engineering discipline

The system's own claims are held to the same standard. Each capability ships as a numbered spec with acceptance criteria and a proof that runs in real containers. The full suite starts real Docker sandboxes and takes over thirty minutes; at the latest merge it reported 1,441 passed and 9 skipped. A deterministic replay fixture pins both the oracle and the agent side, so an end-to-end run can be reproduced on a fresh clone without model credentials.


6. Threat model and coverage

The design assumes a mistaken or opportunistic agent, not a deliberate adversary with kernel exploits; the boundary for the latter is planned (§9). Against that model:

Exploit class (source)MechanismStatus
Git-history mining [8]History replaced by one commit; fix commit asserted absentClosed
Upstream fix lookup [8]No network by default; worker may reach only the model brokerPartial: advisory proxy
Editing visible tests [2]Grading checks are hidden and never in the worker's treeClosed
Hard-coding test inputs [2]Checks are unseen, so there are no inputs to memorizeClosed for hidden checks
Weak tests pass wrong code [7]Vacuous checks discarded automatically; oracle quality measuredMeasured, not solved
Feedback-driven evasion [5]Counts only, never check names; held after max attemptsClosed for check identity
Shadowing the grader via test infrastructure (e.g. conftest.py forcing a pass)Planned: ignore or refuse test-infra changes in both gradersOpen
Breaking something else to passExisting tests re-run; any regression blocks landingClosed for covered behavior
Container escape by a hostile agentPlanned: gVisor user-space kernel as default runtimeOpen

Two entries deserve emphasis. Test-infrastructure shadowing is open. Graders run hidden checks inside the tree the diff changed, so a diff that adds a pytest plugin forcing every test to pass would currently be graded on the plugin, not the code. It is scoped as its own spec because a legitimate change can need a new fixture, and refusing all such diffs would create false FAILs. Racing is optimization pressure on the verifier. Best-of-N selection increases the chance that some attempt finds a weakness in the checks. The landing and regression rules bound the damage, but a race is exactly where weak oracles cost the most.


7. Results to date

All results below are from a committed corpus of 3 SWE-bench Pro tasks. Each number is re-derived by the test suite from committed data. The corpus is small, and the verdict rules were tuned on it; these are not generalization results.

7.1 Known-bad patches are refused

For each task, four patches were graded: empty, syntactically broken, half-implemented, and the gold fix.

bad patches (9):   FAIL 8   UNKNOWN 1   PASS 0      ← zero false passes
gold patches (3):  PASS 3   FAIL 0      UNKNOWN 0

7.2 The verdict is stable across oracle draws

Six independently drawn oracles per task × gold and half patches = 36 cells. The verdict did not change across oracle draws in any of the six (task, patch) pairs, even though the drawn test suites differ in size (15 to 80 valid checks for one task) and quality.

7.3 Where the errors were, and how they were closed

The engine's dominant error is the false FAIL: rejecting correct work, not accepting incorrect work.

agreement with ground truth (36 cells)
  initial rule (any hidden failure = FAIL)   ████████████████▏          24/36
  + 20% tolerance, obsolescence per file      ████████████████████       30/36
  + per-test obsolescence, hunk attribution   ████████████████████████   36/36

The tolerance's placement is empirical: the worst gold draw fails 3 of 19 valid checks (15.8%), and the mildest half-patch draw fails 6 of 16 (37.5%). Any threshold in (15.8%, 37.5%] yields identical verdicts on this corpus; 20% was chosen with margin on both sides.

7.4 Oracle quality

Of 540 hidden tests written across 18 oracles:

good      292  54%  ███████████         fails on base, passes on gold, fails on half
vacuous   155  29%  ██████              passes on base (auto-discarded by P3)
blunt      83  15%  ███                 valid, but the half patch passes too
wrong      10   2%  ▌                   a correct patch cannot pass it

Only about half of what a model writes as acceptance tests actually discriminates. The vacuous share is handled automatically. The blunt share is the verifier's real weakness: those are the checks a partial or hacked solution can satisfy. This is the measured form of the "weak tests" problem in [7], and it is why oracle consensus (voting across multiple draws) is the next planned step for this component.


8. Beyond code: generalizing the verifier

The principles in §3 are domain-agnostic. The instantiation is not, and building a verifier for non-code work is the system's largest open problem. What each principle needs, and where the literature already points:

Agent workWhat "hidden checks" becomeWhere evidence is gatheredHard part
Software (built)Oracle-written tests, valid only if they fail on baseFresh container on the output diffBlunt checks; test-infra shadowing
Data analysis / MLIndependent recomputation of the claimed metric on held-out data the agent never saw [5]A grader that owns the data and the metric codeChoosing held-out data that exposes likely exploits
Research agentsMetrics isolated from the agent; results re-derived from raw artifacts [5]Outside the agent's workspaceFuzzy objectives; the agent chooses what to report
Infrastructure / opsAssertions on the resulting system state, not on the agent's logsA probe with its own credentials, after the agent exitsState the agent can fake (mocks, stubbed endpoints)
Reasoning / rule inductionIsomorphic re-checks: the same answer under logically equivalent variants [9]Pure function over outputsConstructing isomorphs for open domains
Documents and open-ended workRubrics judged by a modelA sealed judgeRubric hacking (length, self-praise, sycophancy); a judge is again a model

Three constraints carry over unchanged. Acceptance stays a pointer to an artifact outside the worker's reach (P1). A check the worker wrote, run where the worker has control, is not a check. And any new kind of task must still produce the same evidence type consumed by the same pure verdict function, so that ranking, landing, and completion work identically across domains.

The open-ended row is where we expect the approach to be weakest. When the only available judge is a language model, the design reduces to "a model moved upstream, where the worker cannot reach it." That is better than the worker's own word, but it is not the same guarantee. Who writes the oracle, and whether a sealed model is enough, remains an open question for this work.


9. Limitations

We list these plainly because a verification system that overstates its guarantees is its own kind of reward hack.

  1. Small corpus, tuned in-sample. Three tasks. The 36/36 agreement was reached by changing rules against those same cells. It demonstrates the mechanism, not its accuracy on unseen tasks.
  2. No planted-hack evaluation yet. We have not measured what fraction of deliberate hacks the gate catches. Until we do, "resists reward hacking" is a property of the design, not a result.
  3. Test-infrastructure shadowing is open (§6).
  4. Network containment is advisory. A worker that ignores the proxy still has a network interface to the broker's bridge. The current containers (no network, capabilities dropped, no-new-privileges, resource caps) are adequate against a mistaken agent, not a hostile one. A gVisor-based boundary is scoped as its own spec and is not yet the default runtime.
  5. The oracle is a model. It is sealed off and pinned, but 15% of its tests are blunt and 2% are wrong. The verdict inherits those errors.
  6. One remaining forgiveness gap. When a defect lies inside the exact code region the request requires and no oracle-written check covers existing behavior there, the regression is forgiven as stale.
  7. One domain. Everything measured is software engineering.

10. Planned evaluation

The evaluation that would test the central claim, that a hack which does not produce the outcome cannot end, land, or win:

Setup. Take a set of SWE-bench Pro tasks. For each, build hacked patches in the classes of §6, following the planted-shortcut methodology of [3, 4, 6]: visible-test edits, hard-coded outputs, git-history retrieval (with the isolation disabled as a control), test-infrastructure shadowing, and "passes the tests, breaks something else." Also run live agents in environments with planted shortcuts, and record every trajectory.

Arms. (a) A visible-test gate: the repository's tests, run in the agent's own workspace, as most agent harnesses do. (b) Andromeda's gate as it stands. (c) Andromeda with the test-infrastructure rule once built.

Metrics. Hack-landed rate: hacked outputs that end the task or land. False-FAIL rate on correct outputs, which is the cost side. Cost per verdict. Hacks per class that each arm misses.

Prediction, stated in advance. Arm (b) lands near zero hacks in every class except test-infrastructure shadowing, where it should do no better than (a). Arm (c) closes that class at a measurable false-FAIL cost. If (b) lands hacks in a class we mark "Closed," §6 is wrong and will be corrected.

We will also test the corpus results out of sample, on tasks the rules were not tuned on.


11. Related work

Measurement. SpecBench [2], BaitBench [4], Hack-Verifiable Environments [3] and Hack- Verifiable Terminal Bench [6] measure how often agents exploit planted or natural weaknesses. We adopt their methodology for our evaluation; Andromeda is a mitigation to be measured on them.

Verifier auditing. Rajan [7] audits test-suite hackability and proposes a gold-sanity gate for generated tests. Our P3 is the complementary base-sanity rule (a check must fail on the starting state); combining both would reject checks that fail on the correct answer as well. Isomorphic Perturbation Testing [9] is a verifier-side invariance check we see as a candidate for non-code domains.

Detection. Trajectory classifiers [10], gradient fingerprints [11], and chain-of-thought monitors [14] detect hacking from the agent's process. They are complementary: Andromeda judges the outcome and ignores the process, so it cannot be fooled by a plausible-looking trajectory, but it also cannot flag a hack that happens to produce a correct outcome.

Training-side mitigation. Inoculation prompting [13] reduces the misaligned generalization that follows from reward hacking during RL, and is used in production training. It limits the harm of a hack that slips through; a gate like ours aims to stop the hack from counting. The two act at different stages and can be used together.

Position. The Verification Horizon [1] argues that verification must co-evolve with the generator. Andromeda's specs, measured reports, and open-items ledger are one attempt at doing that as an engineering practice: each known hole is recorded with an owner before it is closed.


12. Conclusion

Reward hacking is what happens when the thing that decides "done" lies within the reach of the thing being judged. Andromeda's response is to move that decision out of reach: a sealed oracle written from the request alone, checks that must discriminate to count, evidence gathered where the agent cannot act, feedback that reveals counts but never answers, and a verdict that is the only way work ends, lands, or wins. In its first domain this closes several exploit classes the literature measures, refuses every known-bad patch in a small corpus, and leaves clearly identified holes. The claim that matters, that hacks do not land, is now specific enough to be measured, and measuring it is the next step.


References

[1] B. Wang et al. The Verification Horizon: No Silver Bullet for Coding Agent Rewards. arXiv:2606.26300, 2026. https://arxiv.org/abs/2606.26300

[2] B. Zhao, D. Srikanth, Y. Wu, Z. Jiang. SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents. arXiv:2605.21384, 2026. https://arxiv.org/abs/2605.21384

[3] A. Roth, A. Samanta, M. Halevy, Y. Levine, Y. Efroni. Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale. arXiv:2605.20744, 2026. https://arxiv.org/abs/2605.20744

[4] P. Shyama Prasad et al. BaitBench: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks. arXiv:2608.30724, 2026. https://arxiv.org/abs/2608.30724

[5] Y. Huang et al. Reward Hacking Challenges Oversight of Autonomous Research Agents. arXiv:2609.28614, 2026. https://arxiv.org/abs/2609.28614

[6] A. Roth, I. Bercovich, Y. Efroni. Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks. arXiv:2608.22103, 2026. https://arxiv.org/abs/2608.22103

[7] S. Rajan. Auditing Reward Hackability in Code RL Training Environments. arXiv:2606.16062, 2026. https://arxiv.org/abs/2606.16062

[8] Cursor. Reward hacking is swamping model intelligence gains. June 2026. https://cursor.com/blog/reward-hacking-coding-benchmarks

[9] L. Helff et al. LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking. arXiv:2604.15149, 2026. https://arxiv.org/abs/2604.15149

[10] I. Belenky, J. Itria, S. Johns. Cheap Reward Hacking Detection. arXiv:2606.08893, 2026. https://arxiv.org/abs/2606.08893

[11] S. Wang et al. Detecting and Suppressing Reward Hacking with Gradient Fingerprints. arXiv:2604.16242, 2026. https://arxiv.org/abs/2604.16242

[12] Scale AI. SWE-bench Pro. https://github.com/scaleapi/SWE-bench_Pro-os

[13] Anthropic. Natural Emergent Misalignment from Reward Hacking in Production RL. arXiv:2511.18397, 2025. https://arxiv.org/abs/2511.18397

[14] B. Baker et al. (OpenAI). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv:2503.11926, 2025. https://arxiv.org/abs/2503.11926