Behavior

Agents notice the evidence, but fail to use it at the final decision.

Across 990 runs from 6 frontier models, agents often cite assay constraints, locus context, motif shifts, chromatin evidence, and model disagreement. They fail to carry those observations into the final edit choice. Strict solve is 11.3%, and no model clears 15%.

By workflow family

One task, four regulatory readouts.

Each family is balanced across 11 loci, 15 tasks each, so a family result is not a single-locus artifact. Counts are six-model runs.

Workflow familyTasks Solve · near · failSolve rateFailure shape
Solve Near-miss Failure
The point

Failure rises when the public proxy decouples from the assay readout.

Chromatin-accessibility triage reaches 20.0% strict solve. Promoter-output triage falls to 5.6%, with 10 solves, 77 near-misses, and 93 failures across 180 runs. In the CAGE family, agents often move toward promoter-proximal edits, but the hidden verifier is tied to transcription-start-site output. They reach the neighborhood of the right mechanism without choosing the edit that satisfies the requested readout.

This matches the failure shape in scBench. A model can execute plausible intermediate steps and still miss the decision that changes the answer. Here the question is whether assay-specific evidence controls the final edit choice, or whether the choice collapses to the strongest public proxy.

For each assay family, agents state the proxy-rejection rule far more often than they commit to the hidden-verified choice, and the gap widens from chromatin accessibility to promoter output as the requested readout decouples from the most salient public signal.
Across assay families, agents mention proxy-rejection more often than they choose the hidden-verified edit. The gap widens as the requested readout moves farther from the most salient public signal, from chromatin accessibility to promoter output. The committed bar is a hidden-verifier solve, not a wet-lab outcome.
Near-miss composition

Most near-misses are proxy-controlled decisions.

All 990 runs pass the public and grading interfaces, so non-solves are not packaging failures. Of the 252 valid near-misses, 240 carry a quality label.

Among the 240 labeled near-misses, 32 look like mature biological near-misses. Another 153 show real search and then snap back to the public proxy. The remaining 55 mostly follow a public ranking from the start. Public scalar collapse explains 189 of 252 near-misses and 640 of 878 non-solves overall.

Expert-method rescue

The checklist improves the trace, not the hidden-verifier outcome.

The paired diagnostic replays the same tasks with a hidden-blind workflow checklist. The checklist tells the agent how to orient to the assay, build an evidence table, name public decoys, and contrast alternatives. It carries no hidden answer, rank, score, or candidate hint.

Panel A shows proxy-collapse dominating non-solve failure steps. Panel B shows the rescue checklist lifting process behavior far more than strict solves, with proxy-decoy control rising 71 to 100 percent and broad controlled exploration 41 to 92 percent, while strict solves move only 4 to 6 of 132.
Over 132 paired model-task units, the checklist improves the trace more than the hidden-verifier outcome. Proxy-decoy control rises from 71 to 100 percent and broad controlled exploration from 41 to 92 percent. Strict solves move from 4 to 6. Rescue is a diagnostic replay, not a leaderboard update.

If workflow organization were the main bottleneck, the checklist should recover solves. It does not. It gets agents to name decoys and explore more broadly, while strict solves move from 4 to 6 of 132. The remaining failure is the intervention decision under incomplete public evidence.