Behavior

Public evidence does not reliably carry through to the final choice.

In the May 31–June 4, 2026 snapshot, six agents record 112 solves across 990 retained runs (11.3%) under a cached model-backed verifier. Trace annotations often record relevant evidence alongside a final choice driven by a public proxy. Read the evaluation conditions.

By workflow family

Solve rates differ across assay families.

The corpus covers 11 loci with 15 tasks per locus. ATAC and CAGE span two loci each, ChIP four, and RNA-seq three. Counts are six-model runs; family differences also reflect different loci and biological contexts.

Workflow familyTasks Solve · near · failSolve rateAnnotated pattern
Solve Near-miss Failure
The point

Expressed uncertainty exceeds strict solves in every family.

Chromatin-accessibility triage reaches 20.0% strict solve. Promoter-output triage falls to 5.6%, with 10 solves, 77 near-misses, and 93 failures across 180 runs. In the CAGE family, agents often move toward promoter-proximal edits, but the hidden verifier is tied to transcription-start-site output. One interpretation is that the public signals are less informative about the requested readout. This comparison changes loci and cell contexts alongside assay family, so it does not isolate that explanation.

This resembles the distinction between intermediate progress and final success studied in scBench. A model can execute plausible intermediate steps and still miss the decision that changes the answer. Here the question is whether assay-specific evidence controls the final edit choice, or whether the choice collapses to the strongest public proxy.

Expresses proxy uncertaintyStrict solve
Rates from Table 1 of the paper, each divided by its assay family’s retained runs. Expressing proxy uncertainty is a trace annotation; strict solve is a model-verifier outcome. The gap does not by itself identify why the agent failed. Bars share a 0–100% scale.
Near-miss composition

Most near-misses are proxy-controlled decisions.

All 990 retained runs pass the public and grading interfaces; invalid attempts were rerun. Of 252 near-misses, 240 carry one of the three quality labels below. The other 12 are uncertain or outcome-near with poor process evidence.

Among the 240 labeled near-misses, 32 look like mature biological near-misses. Another 153 show real search and then snap back to the public proxy. The remaining 55 mostly follow a public ranking from the start. A separate annotation labels 189 of all 252 near-misses as public scalar collapse: the final choice follows the largest public scalar rather than assay-matched evidence. At the failure-step level, 640 of 878 non-solves receive that label. These counts overlap the quality categories above; they are not additional mutually exclusive outcomes.

Labels use public rationales, sanitized trace spans, deterministic features, and quote-backed review. Existing provider adjudications were retained where available; remaining rows use lower-confidence heuristics. Trace depth differs across harnesses. These are diagnostic annotations, not calibrated measurements of reasoning quality. Appendix D gives the definitions and review scope.

Expert-method rescue

The tested checklist changes process labels more than solves.

The paired diagnostic replays the same tasks with a hidden-blind workflow checklist. The checklist tells the agent how to orient to the assay, build an evidence table, name public decoys, and contrast alternatives. It carries no hidden answer, rank, score, or candidate hint.

Panel A shows proxy-collapse dominating non-solve failure steps. Panel B shows the rescue checklist lifting process behavior far more than strict solves, with proxy-decoy control rising 71 to 100 percent and broad controlled exploration 41 to 92 percent, while strict solves move only 4 to 6 of 132.
Over 132 paired model-task units, the checklist improves the trace more than the hidden-verifier outcome. Proxy-decoy control rises from 71 to 100 percent and broad controlled exploration from 41 to 92 percent. Strict solves move from 4 to 6. Rescue is a diagnostic replay, not a leaderboard update. Open the full-resolution figure.

This checklist changed the recorded process without producing a clear solve-rate gain: strict solves moved from 4 to 6 of 132. It directly prompts behaviors later used as process labels, including naming decoys and contrasting alternatives. Those label gains must therefore be kept separate from the outcome. The test did not resolve the final edit choice under the hidden objective, and it does not rule out other workflow interventions.