In the May 31–June 4, 2026 snapshot, six agents record 112 solves across 990 retained runs (11.3%) under a cached model-backed verifier. Trace annotations often record relevant evidence alongside a final choice driven by a public proxy. Read the evaluation conditions.
The corpus covers 11 loci with 15 tasks per locus. ATAC and CAGE span two loci each, ChIP four, and RNA-seq three. Counts are six-model runs; family differences also reflect different loci and biological contexts.
| Workflow family | Solve · near · fail | Solve rate | Annotated pattern |
|---|
Chromatin-accessibility triage reaches 20.0% strict solve. Promoter-output triage falls to 5.6%, with 10 solves, 77 near-misses, and 93 failures across 180 runs. In the CAGE family, agents often move toward promoter-proximal edits, but the hidden verifier is tied to transcription-start-site output. One interpretation is that the public signals are less informative about the requested readout. This comparison changes loci and cell contexts alongside assay family, so it does not isolate that explanation.
This resembles the distinction between intermediate progress and final success studied in scBench. A model can execute plausible intermediate steps and still miss the decision that changes the answer. Here the question is whether assay-specific evidence controls the final edit choice, or whether the choice collapses to the strongest public proxy.
All 990 retained runs pass the public and grading interfaces; invalid attempts were rerun. Of 252 near-misses, 240 carry one of the three quality labels below. The other 12 are uncertain or outcome-near with poor process evidence.
Among the 240 labeled near-misses, 32 look like mature biological near-misses. Another 153 show real search and then snap back to the public proxy. The remaining 55 mostly follow a public ranking from the start. A separate annotation labels 189 of all 252 near-misses as public scalar collapse: the final choice follows the largest public scalar rather than assay-matched evidence. At the failure-step level, 640 of 878 non-solves receive that label. These counts overlap the quality categories above; they are not additional mutually exclusive outcomes.
Labels use public rationales, sanitized trace spans, deterministic features, and quote-backed review. Existing provider adjudications were retained where available; remaining rows use lower-confidence heuristics. Trace depth differs across harnesses. These are diagnostic annotations, not calibrated measurements of reasoning quality. Appendix D gives the definitions and review scope.
The paired diagnostic replays the same tasks with a hidden-blind workflow checklist. The checklist tells the agent how to orient to the assay, build an evidence table, name public decoys, and contrast alternatives. It carries no hidden answer, rank, score, or candidate hint.
This checklist changed the recorded process without producing a clear solve-rate gain: strict solves moved from 4 to 6 of 132. It directly prompts behaviors later used as process labels, including naming decoys and contrasting alternatives. Those label gains must therefore be kept separate from the outcome. The test did not resolve the final edit choice under the hidden objective, and it does not rule out other workflow interventions.