Across 990 runs from 6 frontier models, agents often cite assay constraints, locus context, motif shifts, chromatin evidence, and model disagreement. They fail to carry those observations into the final edit choice. Strict solve is 11.3%, and no model clears 15%.
Each family is balanced across 11 loci, 15 tasks each, so a family result is not a single-locus artifact. Counts are six-model runs.
| Workflow family | Solve · near · fail | Solve rate | Failure shape |
|---|
Chromatin-accessibility triage reaches 20.0% strict solve. Promoter-output triage falls to 5.6%, with 10 solves, 77 near-misses, and 93 failures across 180 runs. In the CAGE family, agents often move toward promoter-proximal edits, but the hidden verifier is tied to transcription-start-site output. They reach the neighborhood of the right mechanism without choosing the edit that satisfies the requested readout.
This matches the failure shape in scBench. A model can execute plausible intermediate steps and still miss the decision that changes the answer. Here the question is whether assay-specific evidence controls the final edit choice, or whether the choice collapses to the strongest public proxy.
All 990 runs pass the public and grading interfaces, so non-solves are not packaging failures. Of the 252 valid near-misses, 240 carry a quality label.
Among the 240 labeled near-misses, 32 look like mature biological near-misses. Another 153 show real search and then snap back to the public proxy. The remaining 55 mostly follow a public ranking from the start. Public scalar collapse explains 189 of 252 near-misses and 640 of 878 non-solves overall.
The paired diagnostic replays the same tasks with a hidden-blind workflow checklist. The checklist tells the agent how to orient to the assay, build an evidence table, name public decoys, and contrast alternatives. It carries no hidden answer, rank, score, or candidate hint.
If workflow organization were the main bottleneck, the checklist should recover solves. It does not. It gets agents to name decoys and explore more broadly, while strict solves move from 4 to 6 of 132. The remaining failure is the intervention decision under incomplete public evidence.