A Theory on Becoming an Expert
Becoming an expert means deliberately building the mental architecture to judge, question, and understand what models generate.
Essays, working notes, and research reports. For inference and training walkthroughs, see Systems recipes.
Becoming an expert means deliberately building the mental architecture to judge, question, and understand what models generate.
A proposal to train browser models that predict action consequences and test whether their learned dynamics transfer to new goals.
Production agents need shared learning loops where failures become reusable experience across the network instead of one-off human patches.
Institutional knowledge becomes dynamic when every diff, decision, and correction is searchable, reviewable, and available at the moment of use.
Tracing valid continuations in Tiny Aya reveals a gap between next-token control and full-answer repair on Japanese and Korean tasks.
A Qwen3-0.6B adaptation learned an interleaved-thought interface, but blank or generic controls still outperformed its own thoughts on suffix reward.
Models distinguish valid from invalid critique, but reviewer-panel pressure can erase that distinction. Internal readouts transferred; targeted steering lacked specificity.
On KernelBench, correctness-filtered surprisal selection found rare fast kernels, while ranking by surprisal before execution did not reduce evaluation cost.
Across 40 simulated episodes per model, eight security agents took incorrect containment actions in 45–97.5% of episodes. Correct actions alone hid over-triggering.
SFT1 reached 60% Pass@4 versus 59% for GRPO in one Tau2-bench comparison, while GRPO improved sampled Pass@1. Deployment depends on an available verifier.
Online evaluations turn production traces into verified frontier tasks with calibrated difficulty, while anchor sets keep progress comparable as agents improve.