In August I re-ran HAGD, our January circuit-extraction preprint, from the released notebooks, with five seeds, an accuracy gate and three controls, and published the code and results on 21 August. This is what came out.
What runs
The whole pipeline runs on GPT-2 Small fine-tuned for addition modulo 23, on one Tesla T4, in 22.7 minutes end to end. The search itself is the cheap part: 0.09 seconds to narrow 413 active features to a 10-feature circuit. For seed 0 the fine-tuned model reaches 83.8% on held-out prompts. The other four seeds failed the 80% accuracy gate, so only seed 0 went on to circuit extraction.
What the circuit is worth
The test: replace every feature outside a set with its mean at the generation steps, and measure how much of the model's behaviour survives. The repository's control comparison puts the HAGD circuit next to three alternatives.
Behaviour kept at generation steps, seed 0. The full model is 100%.
results/controls_seed0.json). Restricting every position, prompt reading included, drops the circuit to 0%.And the larger numbers in the preprint, 91% behavioural preservation, runs on Pythia and Llama models up to 70B parameters, and cross-architecture transfer, are not in the released code at all. The repository's automated claim audit marks each of them unsupported.
Where it breaks
- Candidate generation is the bottleneck. Of the 10 features with the largest individual necessity scores, 7 never entered the candidate pool. Traversal recall is 20 to 31% across budgets.
- Necessity scores carry little signal here. The 10 least necessary features preserve as much behaviour as the circuit does.
- The first answer token carries the work. Restricting only that token already leaves 30% of the behaviour, not far from the 23% left when every generation step is restricted.
- Features act together. Some add more alongside others than alone, so ranking them one at a time undersells them.
What I changed
The project page reports only numbers the released code produces. The README prints the audit next to the claims it checks. I am revising the preprint to match. The next experiments follow from the diagnostics: fix recall before improving the scoring, measure sufficiency where the computation happens, and select features jointly.
It was not the result I wanted to publish. It is the most useful thing I did this year.