Mohammed Mudassir Uddin

BlogAugust 2026interpretabilityreproducibility

I audited my own paper

HAGD's search is fast. The circuit it found in GPT-2 Small kept no more behaviour than random features, and the preprint's larger results were not in the code. What the audit found, and what I changed.

In August I re-ran HAGD, our January circuit-extraction preprint, from the released notebooks, with five seeds, an accuracy gate and three controls, and published the code and results on 21 August. This is what came out.

What runs

The whole pipeline runs on GPT-2 Small fine-tuned for addition modulo 23, on one Tesla T4, in 22.7 minutes end to end. The search itself is the cheap part: 0.09 seconds to narrow 413 active features to a 10-feature circuit. For seed 0 the fine-tuned model reaches 83.8% on held-out prompts. The other four seeds failed the 80% accuracy gate, so only seed 0 went on to circuit extraction.

What the circuit is worth

The test: replace every feature outside a set with its mean at the generation steps, and measure how much of the model's behaviour survives. The repository's control comparison puts the HAGD circuit next to three alternatives.

Behaviour kept at generation steps, seed 0. The full model is 100%.

The circuit is no better than the controls. From the repository's control comparison (results/controls_seed0.json). Restricting every position, prompt reading included, drops the circuit to 0%.

And the larger numbers in the preprint, 91% behavioural preservation, runs on Pythia and Llama models up to 70B parameters, and cross-architecture transfer, are not in the released code at all. The repository's automated claim audit marks each of them unsupported.

Where it breaks

What I changed

The project page reports only numbers the released code produces. The README prints the audit next to the claims it checks. I am revising the preprint to match. The next experiments follow from the diagnostics: fix recall before improving the scoring, measure sufficiency where the computation happens, and select features jointly.

It was not the result I wanted to publish. It is the most useful thing I did this year.