Mohammed Mudassir Uddin

BlogJuly 2026reproducibilityLLM agents

What our AgentCompress notebook shows, and what it doesn't

The paper's benchmark is not public, so the released notebook runs on synthetic workflows and its numbers are lower. Both columns belong in the README.

On 20 July we published the AgentCompress code as a single Colab notebook: the controller, the variant cache, the training loop, the theoretical bounds, the significance tests and the ablations. By default it runs on TinyLlama, so the controller trains on a CPU. The INT8, INT4 and INT2 paths need a GPU, and Llama-2 7B and 70B need approved access.

What it cannot do

The paper's 290-workflow benchmark and its human difficulty labels are not public. So the notebook cannot reproduce the paper's numbers, and it does not pretend to. It generates synthetic workflows, calibrates its uniform FP16, INT8 and INT4 anchors to the paper's table, and runs the whole pipeline end to end.

Table 1. The paper's benchmark against the released notebook. Different workflows and hardware, so these are two conditions, not a reproduction.
MeasurePaper290 workflowsNotebooksynthetic workflows
AgentCompress, compute saved68.3%50.9%
AgentCompress, task quality96.2%89.6%
Uniform FP16, task quality98.2%97.4%
Uniform INT8, task quality87.5%86.8%
Controller latency per decision≈12 ms, two A100s36.9 ms mean

In the notebook the router saves less and keeps less quality, and its controller takes three times as long per decision on the hardware it ran on. None of that contradicts the paper; the conditions differ. But a reader deserves to see both columns side by side before they have to go looking, so the README and the project page print them together.

To evaluate real workflows, you replace the notebook's two simulation hooks with actual inference and scoring. That is the experiment I would most like someone to run.

The habit I took from this into everything after: publish the gap between what the paper reports and what the public code runs, before someone else finds it.