The ideaSpend precision where the reasoning is
A research agent might survey papers, form a hypothesis, design an experiment, clean data, interpret results and format citations, all with one 70B model. Running every step at FP16 costs about 2,847 TFLOPs per workflow, roughly $127 of cloud time. Quantizing everything to INT8 cuts that by 42% but drops task success from 98.2% to 87.5%, because the reasoning steps get the same cut as the clerical ones.
The first few tokens of a request usually give away how hard it is. “Given conflicting evidence on reaction selectivity…” needs the whole model; “Convert these references to APA” does not. AgentCompress learns that signal and makes the precision choice per step.
What is newFrom one compression setting to one per step
| Standard practice | AgentCompress | |
|---|---|---|
| When precision is chosen | Once, before deployment, for an assumed average workload (GPTQ, SmoothQuant, LLM.int8()) | Per step, at run time, from the step’s first 32 tokens |
| What is optimized | A single forward pass (early exit, layer skipping) | Compute for the whole multi-step workflow, under a quality floor θ = 0.95 |
| Where the policy comes from | A fixed bit width or hand-written rules | A 2.37M-parameter controller meta-trained across many workflows, so it carries over to new domains |
How it worksA small controller in front of four copies of the model
Objective
A workflow W = (τ1, …, τn) runs its steps in order. For each step the policy π picks a configuration: a precision in {FP16, INT8, INT4, INT2} and an attention-pruning ratio in {0, 0.25, 0.5, 0.75}. The target is the cheapest policy whose expected workflow quality stays above a floor:
Reading difficulty from the first 32 tokens
The first k = 32 tokens of the step go through a frozen LLaMA encoder. The mean-pooled final hidden state is refined by a 6-layer transformer:
A 4-head attention layer and a small MLP turn that embedding into a single cognitive-load score:
Three heads read the embedding and the score. One picks the precision (a softmax, relaxed with Gumbel-softmax during training so gradients pass through the discrete choice), one picks an attention-pruning ratio up to 0.75, and one a sparsity target up to 0.9.
Keeping the right copy loaded
The four copies take 140, 70, 35 and 17.5 GB. With 80 GB of GPU memory (two A100 40 GB cards) one copy is resident and the rest wait on NVMe. The resident copy is chosen by a frequency-and-recency priority:
Workflows tend to run similar steps back to back, so 94% of consecutive steps reuse the loaded copy. A swap costs about 2.3 s. Quantizing on the fly instead of caching would cost 340 ms at every switch.
Training
The controller is meta-trained with first-order MAML on batches of 16 workflows for 8,700 iterations (AdamW, learning rate decayed from 10−4 to 10−6). The loss charges for compute but weighs missing the quality floor more heavily:
with λ1 = 0.3 and λ2 = 0.7, so missing the floor costs more than spending compute.
The paper also bounds the savings. If the difficulty estimate is off by at most εc and a fraction plow of steps are easy (load below 0.3), then
Δcost is the FP16-to-INT4 cost gap. In words: savings grow with how varied the workflow is and how accurately the controller reads it.
ResultsClose to the oracle on both cost and quality
The benchmark, ResearchAgent, has 290 multi-stage workflows built from 2020–2024 papers: computer science (87), physics (71), chemistry (64) and biology (68), with 4 to 15 steps each (8.3 on average). Three ML researchers rated step difficulty on a 5-point scale (κ = 0.81) for a stratified subset of steps, and the ratings were propagated to similar steps by task-type clustering. The baselines are uniform FP16, uniform INT8 and static INT4 (GPTQ, group size 128), plus an oracle that knows every step’s difficulty in advance. Results are means over 5 seeds.
AgentCompress costs about as much as static INT4 and keeps quality within 2 points of FP16
| Method | ComputeTFLOPs | Reduction% | Task success% |
|---|---|---|---|
| Uniform FP16 | 2,847.3 | 0.0 | 98.2±1.4 |
| Uniform INT8 | 1,643.7±12.4 | 42.3±0.4 | 87.5±3.8 |
| Static INT4 | 819.6±8.7 | 71.2±0.3 | 63.8±6.2 |
| AgentCompress | 902.1±24.3 | 68.3±0.9 | 96.2±1.9 |
| Oracle | 794.3±11.2 | 72.1±0.4 | 98.7±1.2 |
Against uniform INT8, AgentCompress saves more compute (68.3% vs 42.3%) and keeps 8.7 more points of task success. At the paper’s cloud prices, that takes a $127 session to about $40. Routing adds about 12 ms per step, which is small next to the 0.5–2 s each step spends in inference.
Hypothesis and interpretation steps stay at FP16; formatting and citations drop to 4 bits
Trained on computer science and physics, it holds up on chemistry and biology
(a) Cost reduction
(b) Quality retention
The difficulty estimates also track the human ratings: Pearson r = 0.87 on 150 tasks, with a fitted line of ŷ = 0.94x + 0.03. That slope means hard steps are slightly underestimated and easy ones slightly overestimated.
| Configuration | Cost reduction% | Δ | Task success% | Δ |
|---|---|---|---|---|
| Full AgentCompress | 68.3 | — | 96.2 | — |
| Fixed heuristic controller | 45.3 | −23.0 | 88.9 | −7.3 |
| No meta-learning (trained on ResearchAgent only) | 51.7 | −16.6 | 91.4 | −4.8 |
| No cognitive-load predictor | 54.8 | −13.5 | 89.7 | −6.5 |
| No attention pruning | 59.2 | −9.1 | 95.8 | −0.4 |
| Smaller controller (128-d) | 61.4 | −6.9 | 94.1 | −2.1 |
| No cache (quantize at run time) | 67.8 | −0.5 | 95.9 | −0.3 |
A hand-written rule does worst: it loses 23 points of savings and 7.3 points of quality. Meta-learning across many workflows is worth 16.6 points of savings on its own. Controllers much smaller than 128 dimensions were tried early on and sometimes sent reasoning steps to INT2, which pushed quality below 50%.
LimitationsWhere it goes wrong
About 8% of steps lose more than 10% quality compared with running at full precision. Two patterns account for most of them:
- Hidden difficulty. “Summarize the methodology section” reads as routine, but not when the methodology is a subtle new proof technique.
- Unfamiliar notation. SMILES strings in chemistry workflows look like low-complexity text, so they get compressed too hard.
The paper also lists these limits:
- The controller needs diverse training workflows. A new domain will likely need some fine-tuning.
- A 12 ms decision is too slow for applications that need responses in under 10 ms.
- Only sequential pipelines are supported. Branching or parallel workflows would need architectural changes.
- Ground-truth difficulty comes from human ratings, which are partly subjective.
- The FP16 model (140 GB) does not fit in the 80 GB of GPU memory used. INT8, INT4 and INT2 were run directly; FP16 cost and latency were estimated with a calibrated model.
CodeWhat the public notebook runs
The repository implements the controller, the variant cache, the Algorithm 1 training loop, the bounds, the significance tests and the ablations in one Colab notebook. By default it uses TinyLlama, so the controller and training loop run on a CPU; the INT8, INT4 and INT2 paths need a CUDA runtime, and Llama-2 7B/70B need approved Hugging Face access.
The paper’s 290-workflow benchmark and its human difficulty labels are not public. The notebook therefore runs on a synthetic workflow generator whose FP16, INT8 and INT4 anchors are calibrated to Table 1. It shows the pipeline running end to end; it is not a reproduction of the paper’s numbers. To evaluate real workflows, the notebook’s simulate_cost_batch and simulate_quality_batch hooks are replaced with real inference and scoring.
| Configuration | ComputeTFLOPs | Reduction | Quality |
|---|---|---|---|
| Uniform FP16 | 2,847.3 | — | 97.4% |
| Uniform INT8 | 1,643.7 | 42.3% | 86.8% |
| Static INT4 | 819.6 | 71.2% | 62.5% |
| AgentCompress | 1,398.8 | 50.9% | 89.6% |
| Oracle | 2,235.3 | 21.5% | 95.5% |
The same run measured a mean controller latency of 36.9 ms (p95: 39.7 ms) on its own hardware; the paper’s ≈12 ms is on two A100 GPUs.
CitationCite this work
@article{taha2026agentcompress,
title = {AgentCompress: Task-Aware Compression for Affordable
Large Language Model Agents},
author = {Taha, Zuhair Ahmed Khan and Uddin, Mohammed Mudassir
and Alam, Shahnawaz},
journal = {arXiv preprint arXiv:2601.05191},
year = {2026}
}