Mohammed Mudassir Uddin

AgentCompressarXiv 2601.05191Preprint · January 2026

AgentCompress: Task-Aware Compression for Affordable Large Language Model Agents

Most steps in an agent workflow do not need the full model. AgentCompress reads the first 32 tokens of each step, estimates how hard the step is, and runs it on an FP16, INT8, INT4 or INT2 copy of LLaMA-2-70B to match. On 290 research workflows the paper reports 68.3% less compute at 96.2% task success, against 98.2% for the uncompressed model.

68.3%
less compute per workflow (2,847 → 902 TFLOPs)
96.2%
task success, vs 98.2% with no compression
≈12ms
per routing decision, against 0.5–2 s of inference per step
2.4pts
quality spread across four domains, two of them unseen in training

The ideaSpend precision where the reasoning is

A research agent might survey papers, form a hypothesis, design an experiment, clean data, interpret results and format citations, all with one 70B model. Running every step at FP16 costs about 2,847 TFLOPs per workflow, roughly $127 of cloud time. Quantizing everything to INT8 cuts that by 42% but drops task success from 98.2% to 87.5%, because the reasoning steps get the same cut as the clerical ones.

The first few tokens of a request usually give away how hard it is. “Given conflicting evidence on reaction selectivity…” needs the whole model; “Convert these references to APA” does not. AgentCompress learns that signal and makes the precision choice per step.

Literaturereview Hypothesis Experimentdesign Dataprocessing Interpretation Formatting PER WORKFLOW Uniform FP16 no compression FP16 FP16 FP16 FP16 FP16 FP16 2,847TFLOPs 98.2% task success Uniform INT8 same for every step INT8 INT8 INT8 INT8 INT8 INT8 1,644TFLOPs 87.5% task success AgentCompress chosen per step INT8 FP16 FP16 INT4 FP16 INT4 902TFLOPs 96.2% task success
Figure 1. One six-step workflow under three policies. Bar length is bit width, a proxy for the compute a step costs. The per-step choices shown for AgentCompress are the precision the controller picks most often for that stage (Figure 4). Workflow totals are the paper’s Table 1: ResearchAgent, 290 workflows, LLaMA-2-70B.

What is newFrom one compression setting to one per step

Standard practiceAgentCompress
When precision is chosenOnce, before deployment, for an assumed average workload (GPTQ, SmoothQuant, LLM.int8())Per step, at run time, from the step’s first 32 tokens
What is optimizedA single forward pass (early exit, layer skipping)Compute for the whole multi-step workflow, under a quality floor θ = 0.95
Where the policy comes fromA fixed bit width or hand-written rulesA 2.37M-parameter controller meta-trained across many workflows, so it carries over to new domains

How it worksA small controller in front of four copies of the model

CONTROLLER · 2.37M PARAMETERS · ≈12 MS PER STEP Task τi first k = 32 tokens of the step’s prompt Task encoder frozen LLaMA encoder + 6-layer transformer 512-d, 8 heads, mean-pooled → e(τ) ∈ ℝ⁵¹² Cognitive-load predictor 4-head self-attention MLP 256 → 128 → 1, sigmoid → load c ∈ [0, 1] Policy heads precision: 16, 8, 4 or 2 bits attention pruning ρ ≤ 0.75 sparsity ≤ 0.9 Variant cache 4 GPTQ copies of LLaMA-2-70B FP16140 GB INT870 GB INT435 GB INT217.5 GB highlighted: the copy in GPU memory; rest on NVMe 94% of steps reuse it; a swap costs ≈2.3 s Run step i on the selected copy 2× A100 40 GB 0.5–2 s per step Output yi result of step i, handed to step i + 1 config (q, ρ, s) next step
Figure 2. Architecture. The highlighted controller is the new part: it reads the start of each step, predicts a cognitive load, and picks a configuration. The variant cache and the model itself are standard GPTQ-quantized LLaMA-2-70B. Only one copy fits in the 80 GB of GPU memory, so the cache decides which copy stays loaded.

Objective

A workflow W = (τ1, …, τn) runs its steps in order. For each step the policy π picks a configuration: a precision in {FP16, INT8, INT4, INT2} and an attention-pruning ratio in {0, 0.25, 0.5, 0.75}. The target is the cheapest policy whose expected workflow quality stays above a floor:

min⁡π 𝔼𝒲[∑i=1nCost(π(τi))]s.t.𝔼𝒲[Quality(𝒲,π)] ≥ θ=0.95\min_{\pi}\ \mathbb{E}_{\mathcal{W}}\Big[\sum_{i=1}^{n}\mathrm{Cost}\big(\pi(\tau_i)\big)\Big]\quad\text{s.t.}\quad \mathbb{E}_{\mathcal{W}}\big[\mathrm{Quality}(\mathcal{W},\pi)\big]\ \ge\ \theta = 0.95

Reading difficulty from the first 32 tokens

The first k = 32 tokens of the step go through a frozen LLaMA encoder. The mean-pooled final hidden state is refined by a 6-layer transformer:

e(τ)=1k∑j=1khj(L)∈ℝ512,k=32e(\tau) = \frac{1}{k}\sum_{j=1}^{k} h_j^{(L)} \in \mathbb{R}^{512},\qquad k = 32

A 4-head attention layer and a small MLP turn that embedding into a single cognitive-load score:

c=σ(W2 ReLU(W1 Attn(e(τ)))+b) ∈ [0,1]c = \sigma\big(W_2\,\mathrm{ReLU}(W_1\,\mathrm{Attn}(e(\tau))) + b\big)\ \in\ [0,1]

Three heads read the embedding and the score. One picks the precision (a softmax, relaxed with Gumbel-softmax during training so gradients pass through the discrete choice), one picks an attention-pruning ratio up to 0.75, and one a sparsity target up to 0.9.

Keeping the right copy loaded

The four copies take 140, 70, 35 and 17.5 GB. With 80 GB of GPU memory (two A100 40 GB cards) one copy is resident and the rest wait on NVMe. The resident copy is chosen by a frequency-and-recency priority:

Priority(v)=0.7 Freq(v)+0.3 Recency(v)\mathrm{Priority}(v) = 0.7\,\mathrm{Freq}(v) + 0.3\,\mathrm{Recency}(v)

Workflows tend to run similar steps back to back, so 94% of consecutive steps reuse the loaded copy. A swap costs about 2.3 s. Quantizing on the fly instead of caching would cost 340 ms at every switch.

Training

The controller is meta-trained with first-order MAML on batches of 16 workflows for 8,700 iterations (AdamW, learning rate decayed from 10−4 to 10−6). The loss charges for compute but weighs missing the quality floor more heavily:

ℒ(ϕ)=λ1 𝔼[Cost(𝒲)]+λ2 𝔼[max⁡(0, θ−Quality(𝒲))]\mathcal{L}(\phi) = \lambda_1\,\mathbb{E}\big[\mathrm{Cost}(\mathcal{W})\big] + \lambda_2\,\mathbb{E}\big[\max\big(0,\ \theta - \mathrm{Quality}(\mathcal{W})\big)\big]

with λ1 = 0.3 and λ2 = 0.7, so missing the floor costs more than spending compute.

The paper also bounds the savings. If the difficulty estimate is off by at most εc and a fraction plow of steps are easy (load below 0.3), then

𝔼[Savings] ≥ (1−ϵc)  plow  Δcost\mathbb{E}[\mathrm{Savings}]\ \ge\ (1-\epsilon_c)\;p_{\text{low}}\;\Delta_{\text{cost}}

Δcost is the FP16-to-INT4 cost gap. In words: savings grow with how varied the workflow is and how accurately the controller reads it.

ResultsClose to the oracle on both cost and quality

The benchmark, ResearchAgent, has 290 multi-stage workflows built from 2020–2024 papers: computer science (87), physics (71), chemistry (64) and biology (68), with 4 to 15 steps each (8.3 on average). Three ML researchers rated step difficulty on a 5-point scale (κ = 0.81) for a stratified subset of steps, and the ratings were propagated to similar steps by task-type clustering. The baselines are uniform FP16, uniform INT8 and static INT4 (GPTQ, group size 128), plus an oracle that knows every step’s difficulty in advance. Results are means over 5 seeds.

AgentCompress costs about as much as static INT4 and keeps quality within 2 points of FP16

Figure 3. Compute per workflow against task success on ResearchAgent, from Table 1. Whiskers show one standard deviation over 5 seeds. The oracle (open circle) is the best any router could do with perfect knowledge of step difficulty.
Table 1. Compression strategies on ResearchAgent (mean ± s.d., 5 seeds). All differences from uniform FP16 are significant at p < 0.001 (paired t-tests, Bonferroni-corrected).
MethodComputeTFLOPsReduction%Task success%
Uniform FP162,847.30.098.2±1.4
Uniform INT81,643.7±12.442.3±0.487.5±3.8
Static INT4819.6±8.771.2±0.363.8±6.2
AgentCompress902.1±24.368.3±0.996.2±1.9
Oracle794.3±11.272.1±0.498.7±1.2

Against uniform INT8, AgentCompress saves more compute (68.3% vs 42.3%) and keeps 8.7 more points of task success. At the paper’s cloud prices, that takes a $127 session to about $40. Routing adds about 12 ms per step, which is small next to the 0.5–2 s each step spends in inference.

Hypothesis and interpretation steps stay at FP16; formatting and citations drop to 4 bits

Figure 4. How often the controller picks each precision, by workflow stage (each row sums to 100%). Hypothesis steps run at FP16 78% of the time; citation steps run at INT4 or INT2 80% of the time. The controller also asks for FP16 on some formatting steps that involve non-Latin scripts or unusual rules, which a rule based on stage names would miss.

Trained on computer science and physics, it holds up on chemistry and biology

(a) Cost reduction

(b) Quality retention

Figure 5. Cost reduction and quality retention by domain. The controller saw only computer-science and physics workflows in training. Quality stays between 93.7% and 96.1% in all four domains. Chemistry saves the least (65.3%), which the paper attributes to specialized synthesis reasoning making the controller more cautious.

The difficulty estimates also track the human ratings: Pearson r = 0.87 on 150 tasks, with a fitted line of ŷ = 0.94x + 0.03. That slope means hard steps are slightly underestimated and easy ones slightly overestimated.

Table 2. Ablations on ResearchAgent. Each row removes or changes one component; Δ is the change in percentage points from the full system.
ConfigurationCost reduction%ΔTask success%Δ
Full AgentCompress68.3—96.2—
Fixed heuristic controller45.3−23.088.9−7.3
No meta-learning (trained on ResearchAgent only)51.7−16.691.4−4.8
No cognitive-load predictor54.8−13.589.7−6.5
No attention pruning59.2−9.195.8−0.4
Smaller controller (128-d)61.4−6.994.1−2.1
No cache (quantize at run time)67.8−0.595.9−0.3

A hand-written rule does worst: it loses 23 points of savings and 7.3 points of quality. Meta-learning across many workflows is worth 16.6 points of savings on its own. Controllers much smaller than 128 dimensions were tried early on and sometimes sent reasoning steps to INT2, which pushed quality below 50%.

LimitationsWhere it goes wrong

About 8% of steps lose more than 10% quality compared with running at full precision. Two patterns account for most of them:

  • Hidden difficulty. “Summarize the methodology section” reads as routine, but not when the methodology is a subtle new proof technique.
  • Unfamiliar notation. SMILES strings in chemistry workflows look like low-complexity text, so they get compressed too hard.

The paper also lists these limits:

  • The controller needs diverse training workflows. A new domain will likely need some fine-tuning.
  • A 12 ms decision is too slow for applications that need responses in under 10 ms.
  • Only sequential pipelines are supported. Branching or parallel workflows would need architectural changes.
  • Ground-truth difficulty comes from human ratings, which are partly subjective.
  • The FP16 model (140 GB) does not fit in the 80 GB of GPU memory used. INT8, INT4 and INT2 were run directly; FP16 cost and latency were estimated with a calibrated model.

CodeWhat the public notebook runs

The repository implements the controller, the variant cache, the Algorithm 1 training loop, the bounds, the significance tests and the ablations in one Colab notebook. By default it uses TinyLlama, so the controller and training loop run on a CPU; the INT8, INT4 and INT2 paths need a CUDA runtime, and Llama-2 7B/70B need approved Hugging Face access.

Scope of the notebook

The paper’s 290-workflow benchmark and its human difficulty labels are not public. The notebook therefore runs on a synthetic workflow generator whose FP16, INT8 and INT4 anchors are calibrated to Table 1. It shows the pipeline running end to end; it is not a reproduction of the paper’s numbers. To evaluate real workflows, the notebook’s simulate_cost_batch and simulate_quality_batch hooks are replaced with real inference and scoring.

Table 3. The saved notebook run (synthetic workflows, as published in the repository README).
ConfigurationComputeTFLOPsReductionQuality
Uniform FP162,847.3—97.4%
Uniform INT81,643.742.3%86.8%
Static INT4819.671.2%62.5%
AgentCompress1,398.850.9%89.6%
Oracle2,235.321.5%95.5%

The same run measured a mean controller latency of 36.9 ms (p95: 39.7 ms) on its own hardware; the paper’s ≈12 ms is on two A100 GPUs.

CitationCite this work

@article{taha2026agentcompress,
  title   = {AgentCompress: Task-Aware Compression for Affordable
             Large Language Model Agents},
  author  = {Taha, Zuhair Ahmed Khan and Uddin, Mohammed Mudassir
             and Alam, Shahnawaz},
  journal = {arXiv preprint arXiv:2601.05191},
  year    = {2026}
}