The ideaKeep the channels that separate diseases
A farmer who finds an unfamiliar blight needs an answer in the field, often with no lab nearby and no reliable connection. That rules out a large model on a server and rules out collecting thousands of labelled photos of a new disease. The model has to be small enough for a $35–$100 board, and it has to learn a disease from one to ten examples.
Pruning makes the model small, but standard criteria judge a channel by its weight magnitude or by how well the pruned layer reconstructs its output. Neither asks whether the channel helps tell early blight from late blight, and with only a few examples per class, that is the property worth protecting.
What is newPruning that knows the task, interleaved with meta-learning
| Generic channel pruning | PMP-DACIS | |
|---|---|---|
| What makes a channel worth keeping | Weight magnitude, batch-norm scale or reconstruction error | How well it separates disease classes (Fisher ratio), plus gradient sensitivity and activation variance |
| When pruning happens | Once, before or after training | Before and after meta-learning, with the second pass guided by meta-gradients |
| Labels it assumes | Plenty per class | 1, 5 or 10 labelled leaves per disease, in episodes |
| Where it runs | A GPU | A Raspberry Pi 4, at 142 ms per image |
The paper scopes its own claims carefully. DACIS is presented as an empirically motivated combination of known metrics, PMP as an engineering pipeline, and the shot-specific models as a deployment strategy rather than dynamic run-time switching.
How it worksPrune, meta-learn, prune again
Scoring a channel
For channel c in layer ℓ, DACIS is a weighted sum of three signals. The weights were chosen by nested cross-validation and favour class separation:
with (λ1, λ2, λ3) = (0.3, 0.2, 0.5).
𝒢 is the gradient norm of the loss with respect to the channel’s weights, with a Hessian-trace correction. 𝒱 is the variance of the channel’s globally pooled activation. 𝒟, the part that makes the score disease-aware, is Fisher’s discriminant ratio across the N disease classes: between-class scatter over within-class scatter.
Early layers encode textures shared by every disease and deep layers encode disease-specific features, so the pruning threshold varies with depth. It is also relaxed when the classes in an episode look alike (task complexity Ctask is one minus the mean cosine similarity between class prototypes):
Why three stages
Pruning once, from pretrained weights, ranks channels by what mattered for ImageNet. Pruning only after meta-learning disrupts what meta-learning built. PMP does a conservative first cut (40%), lets meta-learning show which channels few-shot tasks actually use, and then makes the second cut with that information:
A fourth or fifth stage buys 0.3 or 0.1 points at 45% or 77% more training time (Figure 5). Separate static models are trained for the 1-, 5- and 10-shot regimes, because the best amount of compression depends on how many examples are available.
ResultsBest accuracy among pruned models at the same size
Experiments use PlantVillage (54,305 images, 38 classes across 14 crops) with a visual domain shift split: training images have uniform lighting and simple backgrounds, test images have complex backgrounds and variable light. PlantDoc (2,598 in-the-wild images, 27 classes) is the harder field test. All methods are 5-way, use a ResNet-18 backbone and are compared at equal parameter budgets; results are over 1,000 test episodes.
At 30% of the parameters, PMP-DACIS is closest to the full model
(a) 5-way 1-shot accuracy
(b) 5-way 5-shot accuracy
| Method | Params% of 11.2M | 1-shot | 5-shot | 10-shot | DES |
|---|---|---|---|---|---|
| ProtoNet (full) | 100 | 71.2±2.4 | 84.6±2.1 | 89.3±1.8 | 0.42 |
| MAML (full) | 100 | 69.8±2.5 | 82.1±2.2 | 87.6±1.9 | 0.38 |
| Magnitude pruning | 30 | 58.4±2.8 | 72.3±2.5 | 79.1±2.1 | 1.21 |
| γ-threshold (network slimming) | 30 | 61.2±2.7 | 75.8±2.4 | 81.4±2.0 | 1.34 |
| Channel pruning | 30 | 63.7±2.6 | 77.2±2.3 | 83.0±1.9 | 1.45 |
| Meta-Prune | 30 | 65.1±2.5 | 79.4±2.2 | 84.8±1.8 | 1.52 |
| PMP-DACIS | 30 | 68.9±2.1 | 83.2±1.8 | 88.1±1.5 | 1.98 |
| PMP-DACIS | 22 | 66.4±2.2 | 81.0±1.9 | 86.3±1.6 | 2.31 |
The margin over other pruning methods is largest with the fewest examples, which fits the premise: when the support set is tiny, keeping class-separating channels matters most. On in-the-wild PlantDoc images the pruned model even beats the full-size baselines:
| Method | 1-shot | 5-shot | 10-shot |
|---|---|---|---|
| ProtoNet (full) | 42.5±2.8 | 61.3±2.4 | 68.7±2.1 |
| MAML (full) | 40.1±2.9 | 58.9±2.5 | 66.2±2.2 |
| Meta-Prune | 38.4±3.0 | 55.2±2.6 | 62.1±2.3 |
| PMP-DACIS | 45.8±2.6 | 64.1±2.2 | 71.5±1.9 |
Meta-training and the Fisher term contribute the most
Three stages are where extra training stops paying off
| Method | Early → early | Early → late | CSG | FSI |
|---|---|---|---|---|
| ProtoNet (full) | 85.2 | 62.4 | 0.73 | 0.89 |
| Magnitude pruning | 73.8 | 48.1 | 0.65 | 0.76 |
| PMP-DACIS | 82.8 | 68.7 | 0.83 | 0.92 |
On the deviceFrom half a second to 142 milliseconds
| Device | Full ResNet-18ms · FPS | PMP-DACIS, 2.5Mms · FPS | Speed-up |
|---|---|---|---|
| Raspberry Pi 4 | 512 · 1.95 | 142 · 7.0 | 3.6× |
| Jetson Nano | 85 · 11.8 | 45 · 22.2 | 1.9× |
| Pixel 6 | — | 28 · 35.7 | — |
| RTX 3080 | — | 8 · 125 | — |
In the field, a wrong diagnosis costs a crop, so the model also reports when it is unsure. With Monte Carlo dropout over 20 passes, 23% of predictions are flagged for a human to check. Among flagged predictions the error rate is 67%; among the rest it is 4.2% (Spearman ρ = 0.72 between uncertainty and error).
Training still needs a GPU: about 8.5 hours on one RTX 3080. “Edge” here means inference only.
LimitationsWhere it struggles
| Pair | Confusion rate | Why |
|---|---|---|
| Early blight vs late blight | 14.2% | similar brown lesions; the difference is fine texture |
| Bacterial spot vs Septoria leaf spot | 11.8% | small spots with halos in both |
| Healthy vs early-stage disease | 10.4% | first symptoms are subtle |
The pruned and full models make similar mistakes (Spearman ρ = 0.89 between their confusion matrices), so pruning amplifies existing weaknesses rather than adding new ones. Most errors (63%) fall between biologically related pathogens.
- All crops tested are Solanaceae (tomato, potato, pepper). Narrow-leaf cereals and compound-leaf legumes are untested.
- PlantVillage images were taken in a lab. The domain-shift split is a proxy for field conditions, not a field study.
- The baselines are classical: ProtoNet, MAML and structured pruning. Transformer-based few-shot learners, distillation and quantization are not compared.
- The DACIS weights were tuned on these datasets, and the disease taxonomy behind the Fisher term needs expert input. Without the Fisher term, DACIS falls back to the gradient and variance terms: 78.4% at 5-shot, against 72.3% for magnitude pruning.
- Pruning masks are fixed once deployed; the model cannot adjust its size to a harder input.
CitationCite this work
@article{uddin2026pmpdacis,
title = {Meta-Learning Guided Pruning for Few-Shot Plant Pathology
on Edge Devices},
author = {Uddin, Mohammed Mudassir and Alam, Shahnawaz and Pasha, Mohammed Kaif
and Rehman, Tasneem Bano and Taranum, Fahmina and Begum, Afroze},
journal = {arXiv preprint arXiv:2601.02353},
year = {2026}
}