Mohammed Mudassir Uddin

PMP-DACISarXiv 2601.02353Edge ML · few-shot learningUnder review

Meta-Learning Guided Pruning for Few-Shot Plant Pathology on Edge Devices

Generic pruning keeps the channels that matter for ImageNet. Diagnosing a leaf disease from a handful of photos needs the channels that tell one disease from another. DACIS scores every channel by how well it separates disease classes, and the PMP pipeline prunes, meta-learns, then prunes again using the meta-gradients. ResNet-18 goes from 11.2M to 2.5M parameters and runs at 7 frames per second on a Raspberry Pi 4.

2.5M
parameters, down from 11.2M (78% smaller)
142ms
per image on a Raspberry Pi 4, 3.6× faster than the full ResNet-18
+3.8pts
over Meta-Prune at the same size, 5-way 1-shot and 5-shot
98.3%
of the full model’s 5-shot accuracy, with 30% of its parameters

The ideaKeep the channels that separate diseases

A farmer who finds an unfamiliar blight needs an answer in the field, often with no lab nearby and no reliable connection. That rules out a large model on a server and rules out collecting thousands of labelled photos of a new disease. The model has to be small enough for a $35–$100 board, and it has to learn a disease from one to ten examples.

Pruning makes the model small, but standard criteria judge a channel by its weight magnitude or by how well the pruned layer reconstructs its output. Neither asks whether the channel helps tell early blight from late blight, and with only a few examples per class, that is the property worth protecting.

Channel A responses to the two diseases barely overlap channel activation → High Fisher ratio: keep Channel B responds to both diseases about equally channel activation → Low Fisher ratio: prune a magnitude criterion might still keep it Early blight Late blight
Figure 1. The scoring idea, drawn schematically (not measured data). Each curve shows how strongly one channel responds to images of one disease. DACIS prefers channels whose per-class responses are far apart relative to their spread, which is Fisher’s discriminant ratio.

What is newPruning that knows the task, interleaved with meta-learning

Generic channel pruningPMP-DACIS
What makes a channel worth keepingWeight magnitude, batch-norm scale or reconstruction errorHow well it separates disease classes (Fisher ratio), plus gradient sensitivity and activation variance
When pruning happensOnce, before or after trainingBefore and after meta-learning, with the second pass guided by meta-gradients
Labels it assumesPlenty per class1, 5 or 10 labelled leaves per disease, in episodes
Where it runsA GPUA Raspberry Pi 4, at 142 ms per image

The paper scopes its own claims carefully. DACIS is presented as an empirically motivated combination of known metrics, PMP as an engineering pipeline, and the shot-specific models as a deployment strategy rather than dynamic run-time switching.

How it worksPrune, meta-learn, prune again

ResNet-18 ImageNet weights 11.2M params 1 · Initial pruning score channels with DACIS on base-class data remove 40% of parameters short fine-tune → 6.7M params 2 · Meta-learning 5-way, K = 1, 5 or 10 shots inner loop: support set outer loop: query set first-order MAML 60,000 episodes collects meta-gradients 3 · Refinement pruning rescale each DACIS score by its meta-gradient size prune further final fine-tune → 2.5M params (−78%) Edge device Raspberry Pi 4 142 ms / image Jetson Nano 45 ms / image DACIS: disease-aware channel importance Gradient norm 𝒢 sensitivity of the loss λ₁ = 0.3 Variance 𝒱 spread of the activation λ₂ = 0.2 Fisher ratio 𝒟 how well it separates diseases λ₃ = 0.5 weighted sum → layer-adaptive threshold τℓ → keep or prune each channel stage 3 multiplies the score by (1 + γ‖Gₘₑₜₐ‖) scores scores
Figure 2. The Prune-then-Meta-Learn-then-Prune (PMP) pipeline. The highlighted parts are new: DACIS scoring (bottom) drives the pruning in stages 1 and 3, and stage 3 also uses the meta-gradients collected in stage 2. Stage 2 is standard first-order MAML.

Scoring a channel

For channel c in layer ℓ, DACIS is a weighted sum of three signals. The weights were chosen by nested cross-validation and favour class separation:

DACISℓ(c)=λ1 𝒢ℓ(c)+λ2 𝒱ℓ(c)+λ3 𝒟ℓ(c)\mathrm{DACIS}^{(c)}_{\ell} = \lambda_1\,\mathcal{G}^{(c)}_{\ell} + \lambda_2\,\mathcal{V}^{(c)}_{\ell} + \lambda_3\,\mathcal{D}^{(c)}_{\ell}

with (λ1, λ2, λ3) = (0.3, 0.2, 0.5).

𝒢 is the gradient norm of the loss with respect to the channel’s weights, with a Hessian-trace correction. 𝒱 is the variance of the channel’s globally pooled activation. 𝒟, the part that makes the score disease-aware, is Fisher’s discriminant ratio across the N disease classes: between-class scatter over within-class scatter.

𝒟ℓ(c)=∑n=1NNn (aˉℓ,n(c)−aˉℓ(c))2∑n=1N∑x∈𝒞n(aℓ(c)(x)−aˉℓ,n(c))2\mathcal{D}^{(c)}_{\ell} = \frac{\sum_{n=1}^{N} N_n\,\big(\bar a^{(c)}_{\ell,n} - \bar a^{(c)}_{\ell}\big)^2}{\sum_{n=1}^{N}\sum_{x\in\mathcal{C}_n}\big(a^{(c)}_{\ell}(x) - \bar a^{(c)}_{\ell,n}\big)^2}

Early layers encode textures shared by every disease and deep layers encode disease-specific features, so the pruning threshold varies with depth. It is also relaxed when the classes in an episode look alike (task complexity Ctask is one minus the mean cosine similarity between class prototypes):

τℓ=τbase(1+α ℓL) e−β Ctask\tau_\ell = \tau_{\text{base}}\Big(1 + \alpha\,\frac{\ell}{L}\Big)\,e^{-\beta\,C_{\text{task}}}

Why three stages

Pruning once, from pretrained weights, ranks channels by what mattered for ImageNet. Pruning only after meta-learning disrupts what meta-learning built. PMP does a conservative first cut (40%), lets meta-learning show which channels few-shot tasks actually use, and then makes the second cut with that information:

DACIS~ℓ(c)=DACISℓ(c) (1+γ ∥Gmeta,ℓ(c)∥2)\widetilde{\mathrm{DACIS}}^{(c)}_{\ell} = \mathrm{DACIS}^{(c)}_{\ell}\,\big(1 + \gamma\,\lVert G^{(c)}_{\text{meta},\ell}\rVert_2\big)

A fourth or fifth stage buys 0.3 or 0.1 points at 45% or 77% more training time (Figure 5). Separate static models are trained for the 1-, 5- and 10-shot regimes, because the best amount of compression depends on how many examples are available.

ResultsBest accuracy among pruned models at the same size

Experiments use PlantVillage (54,305 images, 38 classes across 14 crops) with a visual domain shift split: training images have uniform lighting and simple backgrounds, test images have complex backgrounds and variable light. PlantDoc (2,598 in-the-wild images, 27 classes) is the harder field test. All methods are 5-way, use a ResNet-18 backbone and are compared at equal parameter budgets; results are over 1,000 test episodes.

At 30% of the parameters, PMP-DACIS is closest to the full model

(a) 5-way 1-shot accuracy

(b) 5-way 5-shot accuracy

Figure 3. Every pruned model keeps 30% of ResNet-18’s parameters (3.36M). The dashed line is the unpruned ProtoNet (11.2M). PlantVillage with the visual domain shift split; episode-level standard deviation is about ±2.3 points.
Table 1. 5-way accuracy (%) on PlantVillage under the visual domain shift split (mean ± episode-level s.d.). DES is the paper’s deployment efficiency score: accuracy × FPS / (parameters × energy).
MethodParams% of 11.2M1-shot5-shot10-shotDES
ProtoNet (full)10071.2±2.484.6±2.189.3±1.80.42
MAML (full)10069.8±2.582.1±2.287.6±1.90.38
Magnitude pruning3058.4±2.872.3±2.579.1±2.11.21
γ-threshold (network slimming)3061.2±2.775.8±2.481.4±2.01.34
Channel pruning3063.7±2.677.2±2.383.0±1.91.45
Meta-Prune3065.1±2.579.4±2.284.8±1.81.52
PMP-DACIS3068.9±2.183.2±1.888.1±1.51.98
PMP-DACIS2266.4±2.281.0±1.986.3±1.62.31

The margin over other pruning methods is largest with the fewest examples, which fits the premise: when the support set is tiny, keeping class-separating channels matters most. On in-the-wild PlantDoc images the pruned model even beats the full-size baselines:

Table 2. 5-way accuracy (%) on PlantDoc, over 1,000 episodes.
Method1-shot5-shot10-shot
ProtoNet (full)42.5±2.861.3±2.468.7±2.1
MAML (full)40.1±2.958.9±2.566.2±2.2
Meta-Prune38.4±3.055.2±2.662.1±2.3
PMP-DACIS45.8±2.664.1±2.271.5±1.9

Meta-training and the Fisher term contribute the most

Figure 4. Drop in 5-way 5-shot accuracy when one component is removed, from 83.2% for the full method at 30% of the parameters. Among alternative class-separation measures, the Fisher ratio also beats maximum mean discrepancy (−2.0), KL divergence (−2.4) and silhouette score (−3.6).

Three stages are where extra training stops paying off

Figure 5. 5-way 5-shot accuracy against training time relative to single-stage pruning, all at 30% of the parameters. Grey points are other orderings: meta-learning first and pruning after (M→P), and pruning continuously during meta-learning.
Table 3. Robustness (5-way 5-shot). CSG is accuracy on late-stage disease divided by accuracy on early-stage disease, for models trained on early-stage samples. FSI is 1 − σ/μ of accuracy over 1,000 random support sets.
MethodEarly → earlyEarly → lateCSGFSI
ProtoNet (full)85.262.40.730.89
Magnitude pruning73.848.10.650.76
PMP-DACIS82.868.70.830.92

On the deviceFrom half a second to 142 milliseconds

Table 4. Single-image inference at 224×224, batch size 1, 100 warm-up and 1,000 timed runs. The Raspberry Pi 4 (4 GB) runs headless with passive cooling.
DeviceFull ResNet-18ms · FPSPMP-DACIS, 2.5Mms · FPSSpeed-up
Raspberry Pi 4512 · 1.95142 · 7.03.6×
Jetson Nano85 · 11.845 · 22.21.9×
Pixel 6—28 · 35.7—
RTX 3080—8 · 125—

In the field, a wrong diagnosis costs a crop, so the model also reports when it is unsure. With Monte Carlo dropout over 20 passes, 23% of predictions are flagged for a human to check. Among flagged predictions the error rate is 67%; among the rest it is 4.2% (Spearman ρ = 0.72 between uncertainty and error).

Training still needs a GPU: about 8.5 hours on one RTX 3080. “Edge” here means inference only.

LimitationsWhere it struggles

Table 5. Most confused disease pairs.
PairConfusion rateWhy
Early blight vs late blight14.2%similar brown lesions; the difference is fine texture
Bacterial spot vs Septoria leaf spot11.8%small spots with halos in both
Healthy vs early-stage disease10.4%first symptoms are subtle

The pruned and full models make similar mistakes (Spearman ρ = 0.89 between their confusion matrices), so pruning amplifies existing weaknesses rather than adding new ones. Most errors (63%) fall between biologically related pathogens.

  • All crops tested are Solanaceae (tomato, potato, pepper). Narrow-leaf cereals and compound-leaf legumes are untested.
  • PlantVillage images were taken in a lab. The domain-shift split is a proxy for field conditions, not a field study.
  • The baselines are classical: ProtoNet, MAML and structured pruning. Transformer-based few-shot learners, distillation and quantization are not compared.
  • The DACIS weights were tuned on these datasets, and the disease taxonomy behind the Fisher term needs expert input. Without the Fisher term, DACIS falls back to the gradient and variance terms: 78.4% at 5-shot, against 72.3% for magnitude pruning.
  • Pruning masks are fixed once deployed; the model cannot adjust its size to a harder input.

CitationCite this work

@article{uddin2026pmpdacis,
  title   = {Meta-Learning Guided Pruning for Few-Shot Plant Pathology
             on Edge Devices},
  author  = {Uddin, Mohammed Mudassir and Alam, Shahnawaz and Pasha, Mohammed Kaif
             and Rehman, Tasneem Bano and Taranum, Fahmina and Begum, Afroze},
  journal = {arXiv preprint arXiv:2601.02353},
  year    = {2026}
}