Reference

Small Models and Classifiers

The working case for a swarm of small, specialised models and classifiers instead of one large generalist calling itself for every step, and a catalogue of what you can actually use.

Document
SLM-REF
Revision
2026.09
Status
Living
Scope
Routing, triage, review
Assumes
Some ML vocabulary
Abstract

Most decisions inside an agent system are narrow: is this finding a blocker, which team handles this, is this input safe, is this noise. A frontier model answers such a question by writing prose that a program must then parse, which is a reliable way to build something that fails silently. This document makes the case for the opposite arrangement, surveys the research that supports it, and catalogues the models and classifiers that can actually hold those positions today, grouped by what they do rather than by who made them.

The case for specialisation

An octopus arm's nerve cord is segmented rather than centrally wired: each segment handles local sucker and muscle control independently, and the central brain is involved only in higher-level coordination.2 That is the mental model for an agent graph. Most nodes should be a small, local decision-maker, and only the parts that need genuinely global judgment should reach for something bigger.

The mental model to keep Do not route every decision through the frontier model. Give each arm of the graph its own local judgment, and reserve the expensive model for cross-cutting synthesis that nothing smaller could do.

Three papers that ground it beyond analogy

Economic

Small models beat large ones on cost, latency and controllability for most agent subtasks, and modern small models are already capable enough at function-calling and tool-use style work.1 The practical reading: use the small model unless the task genuinely needs the big one.

Philosophical

General intelligence as a goal may be incoherent. Human intelligence is itself a bundle of survival-relevant specialisations, and the No Free Lunch result plus biological precedent both favour targeted systems over universal ones.3

Empirical

Large models develop internally modular substructure during training on their own.4 The specialisation the field is now doing by hand, one small model per task, mirrors something the models already do to themselves.

The failure this replaces

The argument is not only about cost. A model asked for an opinion returns prose, and a program that must act on prose ends up searching it for substrings. A review gate that searched its input for the word approve will happily pass the sentence "I do not approve". The bug is not the regular expression; it is that the decision was expressed in prose at all.

  the failure this replaces
  ┌──────────────┐ an essay   ┌──────────────┐ substring    ┌────────────┐
  │  a reviewer  │───────────>│    a grep    │─────────────>│   WRONG    │
  └──────────────┘            └──────────────┘              └────────────┘
    "I do not approve" matched the word approve

  the contract
  ┌──────────────┐ one token  ┌──────────────┐ a field      ┌────────────┐
  │  a reviewer  │───────────>│   a check    │─────────────>│  a decision│
  └──────────────┘            └──────────────┘              └────────────┘
    the decoder cannot emit anything but pass, fail or escalate
move the decision into a field and the parser disappears. This is also what makes a three-billion-parameter model safe in the position.

A 3B model choosing between two enumerated values is a different object from a 3B model asked for an opinion. The first is frequently good enough; the second never is. That distinction, not parameter count, is what makes small models usable as deciders.

Tiny language models

The lineage runs from n-grams through Word2Vec and the RNN and LSTM era to the transformer, and the constraint that motivates going small is blunt: weights at GPT-3 scale are hundreds of gigabytes in half precision and simply do not fit on edge hardware.5 Distillation, quantisation and purpose-built small architectures are the three answers, and all three are now mature enough to use without ceremony.

The cascade, which is the shape that matters

Put a tiny model in front of a large one: the tiny model handles the input when it is confident, and falls through to the big model when it is not. This gives large compute reductions with negligible accuracy loss on vision workloads at ImageNet scale.6 The notable design choice there is a fully independent tiny model rather than an early exit branching off the large model's backbone, which avoids the usual problem that intermediate features lack task-relevant semantics. It is the same structure as a boosted cascade in classical detection.7

        an input
            │
            v
  ┌────────────────────┐  confident     ┌────────────────────┐
  │ tier 1  embeddings │───────────────>│     a decision     │
  │ plus a linear head │                │ milliseconds, CPU  │
  └────────────────────┘                └────────────────────┘
            │
            v  unsure
  ┌────────────────────┐  confident     ┌────────────────────┐
  │ tier 3  a 3B model │───────────────>│     a decision     │
  │ decoder = the enum │                │  about one second  │
  └────────────────────┘                └────────────────────┘
            │
            v  escalate
  ┌────────────────────┐
  │  a frontier model  │
  └────────────────────┘
                          only the hard cases reach the expensive tier
the cascade. Each rung answers what it can and passes the rest up; the escalation edge is a token the next stage already reads.

Non-generative classifiers

Prior art for "do not generate text to make a decision, classify instead". These run in production at scale today, and they are the existence proof that a sub-two-megabyte purpose-built model beats both a hand-written heuristic and a call to a language model.

classifiers already doing this job
systemsizelatencytask
Magika~1 MB, ONNXsingle-digit ms file type from sparse byte samples, 100+ types
YARA-Xcompiled rulesnear-instant malware and structural signature matching
fastText<10 MB compressedmicroseconds language ID across 170+ languages; general linear text classification
TrIDsignature databasefast binary file identification by statistical point scoring
libmagictinyinstant magic bytes at fixed offsets; the deterministic fallback

Two mechanisms here are worth stealing directly, independent of the models. Sparse ingestion: read only the first, middle and last few kilobytes rather than the whole payload, which is how a file classifier stays fast on a large file. A deterministic fallback underneath a learned one: magic bytes catch the easy cases and the model handles the rest, which is the cascade again at a smaller scale.

The deployment detail that dominates everything else Inference is not the expensive part. On a typical workstation, starting a Python interpreter costs around 14 ms and importing an ONNX runtime on top of that costs around 83 ms, while a small compiled binary starts in about 25 ms. A 5 ms classifier invoked as a fresh Python process is therefore roughly 94% overhead. Either amortise one process across a whole batch, or use a compiled command-line build with the weights embedded, which is what Magika and fastText both ship.

The System 1 class

A category named after Kahneman's fast, intuitive cognition: a model that is non-autoregressive, evaluates a whole schema of typed questions in one parallel forward pass rather than generating text token by token, and is trained so its stated probabilities are calibrated. The category reduces decision-making to three question shapes, and those three are the right API to build against whatever sits behind them.

the three question shapes
primitiveasksreturnsconsumed as
choicepick one of Nthe option, and a probability over all Nan exit code: the index
yes/nois X truethe probability of yesa gate: above threshold or not
scorerate on a scalea numbera tally, or a ranking
Discount the headline number The commercial entrant in this category claims a 40 to 200 times speed and cost advantage. A technical teardown8 argues it is largely an inference-time arrangement, parallel per-field decoding with cache reuse over an existing open base model, and reproduces a comparable two to three times speedup with plain constrained single-token decoding on a 1.5B open model. The category is real; the specific multiple is marketing.

The three candidates

what each is actually good for
candidatewhat it isverdict
Jevhosted, proprietary, no open weights, priced per decision Not local a hosted service in a decision path is a dependency, not a component
Laya 421M bidirectional encoder, Apache 2.0, about 33 ms, ONNX runner exists, degrades past roughly 20 options The local pick the right shape with no API key, but an encoder embeds an input, it does not read one
SimpleJev a technique, not a model: read next-token logits per question from any open model Needs raw logits so a runtime that exposes them, not a chat API that hides them

The distinction that decides between them: a bidirectional encoder produces a representation of a text and classifies it, which is ideal for "is this finding a blocker" and useless for "does this change introduce a use-after-free". A position that needs comprehension of the input wants an instruct model with its decoder constrained. The two are complementary rather than competing.

A prior-art claim, corrected A widely circulated claim holds that an earlier reinforcement-learning sales model both predates and matches this approach. Half of that is true. That model9 is legitimate prior art for "a small model trained to output calibrated probabilities instead of generating text", using policy optimisation over conversation embeddings. The second paper usually cited alongside it is about confidence-based routing for hallucination mitigation, a different problem. Mechanically the two differ: one emits a turn-by-turn probability trajectory over sequence embeddings, the other emits choices across a whole schema in one pass.

Security note that generalises. That model ships as a reinforcement-learning framework archive, which contains pickled Python objects. Loading one executes the author's code. The same is true of any .pkl from a source you do not control, including artefacts your own tooling produces. Study a stranger's model in a throwaway environment or not at all.

Compression represents intelligence

The deepest available argument for the small-model thesis is information-theoretic. If capability tracks how well a model compresses, then parameter count is a proxy and not the quantity itself, and a small model trained tightly on a narrow distribution is not a degraded general model but a well-compressed specialist.

The empirical version reports a near-linear relationship between compression ability and benchmark performance across a range of models.10 The epistemic version argues compression is foundational to intelligence rather than merely correlated with it.11 Taken together they reframe the question from how large to how well matched to the distribution, which is exactly the bet a specialised graph makes.

Prompt compression, the same idea pointed the other way

A small encoder-style token classifier can decide which tokens in a prompt carry information and drop the rest, typically two to five times compression with little loss.12 Architecturally that is a small classifier used to shrink a large model's input rather than to replace the large model. Worth measuring anywhere several positions read the same large input, which in a review panel is the common case.

The catalogue

Everything below is something you could actually put behind a decision. Grouped by what it does, because that is how you choose one.

This section ages faster than the rest Parameter counts, version numbers and even family names drift within months, and several figures here are as-advertised rather than measured. Treat it as a map of the categories and verify a specific number against the vendor's own model card before citing it.

Generative small models

For a position that must read something and judge it.

compact generative families
familycompact memberswhy you would pick it
Qwen0.5B / 0.6B / 1.5B / 1.7B / 3B, plus coder variants at 1.5B and 3B the strongest coding ability per byte at the small end, long context, permissive licence. The coder variants are the default reviewer candidate on a 4 GB card
Llama3.2 at 1B and 3B built for on-device, with the widest tooling support, so it runs on anything
PhiPhi-4-mini around 3.8B, Phi-3.5-mini heavily trained on synthetic reasoning data; punches above its size on structured tasks, weaker on open-ended prose
Gemma270M, 1B, 2B, 4B the 270M is small enough to be interesting as a classifier substrate; the 1B and 2B are genuine instruction-followers
SmolLMSmolLM2 1.7B, SmolLM3 3B fully open training data and recipe, which matters if you need to explain or reproduce behaviour
Granite2B and small mixture-of-experts variants aimed at enterprise retrieval and security work; unusually clear licensing and provenance
Ministral3B, 8Bedge-focused with very long context for the size
DeepSeek distilledR1-Distill-Qwen-1.5B and relatives reasoning traces distilled into a tiny model; strong where the answer is verifiable
Falconthe tiny seriesworth watching; check current parameter counts before citing
OLMo1Bfully open corpus and checkpoints; the reference point for reproducibility work

On a 4 GB consumer card with roughly 3.6 GB usable, 3B at 4-bit quantisation is the practical ceiling. Anything larger wants a machine with real memory.

Encoders for classification

For a decision with a stable label set. Fine-tune on a few hundred labelled examples, export to ONNX, and run with no deep-learning framework at all.

the BERT lineage
modelsizenote
BERT base110Mthe reference architecture; everything below is a response to it
DistilBERT66M40% smaller, substantially faster, keeps most of the accuracy. The safe default
TinyBERT14Mwhen 66M is still too much. Edge and microcontroller territory
ALBERT12M+parameter sharing; small on disk, not proportionally faster at inference
ELECTRA small14Mtrained by discriminating replaced tokens rather than masking; strong for the size
RoBERTa125M+better-trained BERT; the base for many ready-made fine-tuned classifiers
DeBERTa v3 small44Mdisentangled attention; consistently the best accuracy per parameter here for language understanding
ModernBERT149M / 395Mthe 2024 refresh: long context, faster inference, drop-in. Prefer it for new work

Embeddings, rerankers and entailment

the retrieval and scoring layer
modelkindnote
all-MiniLM-L6-v2bi-encoder~90 MB. The classic default: fast, good enough, runs anywhere
bge-small-en-v1.5bi-encoder~130 MB. Better retrieval quality at a similar size
E5 small, GTE smallbi-encoderdirect competitors to bge-small. Benchmark all three on your own labels; the ranking flips by domain
bge-m3bi-encodermultilingual and multi-granularity when you need it
nomic-embed-textbi-encoderlong context, open training data
SetFitmethodcontrastive fine-tuning over any embedder; useful accuracy from 8 to 16 examples per class13
bge-rerankercross-encoderscores a query and candidate jointly. Slower than comparing embeddings and much more accurate. This is the score primitive off the shelf
jina-rerankercross-encodersame shape, different training. Worth comparing directly
bart-large-mnlientailmentthe original zero-shot classifier: scores "this text entails LABEL" per label, with no training at all
deberta-v3 zeroshotentailmentsmaller and usually better than the bart baseline for the same job

Guard and policy classifiers

An under-used category, and exactly the right shape for an admission check on untrusted input.

safety and policy
modelnote
Llama Guardclassifies prompts and responses against a taxonomy you can edit
ShieldGemmathe same job on a Gemma base, in several sizes
Prompt Guarda small classifier specifically for injection and jailbreak detection. Tiny, fast, and the right thing in front of untrusted input

Beyond text

small multimodal
modelsizenote
Moondream2~1.8Ba position that can look at a screenshot or a diagram on modest hardware
SmolVLM256M to 2.2Bthe smallest genuinely usable vision-language family
Florence-2230M / 770Mdetection, captioning and OCR in one small model
Whisper tiny / base39M / 74Mspeech to text, if anything needs to read audio

Constrained decoding

Not models, but what makes a small generative model safe as a decider.

guided generation tooling
toolnote
Outlines regex and JSON-schema guided generation; the reference implementation of the idea14
XGrammarfast grammar-constrained decoding, integrated into several serving stacks
llama.cpp GBNFgrammar files native to llama.cpp, and the practical path to reading raw logits that chat APIs hide
Structured output APIsa JSON schema in the request. The easiest version, and enough for most work

In practice

Five tiers, cheapest first

The discipline is to start at the top and descend only when the tier above demonstrably cannot do the job. Every step down costs an order of magnitude of latency and buys correspondingly less determinism.

the tiers
tierwhat it iscosttraining
1sentence embeddings plus logistic regressionmilliseconds, CPU20 to 50 labelled lines per class
1.5a purpose-built decision encodertens of millisecondsnone
2zero-shot entailmentslower than tier 1none
3a small instruct model, decoder constrained to the enumabout a second on a small GPUnone
4a fine-tuned small encoder, exported to ONNXbest accuracy per milliseconda few hundred labelled examples

Four rules for tier 3

  • Temperature zero gives an answer, not a confidence. One deterministic sample is always certain of itself. For a probability, sample five to nine times at a nonzero temperature and count the votes.
  • Put the enumeration in the schema, not the prompt. The prompt is a request; the schema is an enforcement.
  • Ask for a reason field even if you discard it. It costs tokens, it makes a wrong answer diagnosable, and a model required to justify tends to answer better.
  • A large enumeration is the wrong tier. Dozens of labels slow decoding and hurt accuracy; that is tier 1 work.

Calibration, the part nobody checks

A probability is only useful if 0.8 means right about eight times in ten. Logistic regression is roughly calibrated by construction; vote counts from a sampled language model are not, and an entailment score is not a probability at all.

  • Hold out thirty to fifty labelled examples the model has never seen.
  • Bucket predictions by stated confidence in bands of 0.1 and compute the actual accuracy within each bucket. A well-calibrated model's buckets sit on the diagonal.
  • Compute the Brier score,15 the mean squared difference between stated confidence and outcome. Lower is better, and 0.25 is what a coin flip scores.
  • Correct what you find. An overconfident linear head wants stronger regularisation; an overconfident sampled model wants more samples, or a decision threshold raised to 0.7.
BS  =  1N  ∑i (pi − oi)2 (7.1)
The rule No confidence number drives a gate before its reliability table exists. Until then a model's stated certainty is decoration, and a threshold applied to it is a number chosen by vibe.

Choosing, in four questions

picking a tier
questionif yes
Is the label set stable and small?a trained classifier or a purpose-built one. Do not use a language model for this
Does the decision require understanding the input?a generative small model with constrained decoding
Do you need a calibrated probability rather than a label?a decision encoder, a logistic head over embeddings, or a reranker
Is the input untrusted?put a guard classifier in front, as a gate, not as a participant

Glossary

terms used here in a specific sense
termmeaning
SLMa language model small enough to run cheaply and locally, traded against a frontier model's raw capability
System 1 modela non-autoregressive model returning typed decisions in one parallel pass instead of generating text
constrained decodingsampling restricted to tokens that keep the output valid against a schema or grammar, so an enumerated answer is guaranteed rather than requested
logitsraw per-token scores before they become probabilities. Reading them per option gives a true one-pass distribution, which is why an API that hides them cannot do the trick
bi-encoderembeds query and candidate separately and compares vectors. Fast enough to search millions; less accurate than a cross-encoder
cross-encoder / rerankerscores a pair jointly. Slower, much more accurate, and the score primitive available off the shelf
zero-shot NLIscoring "this text entails LABEL" with an entailment model, so a classifier exists before any training data does
guard classifiera small model judging whether input or output violates a policy. Belongs in a gate, never in a reporting position
calibration / Brier scorewhether stated confidence matches observed accuracy; mean of (confidence − outcome) squared, where 0.25 is a coin flip
quantisation (Q4, Q8)storing weights at reduced precision so a larger model fits a smaller card. Q4 is what puts a 3B model on 4 GB
distillationtraining a small model on a large model's outputs. Where most compact reasoning models come from
early exit vs cascadean early exit taps an intermediate layer of one large model; a cascade runs a fully separate small model first, whose features are complete rather than intermediate
sparse ingestionclassifying from a few byte ranges rather than the whole payload
magic bytesfixed-offset signature bytes for deterministic file typing; the fallback beneath a learned classifier

References

  1. Belcak, P., Heinrich, G., Diao, S., Fu, Y., Dong, X., Muralidharan, S., Lin, Y. C., Molchanov, P. Small language models are the future of agentic AI. arXiv:2506.02153, 2025.
  2. Kuuspalu, A., Cody, S., Hale, M. E. Multiple nerve cords connect the arms of octopuses, providing alternative paths for inter-arm signaling. Current Biology, 2022.
  3. Goldfeder, J., Wyder, P., LeCun, Y., Shwartz-Ziv, R. AI must embrace specialization via superhuman adaptable intelligence. arXiv:2602.23643, 2026.
  4. Han, P. et al. Modular cognitive architecture emerges in large language models. Preprint.
  5. Penman, T. Tiny language models. tristanpenman.com, 2025.
  6. Wang, J. et al. Tiny models are the computational saver for large models. arXiv:2403.17726, 2024.
  7. Viola, P., Jones, M. Rapid object detection using a boosted cascade of simple features. CVPR, 2001.
  8. Goedecke, S. Jev means structured output is interesting again. seangoedecke.com, 2026.
  9. DeepMost Innovations. Sales conversion prediction with reinforcement learning. arXiv:2503.23303, 2025.
  10. Huang, Y. et al. Compression represents intelligence linearly. arXiv:2404.09937, 2024.
  11. The information-theoretic imperative: compression and the epistemic foundations of intelligence. arXiv:2510.25883, 2025.
  12. Jiang, H. et al. LLMLingua: compressing prompts for accelerated inference. arXiv:2310.05736, 2023. See also LLMLingua-2, arXiv:2403.12968, 2024.
  13. Tunstall, L. et al. Efficient few-shot learning without prompts (SetFit). arXiv:2209.11055, 2022.
  14. Willard, B. T., Louf, R. Efficient guided generation for large language models. arXiv:2307.09702, 2023.
  15. Brier, G. W. Verification of forecasts expressed in terms of probability. Monthly Weather Review 78(1), 1950.
  16. Curated lists, for going wider: Awesome-LLM and Awesome-SLM, the latter organised into milestone papers, a leaderboard and an open-model catalogue by vendor.