Reference
Small Models and Classifiers
The working case for a swarm of small, specialised models and classifiers instead of one large generalist calling itself for every step, and a catalogue of what you can actually use.
Most decisions inside an agent system are narrow: is this finding a blocker, which team handles this, is this input safe, is this noise. A frontier model answers such a question by writing prose that a program must then parse, which is a reliable way to build something that fails silently. This document makes the case for the opposite arrangement, surveys the research that supports it, and catalogues the models and classifiers that can actually hold those positions today, grouped by what they do rather than by who made them.
The case for specialisation
An octopus arm's nerve cord is segmented rather than centrally wired: each segment handles local sucker and muscle control independently, and the central brain is involved only in higher-level coordination.2 That is the mental model for an agent graph. Most nodes should be a small, local decision-maker, and only the parts that need genuinely global judgment should reach for something bigger.
Three papers that ground it beyond analogy
Economic
Small models beat large ones on cost, latency and controllability for most agent subtasks, and modern small models are already capable enough at function-calling and tool-use style work.1 The practical reading: use the small model unless the task genuinely needs the big one.
Philosophical
General intelligence as a goal may be incoherent. Human intelligence is itself a bundle of survival-relevant specialisations, and the No Free Lunch result plus biological precedent both favour targeted systems over universal ones.3
Empirical
Large models develop internally modular substructure during training on their own.4 The specialisation the field is now doing by hand, one small model per task, mirrors something the models already do to themselves.
The failure this replaces
The argument is not only about cost. A model asked for an opinion returns prose, and a program that must act on prose ends up searching it for substrings. A review gate that searched its input for the word approve will happily pass the sentence "I do not approve". The bug is not the regular expression; it is that the decision was expressed in prose at all.
the failure this replaces ┌──────────────┐ an essay ┌──────────────┐ substring ┌────────────┐ │ a reviewer │───────────>│ a grep │─────────────>│ WRONG │ └──────────────┘ └──────────────┘ └────────────┘ "I do not approve" matched the word approve the contract ┌──────────────┐ one token ┌──────────────┐ a field ┌────────────┐ │ a reviewer │───────────>│ a check │─────────────>│ a decision│ └──────────────┘ └──────────────┘ └────────────┘ the decoder cannot emit anything but pass, fail or escalate
A 3B model choosing between two enumerated values is a different object from a 3B model asked for an opinion. The first is frequently good enough; the second never is. That distinction, not parameter count, is what makes small models usable as deciders.
Tiny language models
The lineage runs from n-grams through Word2Vec and the RNN and LSTM era to the transformer, and the constraint that motivates going small is blunt: weights at GPT-3 scale are hundreds of gigabytes in half precision and simply do not fit on edge hardware.5 Distillation, quantisation and purpose-built small architectures are the three answers, and all three are now mature enough to use without ceremony.
The cascade, which is the shape that matters
Put a tiny model in front of a large one: the tiny model handles the input when it is confident, and falls through to the big model when it is not. This gives large compute reductions with negligible accuracy loss on vision workloads at ImageNet scale.6 The notable design choice there is a fully independent tiny model rather than an early exit branching off the large model's backbone, which avoids the usual problem that intermediate features lack task-relevant semantics. It is the same structure as a boosted cascade in classical detection.7
an input │ v ┌────────────────────┐ confident ┌────────────────────┐ │ tier 1 embeddings │───────────────>│ a decision │ │ plus a linear head │ │ milliseconds, CPU │ └────────────────────┘ └────────────────────┘ │ v unsure ┌────────────────────┐ confident ┌────────────────────┐ │ tier 3 a 3B model │───────────────>│ a decision │ │ decoder = the enum │ │ about one second │ └────────────────────┘ └────────────────────┘ │ v escalate ┌────────────────────┐ │ a frontier model │ └────────────────────┘ only the hard cases reach the expensive tier
Non-generative classifiers
Prior art for "do not generate text to make a decision, classify instead". These run in production at scale today, and they are the existence proof that a sub-two-megabyte purpose-built model beats both a hand-written heuristic and a call to a language model.
| system | size | latency | task |
|---|---|---|---|
| Magika | ~1 MB, ONNX | single-digit ms | file type from sparse byte samples, 100+ types |
| YARA-X | compiled rules | near-instant | malware and structural signature matching |
| fastText | <10 MB compressed | microseconds | language ID across 170+ languages; general linear text classification |
| TrID | signature database | fast | binary file identification by statistical point scoring |
libmagic | tiny | instant | magic bytes at fixed offsets; the deterministic fallback |
Two mechanisms here are worth stealing directly, independent of the models. Sparse ingestion: read only the first, middle and last few kilobytes rather than the whole payload, which is how a file classifier stays fast on a large file. A deterministic fallback underneath a learned one: magic bytes catch the easy cases and the model handles the rest, which is the cascade again at a smaller scale.
The System 1 class
A category named after Kahneman's fast, intuitive cognition: a model that is non-autoregressive, evaluates a whole schema of typed questions in one parallel forward pass rather than generating text token by token, and is trained so its stated probabilities are calibrated. The category reduces decision-making to three question shapes, and those three are the right API to build against whatever sits behind them.
| primitive | asks | returns | consumed as |
|---|---|---|---|
| choice | pick one of N | the option, and a probability over all N | an exit code: the index |
| yes/no | is X true | the probability of yes | a gate: above threshold or not |
| score | rate on a scale | a number | a tally, or a ranking |
The three candidates
| candidate | what it is | verdict |
|---|---|---|
| Jev | hosted, proprietary, no open weights, priced per decision | Not local a hosted service in a decision path is a dependency, not a component |
| Laya | 421M bidirectional encoder, Apache 2.0, about 33 ms, ONNX runner exists, degrades past roughly 20 options | The local pick the right shape with no API key, but an encoder embeds an input, it does not read one |
| SimpleJev | a technique, not a model: read next-token logits per question from any open model | Needs raw logits so a runtime that exposes them, not a chat API that hides them |
The distinction that decides between them: a bidirectional encoder produces a representation of a text and classifies it, which is ideal for "is this finding a blocker" and useless for "does this change introduce a use-after-free". A position that needs comprehension of the input wants an instruct model with its decoder constrained. The two are complementary rather than competing.
Security note that generalises. That model ships as a reinforcement-learning framework archive, which contains pickled Python objects. Loading one executes the author's code. The same is true of any
.pkl
from a source you do not control, including artefacts your own tooling produces.
Study a stranger's model in a throwaway environment or not at all.
Compression represents intelligence
The deepest available argument for the small-model thesis is information-theoretic. If capability tracks how well a model compresses, then parameter count is a proxy and not the quantity itself, and a small model trained tightly on a narrow distribution is not a degraded general model but a well-compressed specialist.
The empirical version reports a near-linear relationship between compression ability and benchmark performance across a range of models.10 The epistemic version argues compression is foundational to intelligence rather than merely correlated with it.11 Taken together they reframe the question from how large to how well matched to the distribution, which is exactly the bet a specialised graph makes.
Prompt compression, the same idea pointed the other way
A small encoder-style token classifier can decide which tokens in a prompt carry information and drop the rest, typically two to five times compression with little loss.12 Architecturally that is a small classifier used to shrink a large model's input rather than to replace the large model. Worth measuring anywhere several positions read the same large input, which in a review panel is the common case.
The catalogue
Everything below is something you could actually put behind a decision. Grouped by what it does, because that is how you choose one.
Generative small models
For a position that must read something and judge it.
| family | compact members | why you would pick it |
|---|---|---|
| Qwen | 0.5B / 0.6B / 1.5B / 1.7B / 3B, plus coder variants at 1.5B and 3B | the strongest coding ability per byte at the small end, long context, permissive licence. The coder variants are the default reviewer candidate on a 4 GB card |
| Llama | 3.2 at 1B and 3B | built for on-device, with the widest tooling support, so it runs on anything |
| Phi | Phi-4-mini around 3.8B, Phi-3.5-mini | heavily trained on synthetic reasoning data; punches above its size on structured tasks, weaker on open-ended prose |
| Gemma | 270M, 1B, 2B, 4B | the 270M is small enough to be interesting as a classifier substrate; the 1B and 2B are genuine instruction-followers |
| SmolLM | SmolLM2 1.7B, SmolLM3 3B | fully open training data and recipe, which matters if you need to explain or reproduce behaviour |
| Granite | 2B and small mixture-of-experts variants | aimed at enterprise retrieval and security work; unusually clear licensing and provenance |
| Ministral | 3B, 8B | edge-focused with very long context for the size |
| DeepSeek distilled | R1-Distill-Qwen-1.5B and relatives | reasoning traces distilled into a tiny model; strong where the answer is verifiable |
| Falcon | the tiny series | worth watching; check current parameter counts before citing |
| OLMo | 1B | fully open corpus and checkpoints; the reference point for reproducibility work |
On a 4 GB consumer card with roughly 3.6 GB usable, 3B at 4-bit quantisation is the practical ceiling. Anything larger wants a machine with real memory.
Encoders for classification
For a decision with a stable label set. Fine-tune on a few hundred labelled examples, export to ONNX, and run with no deep-learning framework at all.
| model | size | note |
|---|---|---|
| BERT base | 110M | the reference architecture; everything below is a response to it |
| DistilBERT | 66M | 40% smaller, substantially faster, keeps most of the accuracy. The safe default |
| TinyBERT | 14M | when 66M is still too much. Edge and microcontroller territory |
| ALBERT | 12M+ | parameter sharing; small on disk, not proportionally faster at inference |
| ELECTRA small | 14M | trained by discriminating replaced tokens rather than masking; strong for the size |
| RoBERTa | 125M+ | better-trained BERT; the base for many ready-made fine-tuned classifiers |
| DeBERTa v3 small | 44M | disentangled attention; consistently the best accuracy per parameter here for language understanding |
| ModernBERT | 149M / 395M | the 2024 refresh: long context, faster inference, drop-in. Prefer it for new work |
Embeddings, rerankers and entailment
| model | kind | note |
|---|---|---|
| all-MiniLM-L6-v2 | bi-encoder | ~90 MB. The classic default: fast, good enough, runs anywhere |
| bge-small-en-v1.5 | bi-encoder | ~130 MB. Better retrieval quality at a similar size |
| E5 small, GTE small | bi-encoder | direct competitors to bge-small. Benchmark all three on your own labels; the ranking flips by domain |
| bge-m3 | bi-encoder | multilingual and multi-granularity when you need it |
| nomic-embed-text | bi-encoder | long context, open training data |
| SetFit | method | contrastive fine-tuning over any embedder; useful accuracy from 8 to 16 examples per class13 |
| bge-reranker | cross-encoder | scores a query and candidate jointly. Slower than comparing embeddings and much more accurate. This is the score primitive off the shelf |
| jina-reranker | cross-encoder | same shape, different training. Worth comparing directly |
| bart-large-mnli | entailment | the original zero-shot classifier: scores "this text entails LABEL" per label, with no training at all |
| deberta-v3 zeroshot | entailment | smaller and usually better than the bart baseline for the same job |
Guard and policy classifiers
An under-used category, and exactly the right shape for an admission check on untrusted input.
| model | note |
|---|---|
| Llama Guard | classifies prompts and responses against a taxonomy you can edit |
| ShieldGemma | the same job on a Gemma base, in several sizes |
| Prompt Guard | a small classifier specifically for injection and jailbreak detection. Tiny, fast, and the right thing in front of untrusted input |
Beyond text
| model | size | note |
|---|---|---|
| Moondream2 | ~1.8B | a position that can look at a screenshot or a diagram on modest hardware |
| SmolVLM | 256M to 2.2B | the smallest genuinely usable vision-language family |
| Florence-2 | 230M / 770M | detection, captioning and OCR in one small model |
| Whisper tiny / base | 39M / 74M | speech to text, if anything needs to read audio |
Constrained decoding
Not models, but what makes a small generative model safe as a decider.
| tool | note |
|---|---|
| Outlines | regex and JSON-schema guided generation; the reference implementation of the idea14 |
| XGrammar | fast grammar-constrained decoding, integrated into several serving stacks |
| llama.cpp GBNF | grammar files native to llama.cpp, and the practical path to reading raw logits that chat APIs hide |
| Structured output APIs | a JSON schema in the request. The easiest version, and enough for most work |
In practice
Five tiers, cheapest first
The discipline is to start at the top and descend only when the tier above demonstrably cannot do the job. Every step down costs an order of magnitude of latency and buys correspondingly less determinism.
| tier | what it is | cost | training |
|---|---|---|---|
| 1 | sentence embeddings plus logistic regression | milliseconds, CPU | 20 to 50 labelled lines per class |
| 1.5 | a purpose-built decision encoder | tens of milliseconds | none |
| 2 | zero-shot entailment | slower than tier 1 | none |
| 3 | a small instruct model, decoder constrained to the enum | about a second on a small GPU | none |
| 4 | a fine-tuned small encoder, exported to ONNX | best accuracy per millisecond | a few hundred labelled examples |
Four rules for tier 3
- Temperature zero gives an answer, not a confidence. One deterministic sample is always certain of itself. For a probability, sample five to nine times at a nonzero temperature and count the votes.
- Put the enumeration in the schema, not the prompt. The prompt is a request; the schema is an enforcement.
- Ask for a reason field even if you discard it. It costs tokens, it makes a wrong answer diagnosable, and a model required to justify tends to answer better.
- A large enumeration is the wrong tier. Dozens of labels slow decoding and hurt accuracy; that is tier 1 work.
Calibration, the part nobody checks
A probability is only useful if 0.8 means right about eight times in ten. Logistic regression is roughly calibrated by construction; vote counts from a sampled language model are not, and an entailment score is not a probability at all.
- Hold out thirty to fifty labelled examples the model has never seen.
- Bucket predictions by stated confidence in bands of 0.1 and compute the actual accuracy within each bucket. A well-calibrated model's buckets sit on the diagonal.
- Compute the Brier score,15 the mean squared difference between stated confidence and outcome. Lower is better, and 0.25 is what a coin flip scores.
- Correct what you find. An overconfident linear head wants stronger regularisation; an overconfident sampled model wants more samples, or a decision threshold raised to 0.7.
Choosing, in four questions
| question | if yes |
|---|---|
| Is the label set stable and small? | a trained classifier or a purpose-built one. Do not use a language model for this |
| Does the decision require understanding the input? | a generative small model with constrained decoding |
| Do you need a calibrated probability rather than a label? | a decision encoder, a logistic head over embeddings, or a reranker |
| Is the input untrusted? | put a guard classifier in front, as a gate, not as a participant |
Glossary
| term | meaning |
|---|---|
| SLM | a language model small enough to run cheaply and locally, traded against a frontier model's raw capability |
| System 1 model | a non-autoregressive model returning typed decisions in one parallel pass instead of generating text |
| constrained decoding | sampling restricted to tokens that keep the output valid against a schema or grammar, so an enumerated answer is guaranteed rather than requested |
| logits | raw per-token scores before they become probabilities. Reading them per option gives a true one-pass distribution, which is why an API that hides them cannot do the trick |
| bi-encoder | embeds query and candidate separately and compares vectors. Fast enough to search millions; less accurate than a cross-encoder |
| cross-encoder / reranker | scores a pair jointly. Slower, much more accurate, and the score primitive available off the shelf |
| zero-shot NLI | scoring "this text entails LABEL" with an entailment model, so a classifier exists before any training data does |
| guard classifier | a small model judging whether input or output violates a policy. Belongs in a gate, never in a reporting position |
| calibration / Brier score | whether stated confidence matches observed accuracy; mean of (confidence − outcome) squared, where 0.25 is a coin flip |
| quantisation (Q4, Q8) | storing weights at reduced precision so a larger model fits a smaller card. Q4 is what puts a 3B model on 4 GB |
| distillation | training a small model on a large model's outputs. Where most compact reasoning models come from |
| early exit vs cascade | an early exit taps an intermediate layer of one large model; a cascade runs a fully separate small model first, whose features are complete rather than intermediate |
| sparse ingestion | classifying from a few byte ranges rather than the whole payload |
| magic bytes | fixed-offset signature bytes for deterministic file typing; the fallback beneath a learned classifier |
References
- Belcak, P., Heinrich, G., Diao, S., Fu, Y., Dong, X., Muralidharan, S., Lin, Y. C., Molchanov, P. Small language models are the future of agentic AI. arXiv:2506.02153, 2025.
- Kuuspalu, A., Cody, S., Hale, M. E. Multiple nerve cords connect the arms of octopuses, providing alternative paths for inter-arm signaling. Current Biology, 2022.
- Goldfeder, J., Wyder, P., LeCun, Y., Shwartz-Ziv, R. AI must embrace specialization via superhuman adaptable intelligence. arXiv:2602.23643, 2026.
- Han, P. et al. Modular cognitive architecture emerges in large language models. Preprint.
- Penman, T. Tiny language models. tristanpenman.com, 2025.
- Wang, J. et al. Tiny models are the computational saver for large models. arXiv:2403.17726, 2024.
- Viola, P., Jones, M. Rapid object detection using a boosted cascade of simple features. CVPR, 2001.
- Goedecke, S. Jev means structured output is interesting again. seangoedecke.com, 2026.
- DeepMost Innovations. Sales conversion prediction with reinforcement learning. arXiv:2503.23303, 2025.
- Huang, Y. et al. Compression represents intelligence linearly. arXiv:2404.09937, 2024.
- The information-theoretic imperative: compression and the epistemic foundations of intelligence. arXiv:2510.25883, 2025.
- Jiang, H. et al. LLMLingua: compressing prompts for accelerated inference. arXiv:2310.05736, 2023. See also LLMLingua-2, arXiv:2403.12968, 2024.
- Tunstall, L. et al. Efficient few-shot learning without prompts (SetFit). arXiv:2209.11055, 2022.
- Willard, B. T., Louf, R. Efficient guided generation for large language models. arXiv:2307.09702, 2023.
- Brier, G. W. Verification of forecasts expressed in terms of probability. Monthly Weather Review 78(1), 1950.
- Curated lists, for going wider: Awesome-LLM and Awesome-SLM, the latter organised into milestone papers, a leaderboard and an open-model catalogue by vendor.