12.6 Small models and on-device AI
You can choose a small model where a large one is waste.
Before:08. Large Language Models
Small language models cover a surprising share of real tasks at a fraction of the cost, and on-device inference — phones, browsers through WebGPU — adds privacy and offline arguments that matter particularly in India's connectivity landscape. It sits in the frontier module as the counterweight to scale. The anchor to drop is leaderboard position: capability per parameter keeps improving, and the right question is whether this model handles this task, not where it ranks.
Work through these
SLM families and capability per parameter
Smaller model families and the observation that capability per parameter has improved considerably. For many tasks a small model is now sufficient, which changes the deployment question.
On-device inference on phones and browsers
Running models directly on phones and inside browsers, without a server. This is a genuine architectural option now rather than a demonstration.
WebGPU and in-browser models
The browser interface for accelerated computation, which is what makes in-browser models practical. It removes the server from the picture entirely.
Privacy and offline arguments
Keeping data on the device and working without a connection are arguments that hold regardless of capability. For some applications they are the deciding factor.
Sign in to keep your progress.
Free resources
We haven't checked most of these for screen reader use yet.
Links last checked 29 Aug 2026.
Stuck here?
Ask a mentor. A real person answers, and they can see exactly which topic you're on. Usually within a couple of working days.
Checking your session…
Topics shown in module order.