advanced Estimated learning time: 3 h

12.6 Small models and on-device AI

You can choose a small model where a large one is waste.

Before:08. Large Language Models

Small language models cover a surprising share of real tasks at a fraction of the cost, and on-device inference — phones, browsers through WebGPU — adds privacy and offline arguments that matter particularly in India's connectivity landscape. It sits in the frontier module as the counterweight to scale. The anchor to drop is leaderboard position: capability per parameter keeps improving, and the right question is whether this model handles this task, not where it ranks.

Work through these

  • SLM families and capability per parameter

    Smaller model families and the observation that capability per parameter has improved considerably. For many tasks a small model is now sufficient, which changes the deployment question.

  • On-device inference on phones and browsers

    Running models directly on phones and inside browsers, without a server. This is a genuine architectural option now rather than a demonstration.

  • WebGPU and in-browser models

    The browser interface for accelerated computation, which is what makes in-browser models practical. It removes the server from the picture entirely.

  • Privacy and offline arguments

    Keeping data on the device and working without a connection are arguments that hold regardless of capability. For some applications they are the deciding factor.

Sign in to keep your progress.

Free resources

Links last checked 29 Aug 2026.

Stuck here?

Ask a mentor. A real person answers, and they can see exactly which topic you're on. Usually within a couple of working days.

Checking your session…

Topics shown in module order.