core Estimated learning time: 3 h

11.3 How speech is produced and classified

You can explain how a speech signal is made and name the sound classes any acoustic model has to tell apart.

Before:06. Deep Learning

Speech is a source driving a filter: the larynx or turbulence makes the sound, the vocal tract shapes it, and the shape changes several times a second. That one model explains why formants identify vowels, why voiced and unvoiced sounds look so different on a spectrogram, and why speech features are built the way they are. It sits just before the audio ML topic so that what the models consume is something whose physics you already understand.

Work through these

  • Trace the production chain: lungs, larynx, vocal tract, radiated sound

    Air from the lungs, a source at the larynx, shaping through the vocal tract, and radiation from the lips. Following the chain once explains most of what a speech signal contains.

  • Use the source-filter model to say what each stage contributes

    Separating what makes the sound from what shapes it is the model that underlies speech coding, synthesis and recognition. It is the single most useful abstraction in the topic.

  • Classify speech sounds - voiced and unvoiced, vowels, plosives, fricatives

    The main categories of speech sound, distinguished by whether the vocal folds vibrate and how the airflow is obstructed. These are the classes an acoustic model has to separate.

  • Recognise each class by its signature in a spectrogram

    Each class has a visible signature in a time-frequency display, and learning to read them is what makes speech data interpretable rather than opaque.

Sign in to keep your progress.

Free resources

Links last checked 29 Aug 2026.

Stuck here?

Ask a mentor. A real person answers, and they can see exactly which topic you're on. Usually within a couple of working days.

Checking your session…

Topics shown in module order.