the.ai

Speech / Representation

verified

Spectrogram

A waveform is tens of thousands of numbers a second and almost none of them mean anything alone. Chop the signal into short overlapping windows, ask each window which frequencies it holds, and stack the answers left to right. Time runs across, pitch runs up, and speech becomes something you can look at. Nearly every speech system starts here rather than at the raw samples.

Viz primitive · update-spectrumtime-resolution = 0.15 · bars = 16

16 values. The left group decays steeply; the right group is 28% of the way to flat, and reads flatter than the left.

One frame's energy across frequency, at this node's analysis window and at a shorter one. Drag the time resolution up to watch the frequency detail smear toward its own average — sharpening one axis is exactly what blunts the other.

0.15

Reviewed by opendroid · 2026-08-18

  • arXiv:1512.02595 — Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
  • arXiv:1904.08779 — SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Origin · not linkable

  • Griffin & Lim 1984 — Signal Estimation from Modified Short-Time Fourier Transform · IEEE Transactions on Acoustics, Speech, and Signal Processing 32(2) · doi:10.1109/TASSP.1984.1164317