Other meanings of Phase vocoder
Audio signal processing
Phase vocoder is an audio signal-processing technique for time stretching and pitch shifting without directly changing a recording’s sample rate. It analyzes short, overlapping frames with a Fourier transform, modifies their spectral timing or frequency relationships, and reconstructs the result while tracking phase continuity.1 The method is powerful for sustained, harmonic sounds but can produce characteristic smearing, reverberant transients, or “phasiness” when its assumptions fail.
The phase vocoder represents sound as evolving spectral magnitude and phase, making it possible to manipulate duration and pitch independently.1 In the classic design, a signal is divided into overlapping windows and transformed into frequency bins by the short-time Fourier transform (STFT). Each bin supplies an amplitude and a phase; changes in phase from one frame to the next estimate the component’s instantaneous frequency rather than merely its nominal bin center.
Flanagan and Golden introduced the technique in 1966 for speech analysis and synthesis.1 Later work made it practical for music and general audio, especially through improved phase estimation, peak tracking, and transient handling. The name describes the analysis-and-resynthesis method, not a particular software product or a vocal effect.
Time stretching is achieved by changing the spacing of analysis frames while preserving estimated component frequencies during synthesis.2 If frames are synthesized farther apart, the sound lasts longer; if they are packed closer, it becomes shorter. The process attempts to preserve the phase trajectories that convey sinusoidal motion, rather than simply repeating waveform segments.
Pitch shifting normally combines two operations: stretch or compress the duration, then resample to restore the original length. Resampling alone changes pitch and duration together, whereas the phase vocoder separates those dimensions. The method is especially effective for polyphonic, sustained material such as pads, strings, and speech vowels, but pitch changes may require frequency reassignment or peak-based processing to avoid blurred harmonics.
Accurate reconstruction depends on consistent phase propagation across overlapping frames.2 A basic implementation unwraps phase differences, corrects for the expected advance between frames, and integrates the resulting frequency estimates during synthesis. The overlap-add stage then combines windowed frames into a continuous waveform.
Artifacts arise because a frequency bin is not necessarily one stable sinusoid. Transients spread energy across many bins, and independent manipulation of those bins can destroy attacks or make percussion sound watery. Rapidly changing mixtures can also develop phasiness when unrelated components acquire artificial coherence. Identity phase locking, sinusoidal peak tracking, transient detectors, and multiresolution analysis address these limitations. Modern systems may combine phase-vocoder analysis with time-domain or source-separated processing rather than rely on one representation for every sound.
The phase vocoder was originally a speech technology before becoming a familiar music-production tool.1 Its early motivation was not creative pitch manipulation but a compact description of speech’s changing spectral envelope and excitation-related structure. This origin explains why speech remains an important test case: voiced segments are comparatively harmonic, while consonants expose timing and transient weaknesses.
A less obvious design choice is the trade-off between time and frequency resolution. Long windows distinguish nearby pitches but smear short events; short windows preserve attacks but provide poorer frequency discrimination. Some implementations use separate analysis paths or adaptive windows for bass, vocals, and percussion. Phase-vocoder concepts also appear in speech coding, forensic audio, restoration, sound design, and music information retrieval. The technique therefore remains a general framework for interpreting and resynthesizing changing spectra, not merely a tempo-control effect.
Terminology varies across implementations: some systems reserve “phase vocoder” for STFT-based analysis and resynthesis, while others use it broadly for related spectral time-scale modification methods.
Help improve the encyclopedia. Reports go straight to the site manager.