← New search

Audio Technology

Spatial audio

Spatial audio is a set of techniques that reproduce sound with a sense of three-dimensional space, placing sounds anywhere around the listener. Unlike conventional stereo, which creates a flat soundstage in front of the listener, spatial audio aims to create an immersive, 360-degree sound field, making listeners feel as though they are inside the recorded or synthesized environment. The technology relies on principles of psychoacoustics and signal processing to encode directional cues, such as interaural time differences and spectral filtering, that the human brain uses to localize sounds. Over recent years, spatial audio has moved from specialty cinema formats to consumer headphones, smartphones, and home theater systems.

360°
Degrees of surround field
Soundfield coverage
3D
Spatial dimensions
Immersive reproduction
5.1
Common home cinema channels
Surround format
64
Maximum channels in some formats
Object-based mixing
1

Core principles and psychoacoustics

Spatial audio exploits the human auditory system's ability to determine the direction of a sound source. The primary cues are interaural time differences (ITD), which arise because a sound reaches the closer ear slightly earlier, and interaural level differences (ILD), which result from the head's acoustic shadow at higher frequencies.1 The outer ear, or pinna, also filters sounds in a direction-dependent manner, imparting spectral cues that help localize elevation and front-back position.1 Binaural recording, which uses microphones placed in a dummy head's ear canals, captures these cues naturally; when played over headphones, the listener perceives a three-dimensional sound field. Synthetic spatial audio systems, however, must simulate these cues through signal processing by filtering sounds with head-related transfer functions (HRTFs).2 These HRTFs are measured from physical mannequins or listeners and encode the direction-specific filtering of the human anatomy.

2

Ambisonics and object-based audio

Classic surround sound formats, such as 5.1 and 7.1, assign audio to a fixed number of loudspeakers arranged around a listener, a channel-based approach that does not adapt well to non-standard playback layouts. Ambisonics offers a more flexible hierarchical representation of a sound field independent of the playback array. It encodes the field as spherical harmonic coefficients up to a given order, which can be decoded for any speaker configuration or over headphones.3 Object-based audio goes further by representing each sound as an individual object, accompanied by positional metadata, which allows the renderer to place objects arbitrarily in space and scale them to the available speakers. Dolby Atmos and MPEG-H are prominent object-based formats that combine objects, traditional beds, and metadata to enable scalable immersive mixes.4 These standards simplify the creation pipelines in cinema, gaming, and music production, allowing artists to position sounds dynamically in space.

3

Consumer applications and headphone rendering

Consumer adoption has accelerated with the rise of virtual and augmented reality, as well as streaming services. Smartphones and tablets now support spatial audio via built-in head-trackers, which rotate the rendered sound field to correspond with the listener's head movements, enhancing the sense of presence.5 Streaming platforms, notably Apple Music and Tidal, have introduced spatial audio mixes, often in Dolby Atmos, that are downmixed to binaural for headphone listening. Personalized HRTFs, measured using smartphone cameras or approximated from anthropometric parameters, address the inter-individual variability in pinna shape that reduces the accuracy of generic HRTF rendering.2 Gaming engines integrate spatial audio natively for 3D sound cues, which have been shown to improve directional awareness and reaction times in virtual environments.6 Audiovisual streaming platforms and broadcasters are also transitioning to immersive audio for live sports and concerts.

4

Lesser-known aspects

Early efforts at spatial audio date back to the 1930s, when Alan Blumlein patented binaural and stereo recording techniques in the United Kingdom. In the 1970s, the German company Carsten & Scannerz developed the first commercial ambisonics system, known as Harplay.3 The development of the Apple Music's spatial audio feature leveraged binaural rendering from a virtual stereo mix, using directional metadata embedded in the original recordings.5 Psychoacoustic research has revealed that head-tracking can counteract the front-back confusion often experienced with static HRTF rendering, because small movements help resolve ambiguity.1 Binaural room impulse responses (BRIRs) extend HRTFs to include the acoustics of a particular room, enabling realistic rendering of reverberation in environments like churches or concert halls.1 In broadcasting, the BBC has developed free experimental spatial audio plugins for mixing content, highlighting the technology's democratization among hobbyists.7

Glossary

Interaural time difference (ITD)
The delay between the arrival times of a sound at the two ears, a major cue for horizontal sound localization.
Head-related transfer function (HRTF)
A filter that describes how the anatomy of the head, pinna, and torso transforms sound as it travels from a direction to the eardrum.
Ambisonics
A technique for representing a full sphere of sound using spherical harmonic coefficients, independent of the playback array.
Object-based audio
An audio format in which sounds are treated as individual objects with positional metadata, rendered dynamically on the playback system.
Downmix
The process of converting a multichannel audio mix to a smaller number of channels, e.g., from spatial to binaural stereo.