Introduction to Spatial Audio
- Spatial audio encompasses the strategies that let an audio stream recreate a virtual environment, or place virtual sources within a virtual space.
- It evolved mono → stereo → surround → immersive, along two delivery paths: loudspeakers and headphones.
- Everything reduces to three representations: channel-based, object-based and scene-based. And to one idea: encode, then decode.
- This chapter is for beginners; it assumes no prior knowledge.
This page is the front door to the whole guide. By the end you will know what we mean by “spatial”, why it is worth the trouble, how the field grew from a single channel to fully immersive systems, the two physical paths sound can take to your ears, the three ways engineers represent a spatial scene, the single unifying idea, encode, then decode, that ties every technique together, and finally how a position is written in numbers. Later pages go deep; this one gives you the map. Here is that map, and how to read it.
How to use this guide
You can read this guide straight through, like a course: each part assumes the one before, and the level rises steadily from this introduction to the working detail a practising engineer needs. You can equally arrive from a search, read one chapter and leave: each chapter states what it assumes and links back to whatever it depends on.
- Part I, Fundamentals (you are here) builds the foundation. After this introduction, Spatial Psychoacoustics explains how the ear and brain locate sound: the interaural cues, the role of the outer ear, the limits of human hearing. Every technique is an exploitation of those mechanisms. Multichannel & Immersive Formats then makes the three representations precise, and the Glossary collects the vocabulary.
- What follows. Techniques (stereo, amplitude panning, surround and matrix, object-based audio, Ambisonics, binaural, transaural, wave field synthesis); the sound field and the room (distance and air, source directivity, Doppler, reverberation, envelopment); recording and capture; systems, calibration and installation; production workflows; artistic practice; and worked case studies.
- What you can read today. The guide is published progressively: chapters will go online one after another over the coming months, so not everything listed above is available yet. The sidebar always shows exactly what can be read today: it is the reliable map, not this list.
Mathematics is not a prerequisite: the prose carries the meaning on its own. Equations appear where a relationship has to be exact, and every mechanism is backed by a worked numeric example. Unfamiliar symbols are explained where they occur.
One orientation before the material begins. Everything that follows rests on a single fact of perception: with two ears, a listener reconstructs an entire world of directions and distances. Every microphone technique, every loudspeaker array and every line of rendering code exists to deliver to those two ears, directly or through the air of a room, the signals that will make a scene appear. Keep that in view and the rest of the guide has a single subject rather than a catalogue of methods.
What “spatial” actually means
Sound is spatial when it comes from somewhere. A single loudspeaker plays everything from one point; spatial audio gives each sound its own place again: a direction, a distance, and a room around it. More loudspeakers alone do not make a system spatial: what counts is that each sound has its place.
Behind spatial audio lies a simple wish: to create experiences that feel close to real, and to carry listeners into worlds they could not visit otherwise. Modern tools make both possible, whether by placing sources realistically over a set of loudspeakers, or by building virtual spaces from scratch.
For a long time this belonged to a few specialist rooms. In about a decade it has become ordinary: cinemas mix with objects, streaming services deliver immersive music to phones and headphones, concert stages use spatial mixing systems, and VR headsets and earbuds follow the movements of the head.
So what is it for, and why so much attention now?
Why spatial audio matters
The question is fair: mono recordings carried the twentieth century’s music perfectly well. Three reasons explain why space is worth the engineering effort, and each shows up in everyday listening.
Realism and immersion. The physical world is spatial, so a spatial reproduction stands closer to the thing it represents. A storm passing overhead, a vehicle crossing from front-left to rear-right, the acoustic signature of a hall: each is convincing because it matches the cues the ear expects. Immersion is not only spectacle: the sense of being enveloped by a space, surrounded by its reverberant field, is what separates the reproduction of a room from the experience of being in one. Envelopment has its own chapter, Direct, Diffuse & Envelopment.
Clarity through separation. Spatial separation is of engineering interest even where realism is not the objective. Two sounds sharing a location and a frequency range compete: the louder masks the quieter, making it harder or impossible to resolve. Separating them in space relaxes that competition. The best-documented case is the cocktail party effect, named by Cherry in 1953: a listener follows one conversation among many. Direction is one of the cues that make this possible, and separating two talkers in space yields a measurable gain in intelligibility, alongside differences in pitch, in onset timing, and in familiarity with the voice. In practice, dialogue anchored at the centre, music spread to the sides and effects placed overhead are each easier to follow than the same elements stacked in one position. Spatialisation is, among other things, a tool for intelligibility.
Scale. The third case concerns extent. Stadium concerts, outdoor festivals, planetaria and themed attractions distribute many loudspeakers over a large area, not only to reach a required level but to give an audience of thousands a coherent scene. There is a second reason, and it is the decisive one at that size. Stereo and its descendants build an illusion optimised for one position: the image holds in a narrow sweet spot and decays away from it. Several of the techniques in this guide instead reconstruct the field itself across an area, so the listening zone widens from a seat to a room. That is why they have taken hold in live sound, in scenography and in museum work, where the audience is spread out and moving, and where an image that only holds at one seat is no image at all. At that scale the room and the playback system stop being neutral and become part of the instrument, a theme running through Part III, The Sound Field & the Room, and through Systems, Calibration & Installation.
Space earns its place for three reasons: realism is about matching the world, clarity about organising a mix, scale about filling a venue. Most projects want a blend of all three.
From one channel to a whole space
The history of reproduction is the story of adding spatial dimensions one at a time, each addition answering a constraint of its era rather than a sudden ambition. Following it in order is the quickest way to understand why today’s formats look the way they do.
Mono: one channel. A single loudspeaker reproduces one stream of audio. Everything arrives from the same point, so nothing in the signal carries a location: the position of every instrument is the position of the speaker. The constraint was technical before it became a style: one microphone, one channel, one groove. Its consequence has outlived its cause. As soon as a mix grows dense, sources sharing a point and a frequency range mask one another and the result turns opaque; the practices that grew out of mono (sparse arrangements, parts answering rather than sounding together, equalisation carving room for each element) are all ways of buying back by other means the separation that space would have supplied. Mono remains the correct choice wherever intelligibility outweighs image, in much broadcast and public address, and it is not in itself a limit on quality.
1881: Ader’s binaural demonstration. At the International Exposition of Electricity in Paris, Clément Ader placed some eighty telephone transmitters across the front of the Opéra stage and fed them in pairs down telephone lines to listeners holding one receiver to each ear. Listeners reported that performers seemed to move across an audible stage. This is the first documented demonstration of binaural transmission, two channels one per ear, and it establishes the founding insight of the field half a century before stereophony: deliver the right signal to each ear and the brain reconstructs space. The name usually attached to it, théâtrophone, belongs to what came after: the Compagnie du Théâtrophone was founded in 1890 to sell the experience by subscription, and ran until 1932.
1931: Blumlein, and the phantom image. Alan Blumlein filed a patent that set out two-channel stereophony with remarkable completeness: how to capture a stage with coincident microphones, how level and phase differences between two channels produce a stable image, and how to cut the result into a record groove. The mechanism he described is the one every stereo recording still relies on. Feed two loudspeakers in front of a listener with related but distinct signals, and a sound sent equally to both is heard at neither but between them: the phantom image, a source with no physical existence heard at a definite place. Adjusting the level, and within limits the timing, slides that image along the line joining the speakers, turning two channels into a continuous horizontal stage. The level difference is a matter of geometry: with two figure-of-eight capsules crossed at 90°, a source 20° off centre reaches one of them 6.6 dB louder than the other. Blumlein also understood that loudspeaker stereo works on a different principle from Ader’s earphones, because each ear hears both speakers. It has its own chapter, Stereo & the Phantom Image.
Two channels, and what they do not guarantee. A second channel adds no information by itself. A release carrying the same signal on both sides, double mono, has two channels and one position. The opposite failure was just as common: the two-track pop stereo of the early 1960s put the voice hard on one side and the band on the other, which is two channels and still no coherent scene. Channels are a transport; space is what the differences between them encode.
Adoption took far longer than the technology. Stereo recording was working by 1954, at RCA Victor in the United States and, three months later, at Decca in Geneva, with the microphone array still known as the Decca tree. The first commercial stereo LPs followed in 1958, on classical and operatic repertoire, where the format was sold as a mark of prestige. Popular music kept mono as its reference for another decade: at Abbey Road the mono mix was the one the artists supervised and the stereo a rushed afterthought, and the last Beatles album mixed in mono was The Beatles, in 1968. Records labelled “electronically reprocessed for stereo” were still on sale in the 1970s. A format existing and a format being used are two different dates, and this guide will meet that gap again.
Capturing a scene, or building one. Blumlein answers one question: how to capture a scene that already exists, with two microphones close together whose difference in level does the work. A second family of answers grew beside it, using spaced microphones, where the difference is one of arrival time rather than level. Both are taken apart in Recording & Capture. The other question is how to build a scene that never existed, and its first electronic answer arrived with the next entry on this list.
1940: Fantasound, and the pan pot. Disney, with RCA, built a custom system to present Fantasia: separate optical soundtracks drove speakers placed around the auditorium. Its lasting contribution turned out to be a control rather than a format. To move a sound between those speakers, the engineers built a panoramic potentiometer, a differential network that splits one signal between several outputs from a single gesture. The pan pot is still found on every channel of a mixing console, and it is the first electronic device for placing a sound where no microphone ever stood. With it, the split named above becomes concrete: a scene can be captured, or it can be constructed, and from 1940 onward both are available at once. Fantasound itself ran only as a roadshow, on a handful of installations, and few audiences heard it; Disney, Garity, Hawkins and RCA received an honorary Academy Award for it in 1942.
1973: Ambisonics. Michael Gerzon and colleagues in Britain proposed a principled alternative to counting channels. Rather than assigning sounds to particular speakers, Ambisonics describes the whole sound field at a point as a set of spherical-harmonic components, which a decoder then renders on any layout, including one with height. Accuracy is bought in components: a first-order decoder aims a beam about 65° wide at each loudspeaker, a third-order one about 39°. It was ahead of its market in the analogue era, but its scene-based philosophy underpins much of today’s virtual reality and 360° audio. It has its own chapter, Ambisonics.
1976: Dolby Stereo. Dolby’s optical system encoded four channels (left, centre, right and surround) into the two-channel optical track that 35 mm film already carried, using a matrix that folded the extra channels in and unfolded them on playback. An earlier three-channel Dolby optical release had appeared with Lisztomania in October 1975; A Star Is Born, in December 1976, was the first to carry the surround channel and therefore the full four-channel matrix described here. Upgrading a cinema no longer meant rewiring its projectors, though it did require a Dolby processor and, usually, surround speakers and their amplification. Star Wars (1977) and the films that followed turned cinematic surround from a curiosity into an expectation.
1992: Dolby Digital, and what “discrete” means. With Batman Returns, Dolby introduced a fully discrete 5.1 format: five full-range channels plus a dedicated low-frequency channel, each carried independently rather than matrixed together. The notation is worth decoding, because it is not a count: the LFE is written as a fraction because it is not a full channel. It is band-limited, typically to about 120 Hz, and carries no direction of its own. Discrete channels meant clean separation and reliable steering, and 5.1 became the backbone of DVD, digital television and home cinema for two decades. As with Fantasound, the format arrived long before the installed base: eleven screens in North America were equipped at release, so almost every 1992 audience heard the analogue mix.
2010 onward: height, and objects. The most recent step adds the dimension stereo and classic surround both lack: elevation. Auro 11.1 was installed at Galaxy Studios in May 2010 and publicly launched that October, and Red Tails, mixed in Auro, reached cinemas in January 2012. Auro is a fixed layout of three layers rather than objects; object handling arrived later under the name AuroMax. Dolby Atmos followed with Brave in June 2012, and married height to a deeper change: instead of mixing every sound permanently into fixed channels, individual sounds became objects, each carrying its own position, to be placed by the playback system on whatever speakers are actually present. A helicopter can be told to circle overhead, and a seven-speaker living room and a sixty-four-speaker cinema each render that instruction as well as their hardware allows. DTS:X and MPEG-H arrived in 2015. These are competing delivery formats, each with its own codec and licensing, not variants of one another.
What they share sits upstream of delivery. A production mix is described in the Audio Definition Model (ITU-R BS.2076), an open XML metadata model naming each audio object, its content and its position, so one session can be exported to any of them. The encode step is therefore the same in principle across the whole era, sources plus metadata, and only the decode differs, per cinema, per living room, per pair of headphones. The mix and the playback layout have become two separate things, and the second is not known when the first is made. Object-based audio has its own chapter, Object-Based Audio & Rendering.
The pattern worth carrying out of this section is not the growing number of channels. It is what sound was tied to. Early systems bound it to specific physical channels: a Blumlein groove, a Fantasound track, a Dolby matrix, an AC-3 stream. Ambisonics broke that tie in the 1970s by describing the field instead of the speakers, and object-based audio broke it again in the 2010s by describing the sources. The pendulum has swung twice and has not settled: current tools work in either philosophy, and most real projects use both at once. That is why the rest of this guide is organised around techniques rather than products.
Two paths to the ears
However a spatial scene is built, it ultimately reaches a listener by one of two physical routes, and the difference between them shapes everything.
Loudspeakers. Sound radiates into a room from two or more speakers, mixes in the air, and reaches both of the listener’s ears. Each ear therefore hears every speaker: the near one directly, and the far one too, a fraction of a millisecond later and darkened by the shadow of the head. That second, unwanted path, left speaker to right ear and right speaker to left, is what crosstalk names, and stereo and surround are designed around its presence. Loudspeaker reproduction is shared by construction, a roomful of people hears it at once; it engages the body with real acoustic energy; and it lets the room’s own acoustics contribute to envelopment. Its weakness is that the result depends on speaker placement, room treatment and where the listener sits: with the conventional layouts the image is best at one sweet spot, the limit that the field-reconstruction techniques of Part II set out to lift.
Headphones. Each earpiece feeds one ear directly, with no crosstalk and no room. This is precise and private, and it makes the listener’s head the entire acoustic stage. But it introduces a problem: with no head and outer ears in the signal path, naively panned sound tends to collapse inside the head rather than out in the world.
The answer is binaural audio, which bakes the filtering effect of a head and outer ears into the signal so that headphones can place sounds outside the listener: overhead, behind, at a distance. It is an answer rather than a fix: with a generic set of filters, sounds still often stay inside the head. Reliable externalisation usually needs three further things: some room response, head tracking, and filters matched to the individual listener. All three are covered in Binaural — Headphone Spatialization.
Headphones are now the dominant listening device for recorded music and for much video, and binaural rendering has become the route by which immersive audio most often reaches a listener. Recreating the same control over loudspeakers, by cancelling crosstalk, is the subject of Transaural - Binaural over Loudspeakers.
A single immersive master is increasingly expected to serve both paths: decoded to a speaker array in a cinema, and binaurally rendered to headphones for everyone watching at home. Keeping the two delivery routes distinct in your thinking explains many otherwise puzzling design choices later in the guide.
The three representations
There are three fundamentally different ways to store and transmit a spatial scene. Almost every format in existence is one of these, or a hybrid. The full treatment is in Multichannel & Immersive Formats; here is the preview.
Channel-based. The scene is stored as one audio signal per loudspeaker. Stereo, 5.1, and 7.1 are channel-based: “this is what the left speaker plays, this is the right”, and so on. It is simple, predictable, and as old as stereo itself, but it assumes the playback system has exactly the speakers the mix was made for.
Play a 5.1 mix on a stereo system and something has to fold it down. The losses are specific: the centre channel is split into left and right, so dialogue stops being anchored and shifts in level relative to the music; the surrounds collapse forward, flattening the depth; and the LFE is usually discarded outright. Knowing which of those matters for the material is what decides whether you accept the automatic downmix or make your own.
Object-based. The scene is stored as individual sounds plus metadata describing where each one should be: its direction, distance and size. There are no fixed channels; at playback time a renderer reads the metadata and computes, in real time, which speakers should produce each sound and how loudly, given the speakers actually present. This is what lets one Atmos master play correctly on a phone and in a cinema. The cost is that playback now requires computation, not just routing.
Scene-based. The scene is stored as a description of the whole sound field at a point in space, independent of any sources or speakers. Ambisonics is the canonical example: a compact set of components captures sound arriving from all directions, and a decoder reconstructs that field on whatever array is available. Scene-based formats are elegant for rotation (turn your head in VR and the whole field rotates by a simple matrix) and for capture (one special microphone records the entire field), at the cost of spatial sharpness that grows only as you add more components.
Channel-based says what each speaker plays, object-based says where each sound is, and scene-based says what the field looks like everywhere.
None is universally best. The right choice depends on whether you are recording reality, building a mix, or targeting playback systems you cannot know in advance.
The unifying idea: encode, then decode
Underneath all this variety lies one pattern that, once seen, makes the whole field click. Almost every spatial-audio technique can be read as two stages:
- Encode. Take the spatial scene, whether sources at given positions or a captured field, and represent it in some intermediate form: a matrix, a set of Ambisonic components, a list of objects with metadata.
- Decode. Take that representation and turn it into actual signals for the actual output, loudspeaker feeds or a binaural pair for headphones, using whatever the playback situation provides.
Ambisonics encodes a field as spherical-harmonic components and decodes them to a speaker array. Object audio encodes positions as metadata and decodes them by panning to present speakers. Binaural encodes direction as ear filters and decodes, trivially, straight to headphones.
Channel-based formats, stereo and its multichannel descendants (5.1, 7.1, 7.1.4), are the nuance. They need no encoding or decoding step: each channel already is the feed of one loudspeaker, and a pan law used at mixing places sounds between them. What replaces the decoder is a convention fixed in advance: the loudspeaker positions are frozen in a standard (ITU-R BS.775 for stereo and 5.1, ITU-R BS.2051 for layouts with height), and the signal is read back correctly only if the speakers stand where that standard puts them. The layout itself is the key that turns the channels back into positions; move the speakers and the scene moves with them.
| technique | encode | representation | decode | output |
|---|---|---|---|---|
| Stereo | none: pan law at mixing | L, R | none: speakers at ±30° | phantom image |
| Channel-based (5.1, 7.1) | none: mixed straight to speaker feeds | one signal per speaker | none: speakers at the standard positions | the layout it was mixed for |
| Ambisonics | spherical harmonics | ACN/SN3D components | decoder matrix | any array |
| Object audio | audio + position metadata | ADM / BW64 | renderer | any array |
| Binaural | HRTF filters | two ear signals | — | headphones |
Separating encode from decode buys flexibility: one encoded master can be decoded many ways, for many rooms and devices, including ones that did not exist when the master was made.
The flexibility stops at channel-based files: as the downmix above showed, an n-channel file on fewer than n outputs has to be folded down, or loses channels outright. Scene-based and object-based formats do not have this failure mode: the decoder reconstructs the scene for whatever layout it finds, rather than discarding channels it cannot use. That difference is the strongest practical argument for describing a scene instead of a set of speakers.
This encode/decode lens is how Part II, Spatialization Techniques, is organised: each technique is a particular choice of how to encode a scene and how to decode it again, or, for channel-based formats, of which loudspeaker layout to fix, and once the pattern is recognised even unfamiliar formats become easy to place.
Geometry of a position: Cartesian and polar coordinates
The chapter has placed sounds in words: left and right, overhead, near and far. Working with them takes numbers, and numbers take a frame.
To say where a sound is, you need coordinates, and there are two families to choose from. Cartesian coordinates give three distances along three fixed axes: left/right, back/front and bottom/top. Polar coordinates, spherical in three dimensions, give instead two angles, azimuth and elevation, and a distance, all measured from a single origin.
Spatial audio almost always uses the second, and not as a matter of taste: the origin is the reference point, often the listening position. A position is then written the way hearing itself works: a direction the ears can resolve and a distance from that point, rather than as three offsets that would have to be recomputed every time the listener turns or moves. In that system, every source is described by three quantities.
- Azimuth is the horizontal angle: how far left or right the source sits, measured around you. Directly in front is
0°, to your right is+90°, behind is±180°, and to your left is−90°; the scale runs from−180°to+180°. Azimuth is the dimension stereo and surround systems handle best, because hearing is sharpest in the horizontal plane, and sharpest of all straight ahead, where a listener can resolve a change of direction of about 1°, several times finer than towards the sides (Mills 1958). - Elevation is the vertical angle: how far above or below ear level the source sits. The horizon is
0°, straight overhead is+90°, straight below is−90°. Adding convincing elevation is one of the defining features of modern immersive audio. - Distance is how far away the source is, measured in metres from the reference point.
Polar coordinates are therefore the natural way to say where a sound comes from.
Cartesian coordinates are another way of describing geometry in space. Their strength is a more absolute reference: the origin is usually fixed at the centre of the listening space rather than on one listener, so a position stays the same whoever is listening and wherever they stand. Loudspeakers, sources and listeners can then all be placed in one frame that belongs to the room.
The two descriptions are interchangeable: any position written in one can be converted into the other, as long as both are measured from the same origin. With the axes this guide uses (set out below), going from polar to Cartesian and back reads:
x = r cos φ sin θ, y = r cos φ cos θ, z = r sin φ
r = √(x² + y² + z²), θ = atan2(x, y), φ = arcsin(z / r)
What the formulas make explicit matters more than the formulas themselves: direction (the two angles) and distance (the radius) are independent quantities, and a spatial format has to carry all three. Most consumer systems reproduce azimuth well, elevation partially, and distance only by implication.
Whichever family is used, a set of coordinates only means something once its reference point is known: the origin from which every distance and angle is measured. By default it is either the centre of the room or the listening position, but the two do not give the same numbers, so a position that arrives from a file, a plugin or another tool always has to be checked for where it is measured from. The same goes for the direction of the axes and the sense in which angles are counted: tools disagree by a 90° rotation, a sign flip or both, and a position copied from one to another without checking is one of the most common bugs in spatial audio.
This guide sets its own conventions and keeps them throughout: the origin is the reference point, taken at the centre of the room unless a chapter says otherwise, x to the right, y forward, z up, with azimuth θ counted from straight ahead, negative to the left and positive to the right, over the range −180° to +180°; elevation φ runs from −90° straight below to +90° straight overhead, and distance is the ordinary one, in metres from the reference point. It is the frame our own tools speak, and the one most digital audio workstations and game engines use. In this frame, a source three metres away at θ = +40° and φ = +25° sits at x = 1.75, y = 2.08, z = 1.27 m: a little over two metres ahead, one and three quarters to the right, and a metre and a quarter up.
Several conventions coexist, and none of them is wrong: Ambisonics and IRCAM’s Spat, for instance, count angles the other way. What is wrong is using one without saying which. This guide states its convention once and keeps it; the table places it beside the two conventions you are most likely to meet elsewhere, so a position copied between them can be converted rather than guessed.
| this guide · our tools · most DAWs | Ambisonics (AmbiX, ACN/SN3D) · IRCAM Spat | ITU-R BS.2051 speaker labels | |
|---|---|---|---|
| x | right, left is negative | forward | no Cartesian axes |
| y | forward, behind is negative | left | no Cartesian axes |
| z | up | up | no Cartesian axes |
| azimuth 0° | straight ahead | straight ahead | straight ahead |
| positive azimuth | to the right | to the left | to the left |
| usual range | −180° … +180° | −180° … +180°, or 0° … 360° | M+030 means 30° to the left |
| elevation | −90° … +90°, up positive | −90° … +90°, up positive | U, M and B layers, plus a signed angle |
Going from this guide to a B-format file is therefore one axis swap and one sign flip, to be done deliberately rather than assumed: X = y, Y = −x, Z = z, and the azimuth changes sign. FuMa, the older B-format, differs again, but on channel order and normalisation rather than on geometry.
One more thing has to be settled before Part II. Every encoding in this chapter aims at the same target, two eardrums. A panning law works because of how the ear weighs two arrival times against two levels; a matrix survives because of what the ear forgives; binaural exists because the outer ear colours sound according to where it comes from. None of it can be judged, or designed, without knowing what those two ears actually measure. That is the subject of the next chapter, Spatial Psychoacoustics, and everything after it rests on what that chapter establishes.
References
Books
- Rumsey, F. (2001). Spatial Audio. Focal Press.
- Roginska, A., & Geluso, P. (Eds.). (2017). Immersive Sound: The Art and Science of Binaural and Multi-Channel Audio. Routledge.
- Holman, T. (2008). Surround Sound: Up and Running (2nd ed.). Focal Press.
- Zotter, F., & Frank, M. (2019). Ambisonics: A Practical 3D Audio Theory. Springer. Open access.
- Toole, F. (2017). Sound Reproduction: The Acoustics and Psychoacoustics of Loudspeakers and Rooms (3rd ed.). Routledge.
Perception
- Blauert, J. (1997). Spatial Hearing: The Psychophysics of Human Sound Localization (rev. ed.). MIT Press.
- Rayleigh, Lord (1907). “On our perception of sound direction.” Philosophical Magazine, 13(74), 214–232.
- Mills, A. W. (1958). “On the minimum audible angle.” JASA, 30(4), 237–246.
- Wallach, H. (1940). “The role of head movements and vestibular and visual cues in sound localization.” Journal of Experimental Psychology, 27, 339–368.
- Cherry, E. C. (1953). “Some experiments on the recognition of speech, with one and with two ears.” JASA, 25(5), 975–979.
- Bronkhorst, A. W. (2000). “The cocktail party phenomenon: A review.” Acta Acustica, 86, 117–128.
- Zahorik, P., Brungart, D., & Bronkhorst, A. (2005). “Auditory distance perception in humans: A summary of past and present research.” Acta Acustica, 91, 409–420.
- Møller, H., et al. (1995). “Head-related transfer functions of human subjects.” JAES, 43(5), 300–321.
- Wightman, F., & Kistler, D. (1989). “Headphone simulation of free-field listening.” JASA, 85(2), 858–878.
Stereo
- Blumlein, A. D. (1931). “Improvements in and relating to sound-transmission, sound-recording and sound-reproducing systems.” British Patent 394,325.
- Leakey, D. M. (1959). “Some measurements on the effects of interchannel intensity and time differences in two channel sound systems.” JASA, 31(7), 977–986.
- Bauer, B. B. (1961). “Phasor analysis of some stereophonic phenomena.” JASA, 33(11), 1536–1539.
History
- Scientific American (1881). “The telephone at the Paris Opera.” Contemporary account of Ader’s binaural demonstration.
- Garity, W. E., & Hawkins, J. N. A. (1941). “Fantasound.” Journal of the SMPTE, 37(8), 127–146.
- Gerzon, M. A. (1973). “Periphony: With-height sound reproduction.” JAES, 21(1), 2–10.
- Fellgett, P. (1975). “Ambisonics. Part one: General system description.” Studio Sound, 17(8).
- Allen, I. (1975). “The production of wide-range, low-distortion optical soundtracks utilising the Dolby noise reduction system.” Journal of the SMPTE, 84.
- Todd, C., et al. (1994). “AC-3: Flexible perceptual coding for audio transmission and storage.” AES 96th Convention.
- Herre, J., et al. (2015). “MPEG-H 3D Audio — The new standard for coding of immersive spatial audio.” IEEE JSTSP, 9(5), 770–779.
Standards
- ITU-R BS.2076-3 (2025). Audio Definition Model. The open production metadata model behind every immersive format.
- ITU-R BS.2127 (2019). Audio Definition Model renderer for advanced sound systems.
- ITU-R BS.775. Multichannel stereophonic sound system with and without accompanying picture. The reference stereo and 5.1 loudspeaker layouts.
- ITU-R BS.2051-3. Advanced sound system for programme production. The reference loudspeaker layouts.
- ITU-R BS.2088. Long-form file format for the international exchange of audio programme materials (BW64).
- EBU Tech 3364. Audio Definition Model — metadata specification.
- SMPTE ST 2098-2. Immersive Audio Bitstream specification.
- ATSC A/52. Digital Audio Compression (AC-3, E-AC-3) Standard.