Glossary
- The terms used throughout the guide, each with a short, self-contained definition.
- Grouped by theme, from perception to units; a formula is given where it makes an idea exact.
- Positions follow the guide’s convention: reference point at the centre of the room, x to the right, y forward, z up, azimuth positive to the right.
This glossary collects the vocabulary of the guide. Each entry stands on its own and can be read out of order. The chapters treat most of these terms in depth: Spatial Psychoacoustics for perception, Multichannel & Immersive Formats for representations, and the parts on techniques, recording, the sound field and the room, and systems for the rest.
A–Z index
0–9 · 3-to-1 rule
A · ACN · ADM · AES67 · A-format · Air absorption · Ambisonic Channel Number · Ambisonic microphone · Ambisonic order · Ambisonics · AmbiX · Apparent source width · ASW · Audio Definition Model · Audio Video Bridging · AVB · Azimuth
B · Bass management · Bed · B-format · Binaural · Binaural Room Impulse Response · Blumlein pair · BRIR · Broadcast Wave 64 · BW64
C · C80 and C50 · Calibration · Cartesian coordinates · Channel-based · Cocktail party effect · Coincident pair · Comb filtering · Cone of confusion · Convolution reverb · Coordinate convention · Critical distance · Crossover · Crosstalk · Crosstalk cancellation
D · DBAP · Decibel · Decoder · Decorrelation · DI · Diffuse field · Directivity · Directivity index · Direct sound · Direct-to-reverberant ratio · Discrete · Distance · Distance-Based Amplitude Panning · Distance perception · Doppler effect · Downmix · DRR · Dummy head · Duplex theory
E · Early Decay Time · Early reflections · Echo threshold · EDT · Elevation · Encoder · Externalisation
F · Fast Fourier transform · FDN · Feedback delay network · FFT · FOA and HOA · Focused source · FOH · Front of house · FuMa
G · Generic and individual HRTFs · Group delay
H · Head-Related Impulse Response · Head-Related Transfer Function · Head tracking · Height layer · HRIR · HRTF · Huygens–Fresnel principle
I · IACC · ICTD and ICLD · ILD · Immersive audio · Immersive delivery formats · Impulse response · In-head localisation · Interaural Cross-Correlation Coefficient · Interaural Level Difference · Interaural Time Difference · Inter-channel coherence · Inverse-distance law · IR · ITD
L · Latency · LEV · LFE · Listener envelopment · Localisation · Low-Frequency Effects · LUFS and LKFS
M · MAA · Masking · Matrix encoding · MDAP · Measurement signals · Mid/side · Minimum audible angle · Minimum phase · Mono · M/S · Multiple-Direction Amplitude Panning
N · Near-coincident pair · Near field and far field
O · Object-based · Open Sound Control · OSC
P · Panning · Panning law · Pan pot · Partitioned convolution · Phantom image · Pinna · Polarity · Polar pattern · Polar (spherical) coordinates · Precedence effect · Precision Time Protocol · PTP
R · Reference point · Renderer · Room modes · RT60
S · Scene-based · Schmidt semi-normalisation · Schroeder frequency · SN3D · SOFA · Sound field · Sound pressure level · Spaced array · Spatial aliasing · Spatial averaging · Spatially Oriented Format for Acoustics · Speed of sound · Spherical harmonics · SPL · Spot microphone · Standard layouts · Stereo · Subwoofer · Summing localisation · Surround · Sweet spot
T · Tangent law · Time alignment · Timecode · Transaural · True peak
V · VBAP · Vector Base Amplitude Panning · Virtual loudspeaker
W · Wave field synthesis · Wavelength · WFS
X · X-curve · XTC · XY · X.Y.Z notation
Perception
Localisation. The perception of where a sound comes from: its direction and its distance. The auditory system derives it from the differences between the two ears, the filtering of the outer ear, and the changes produced by head movements.
ITD (Interaural Time Difference). The difference in arrival time of a sound at the two ears. It is the dominant cue for left–right position below about 1.5 kHz, and reaches roughly 0.65 ms for a source directly to one side. For a spherical head, Woodworth’s approximation gives , with head radius , speed of sound and azimuth in radians, from 0 in front to at the side.
ILD (Interaural Level Difference). The difference in level between the two ears caused by the shadow of the head. It becomes the dominant left–right cue above about 1.5 kHz, where the head is large compared with the wavelength, and can reach 20 dB or more at high frequencies.
Duplex theory. Lord Rayleigh’s account (1907) of left–right localisation: time differences serve at low frequencies, level differences at high frequencies, with a transition around 1–1.5 kHz where neither cue is fully reliable.
Pinna. The visible outer ear. Its folds add direction-dependent reflections and spectral notches, which are the main cues for elevation and for telling front from back.
HRTF (Head-Related Transfer Function). The filter describing how the head, torso and outer ear transform a sound arriving from a given direction before it reaches the eardrum, one filter per ear. It carries the ITD, the ILD and the spectral cues of the pinna in a single direction-dependent response.
HRIR (Head-Related Impulse Response). The time-domain form of the HRTF. Convolving a mono signal with the left and right HRIRs for a direction places that sound in that direction over headphones.
Generic and individual HRTFs. A generic HRTF set is measured on a dummy head or averaged across listeners; an individual set is measured on, or fitted to, one listener. Generic sets work reasonably for many people, but produce more front/back and up/down confusions and larger elevation errors; their effect on externalisation is smaller and less consistent.
BRIR (Binaural Room Impulse Response). An HRIR measured in a room rather than in free field, so that it also carries the room’s reflections and reverberation. Rendering with BRIRs helps externalisation.
SOFA (Spatially Oriented Format for Acoustics). The AES69 standard file format for HRTF sets and other direction-dependent acoustic data, so that measured datasets can be exchanged between tools.
Cone of confusion. The set of directions that produce nearly the same ITD and ILD, roughly a cone around the axis through the two ears. Within it, front/back and up/down confusions arise unless spectral cues or head movements resolve them.
Head tracking. Measuring the orientation, and sometimes the position, of the listener’s head so that a binaural rendering can counter-rotate the scene. Sources then stay fixed in the world as the head turns, which resolves front/back confusions and helps externalisation.
Externalisation. The perception that a sound lies outside the head, in the surrounding space. Over headphones it is helped mainly by some room response, and further by plausible HRTF cues and head tracking.
In-head localisation. The opposite of externalisation: the image sits inside the head. It is the usual result of ordinary stereo, and often of generic binaural, over headphones.
Precedence effect. The dominance of the first-arriving wavefront. When a sound is followed by a delayed copy, from about 1 ms up to a few milliseconds for clicks and a few tens of milliseconds for speech and music, the two fuse into one event heard in the direction of the first. Also called the law of the first wavefront; the Haas effect names Haas’s 1951 study of the same phenomenon with speech.
Echo threshold. The delay beyond which a delayed copy stops fusing with the original and is heard as a separate echo: a few milliseconds for clicks, several tens of milliseconds for speech and music.
Summing localisation. When coherent signals from several loudspeakers reach the listener within about 1 ms of one another, the auditory system fuses them into a single phantom image whose position depends on their relative level and timing.
Phantom image. A sound heard between loudspeakers, where no loudspeaker stands, produced when they play related signals. It is the basis of stereo and of amplitude panning.
Minimum audible angle (MAA). The smallest change of direction a listener can detect: about 1° straight ahead, several times coarser towards the sides (Mills 1958), and about 3 to 4° in the vertical plane in front (Perrott and Saberi 1990).
Masking. The reduction in audibility of one sound caused by another. It is strongest when the two share a frequency range and a direction; separating them in space reduces it, an effect called spatial release from masking.
Cocktail party effect. The ability to follow one talker among many, named by Cherry in 1953. Direction is one of the cues involved, together with differences in pitch, in timing and in familiarity with the voice.
Distance perception. Hearing how far away a source is. With no dedicated cue of its own, it is inferred from level, from the direct-to-reverberant ratio, from the loss of high frequencies over long paths and, very close to the head, from large interaural level differences.
Apparent source width (ASW). The perceived broadening of a source beyond its physical size, increased mainly by early reflections arriving from the sides. With listener envelopment, it makes up the impression of spaciousness.
Listener envelopment (LEV). The sense of being surrounded by sound, driven mainly by late reverberant energy, roughly after 80 ms, arriving from the sides and from many directions.
IACC (Interaural Cross-Correlation Coefficient). The maximum of the normalised cross-correlation between the signals at the two ears, taken over lags of up to ±1 ms: . Low values go with a wider, more enveloping impression.
Geometry and coordinates
Reference point. The origin from which positions are measured. In this guide it is the centre of the room unless a chapter says otherwise; many tools use the listening position instead. Coordinates mean nothing until their reference point is known.
Coordinate convention. The choices that give coordinates a meaning: the origin, the direction of the axes and the sense in which angles are counted. Tools differ by a 90° rotation, a sign flip or both, and a position copied between them without conversion is one of the most common errors in spatial audio.
Cartesian coordinates. A position given as three distances along fixed axes. In this guide x points to the right, y forward and z up, in metres from the reference point.
Polar (spherical) coordinates. A position given as two angles, azimuth and elevation, and a distance, all measured from the reference point. With the guide’s axes, , and .
Azimuth. The horizontal angle of a source. In this guide 0° is straight ahead, positive to the right and negative to the left, over −180° to +180°. Ambisonics and ITU-R BS.2051 count the other way, positive to the left.
Elevation. The vertical angle of a source above or below the horizontal plane through the reference point, usually set at ear height: −90° straight below, 0° on the horizon, +90° straight overhead.
Distance. How far a source is from the reference point, in metres.
Reproduction concepts
Mono. Reproduction from a single channel. Everything is heard from the same point, the loudspeaker.
Stereo (2.0). Two-channel reproduction, classically over two loudspeakers at ±30° forming an equilateral triangle with the listener. Level and time differences between the channels place phantom images along the line between the speakers.
Surround. Channel-based reproduction with loudspeakers beside and behind the listener, such as 5.1 and 7.1, in layouts standardised by ITU-R BS.775 and BS.2051.
Immersive audio. Reproduction that adds height to surround, with loudspeakers above and sometimes below the listener, or a binaural rendering that includes elevation. Also called 3D audio.
Height layer. The ring of loudspeakers above ear level in an immersive layout: the “.4” in 7.1.4.
Sweet spot. The listening position, or small area, where a loudspeaker system’s intended image, balance and timing hold. The image degrades as the listener moves away from it.
Crosstalk. Over loudspeakers, the path from each loudspeaker to the ear on the far side: left loudspeaker to right ear, right loudspeaker to left. Stereo is designed around it; binaural signals played over loudspeakers are corrupted by it.
Crosstalk cancellation (XTC). Filtering of the loudspeaker feeds so that the crosstalk paths cancel at the ears, letting each ear receive mainly its own signal. It is the basis of transaural reproduction, and holds only in a small area.
ICTD and ICLD (inter-channel time and level differences). The time and level differences between loudspeaker signals, used by panning and by microphone techniques to place a phantom image. Distinct from ITD and ILD, which are measured at the ears.
Inter-channel coherence. The normalised correlation between two reproduction channels. High coherence gives a narrow, stable phantom image; low coherence gives width and diffuseness. Distinct from IACC, which is measured at the ears.
Decorrelation. Deliberately reducing the similarity between channels, with delays, all-pass filters or phase randomisation, to widen an image, increase envelopment, or reduce comb filtering between loudspeakers.
Comb filtering. The regularly spaced peaks and notches produced when a signal is added to a delayed copy of itself, as when the same sound reaches a listener from two loudspeakers at different distances. For a delay , the first notch falls at .
Virtual loudspeaker. A loudspeaker that exists only in the rendering: a pair of binaural filters that makes a speaker appear at a position over headphones, or an intermediate position to which a decoder renders before mapping onto the real loudspeakers.
Encoder. The stage that turns a scene, whether sources with positions or a captured field, into an intermediate representation such as Ambisonic components or objects with metadata. The counterpart of the decoder.
Decoder. The stage that turns an encoded representation, Ambisonic components or a matrix-encoded signal, into feeds for a concrete loudspeaker layout or for headphones.
Renderer. The software or hardware that turns a scene description, objects with metadata or an Ambisonic signal, into loudspeaker or headphone feeds for the layout actually present.
Upmix. Deriving more output channels than the source contains, for example an immersive presentation from a stereo recording, by extracting ambience and steering direct sound.
Downmix. Combining a signal into fewer channels, for example 5.1 to stereo, with defined coefficients: typically the centre and surrounds are added to left and right at about −3 dB, and the LFE is discarded.
Representations and formats
Channel-based. A representation with one signal per loudspeaker, each tied to a position fixed by a standard layout. Stereo, 5.1 and 7.1.4 are channel-based; the signal is read back correctly only if the loudspeakers stand where the standard puts them.
Object-based. A representation in which sources are stored as audio plus metadata describing their position and behaviour, and rendered at playback to the loudspeakers actually present.
Scene-based. A representation of the whole sound field at a point, independent of sources and loudspeakers, as a set of components such as Ambisonics’ spherical harmonics.
Discrete. Channels carried independently, each with its own signal, rather than folded together by a matrix.
Matrix encoding. Folding extra channels into fewer through amplitude and phase relationships, to be unfolded by a matched decoder, as in the 4:2:4 systems of Dolby Stereo.
Bed. A channel-based layer in an object-based mix, for example a 5.1 or 7.1.2 bed (the largest bed Dolby Atmos allows), carrying ambience and content that does not need to move, over which objects are placed.
LFE (Low-Frequency Effects). A dedicated band-limited channel, the “.1” in 5.1, reproduced by a subwoofer. It is typically limited to about 120 Hz and played 10 dB above the main channels, and it is an effects channel added to the mix, not the place where their bass is sent.
Bass management. Filtering the low frequencies out of the main channels and sending them, summed with the LFE, to one or more subwoofers, so that full-range content plays correctly on loudspeakers of limited extension. The usual crossover frequency is 80 Hz.
X.Y.Z notation. The shorthand for a loudspeaker layout: X full-range channels at ear level, Y LFE channels, Z height channels. 7.1.4 means seven, one and four.
Standard layouts. The loudspeaker positions fixed by ITU-R BS.775 for stereo and 5.1, and by ITU-R BS.2051 for immersive layouts. BS.2051 names each loudspeaker by its layer, U for upper, M for middle, B for bottom, T for top, and its azimuth: M+030 is on the middle layer, 30° to the left.
ADM (Audio Definition Model). An open metadata model, ITU-R BS.2076, describing the content, format and position of the channels, objects and scene-based signals in a programme, so that one production can be exchanged and exported to any delivery format.
BW64 (Broadcast Wave 64). The WAV-based file format of ITU-R BS.2088, which allows files larger than 4 GB and embeds ADM metadata. A file carrying ADM this way is often called an ADM BWF.
Immersive delivery formats. Dolby Atmos, DTS:X, MPEG-H 3D Audio and Auro-3D: competing systems for delivering immersive audio, each with its own codec and licensing. Atmos, DTS:X and MPEG-H carry objects; Auro-3D began as a family of channel-based layouts built in layers (ear level, height and, from Auro 10.1, a top loudspeaker), and added objects later with AuroMax (2015).
Spherical harmonics. The family of functions on the sphere that Ambisonics uses to describe how a sound field varies with direction. Up to order there are of them.
Ambisonic order. The highest order of spherical harmonics kept in an Ambisonic signal. A higher order gives finer spatial resolution at the cost of more channels: for a full sphere.
FOA and HOA. First-order Ambisonics, four channels, and higher-order Ambisonics, order 2 and above.
B-format. The first-order Ambisonic signals: W, an omnidirectional pressure signal, and X, Y and Z, three figure-of-eight signals along the axes. By extension, any Ambisonic signal set in its spherical-harmonic form.
A-format. The raw signals of the capsules of an Ambisonic microphone, classically four capsules on a tetrahedron, before conversion to B-format.
ACN (Ambisonic Channel Number). The standard ordering of Ambisonic channels: the component of order and index , with , takes the number .
SN3D (Schmidt semi-normalisation). The normalisation of Ambisonic channels used with ACN in the AmbiX format. N3D, full normalisation, is the other common convention; the two differ in the relative gain of each order.
AmbiX. The de facto interchange convention for Ambisonics, proposed by Nachbar, Zotter, Deleflie and Sontacchi in 2011: ACN channel order with SN3D normalisation.
FuMa (Furse–Malham). The older Ambisonic convention, with its own channel order and a W channel attenuated by 3 dB. It is still found in first-order material and must be converted before use in AmbiX tools.
UHJ. A stereo-compatible matrix encoding of Ambisonics into two or more channels, devised to distribute Ambisonic recordings on conventional media.
Spatialisation techniques
Panning. Distributing one source’s signal across two or more loudspeakers to place a phantom image between them, most often by adjusting their relative gains.
Pan pot. The panoramic potentiometer: a control that splits one signal between several outputs from a single gesture. It was first built for Disney’s Fantasound in 1940 and is still found on every channel of a mixing console.
Panning law. The rule relating a pan position to the gains of each loudspeaker, chosen so that loudness stays constant across the pan. The centre is typically attenuated by 3 dB (constant power, , for a pan parameter from 0 to , so that ), or by 4.5 or 6 dB.
Tangent law. The relation, derived from a low-frequency model of summing localisation and valid mainly below about 700 Hz, between the gains of two loudspeakers at ±θ₀ and the direction θ of the phantom image: , with θ positive to the right.
VBAP (Vector Base Amplitude Panning). Pulkki’s method (1997) that places a source by computing gains for the two or three loudspeakers surrounding its direction, writing the source direction as a non-negative combination of their direction vectors.
MDAP (Multiple-Direction Amplitude Panning). An extension of VBAP (Pulkki 1999) that spreads a source over several nearby directions, so that its width and loudness stay more even as it moves across and between loudspeakers.
DBAP (Distance-Based Amplitude Panning). A method (Lossius and colleagues, 2009) that sets each loudspeaker’s gain from its distance to the virtual source, with no assumption of a sweet spot or of a regular layout.
Ambisonics. A scene-based technique, proposed by Gerzon and colleagues in 1973, that encodes a sound field as spherical-harmonic components and decodes it to any loudspeaker layout or to binaural, separating production from playback.
Binaural. Reproduction that aims to recreate the signals at the listener’s two ears, usually over headphones through HRTF filtering, so that sounds can be heard in any direction, including above and behind.
Transaural. Delivering binaural signals over loudspeakers, using crosstalk cancellation so that each ear receives mainly its intended signal.
Wave field synthesis (WFS). A technique that uses a dense array of loudspeakers to reconstruct a target wavefront over an extended area, so that a virtual source keeps a consistent position across the audience, within the limits set by the array’s size and spacing.
Huygens–Fresnel principle. The principle that a wavefront can be reconstructed by treating every point on a surface as a secondary source. Formalised by the Kirchhoff–Helmholtz integral, it is the physical foundation of wave field synthesis, which in practice uses the simpler Rayleigh integral with monopole loudspeakers.
Spatial aliasing. In loudspeaker and microphone arrays, the errors that appear once the element spacing exceeds half a wavelength. Above the aliasing frequency, roughly in the worst case, the reconstruction becomes increasingly inaccurate, so the spacing limits the array’s accurate bandwidth.
Focused source. In wave field synthesis, a virtual source rendered in front of the loudspeaker array, between the array and the listener, by a wavefront that converges on the intended point before spreading out again.
Recording and capture
Polar pattern. The directional sensitivity of a microphone. Omnidirectional picks up equally from all sides; cardioid rejects sound from the rear, with its null at 180°; supercardioid and hypercardioid are narrower, with nulls near 126° and 110°; figure-of-eight picks up front and back in opposite polarity, with nulls at 90°.
Coincident pair. Two directional microphones with their capsules at the same point, angled apart. They differ essentially in level only, with negligible time differences. XY and the Blumlein pair are coincident.
XY. A coincident pair of cardioid microphones, typically angled 90° to 135° apart.
Blumlein pair. Two figure-of-eight microphones crossed at 90°, coincident, as described in principle in Alan Blumlein’s 1931 patent.
Mid/side (M/S). A coincident technique that pairs a forward-facing mid microphone, often cardioid, with a sideways figure-of-eight. Left and right are obtained as M + S and M − S, so the stereo width can be adjusted after recording by changing the level of S.
Near-coincident pair. Two directional microphones a short distance apart, combining level and time differences. ORTF, two cardioids 17 cm apart and angled 110°, is the best known.
Spaced array. Microphones several tens of centimetres to several metres apart, usually omnidirectional, whose signals differ mainly in arrival time. The Decca tree, three omnidirectional microphones in a T, is a spaced main array; the Hamasaki square, four sideways figure-of-eights about 2 m apart, captures ambience for surround.
Spot microphone. A microphone placed close to one instrument or section to add presence or detail to the main pickup.
3-to-1 rule. A rule of thumb for close microphones: keep two microphones at least three times as far from each other as each is from its source, so that the leakage into the other is about 9 dB lower and comb filtering stays low.
Ambisonic microphone. A microphone with several capsules, classically four on a tetrahedron, whose A-format output is converted to B-format. Higher-order models use many capsules on a sphere.
Dummy head. A model of a human head, often with torso and pinnae, with microphones at the ear canals, used for binaural recording and HRTF measurement. KEMAR is a widely used standard manikin.
The room and the field
Sound field. The distribution of acoustic pressure and particle velocity in a region over time. Spatial audio aims to recreate or evoke a chosen sound field at the listener.
Direct sound. The sound arriving first, along the straight path from source to listener. It fixes the perceived direction of the source.
Early reflections. The first reflections from the room’s surfaces, arriving roughly within 50–80 ms of the direct sound. They shape perceived width, timbre and the impression of the room’s size.
Diffuse field. A field in which energy arrives equally from all directions with random phase. The late reverberant tail of a room approaches it.
Near field and far field. Close to a source, pressure and particle velocity are not in phase, tending towards quadrature very close to it (the acoustic near field, within about a wavelength), and the source’s size makes level vary irregularly with distance (the geometric near field). Beyond both lies the far field, where pressure falls as .
Direct-to-reverberant ratio (DRR). The ratio of direct to reverberant energy at the listener, in dB. It falls with distance and is one of the main distance cues.
Critical distance. The distance from a source at which direct and reverberant energy are equal: for an omnidirectional source, with the room volume in m³, multiplied by for a source of directivity factor . Beyond it the reverberant field dominates.
RT60. The reverberation time: how long the sound energy takes to decay by 60 dB after a source stops. Sabine’s estimate is , with volume in m³ and total absorption in m². In practice it is measured over part of the decay, from −5 dB to −25 dB (T20) or to −35 dB (T30), and extrapolated to 60 dB.
EDT (Early Decay Time). The reverberation time derived from the first 10 dB of the decay, scaled to 60 dB. It matches perceived reverberance better than RT60.
C80 and C50 (clarity). The ratio, in dB, of early to late energy in an impulse response, with the boundary at 80 ms for music (C80) or 50 ms for speech (C50): .
Room modes. The resonances of a room, at frequencies set by whole numbers of half-wavelengths along its dimensions, singly (axial modes) or in combination (tangential and oblique modes). They dominate the low-frequency response, below the Schroeder frequency.
Schroeder frequency. The frequency above which room modes overlap densely enough to be treated statistically: , in Hz, with in m³. Below it the room behaves as a set of individual modes.
Impulse response (IR). The response of a system, whether a room, a loudspeaker or a filter, to an ideal impulse. It contains the system’s whole linear behaviour, and convolving a signal with it reproduces that behaviour.
Convolution reverb. Reverberation produced by convolving a signal with a measured impulse response, reproducing the reflections and decay of the captured space.
Partitioned convolution. Convolution with a long impulse response split into blocks and processed in the frequency domain, to keep latency low. It is the usual way to run convolution reverb in real time.
Feedback delay network (FDN). An efficient reverberation algorithm made of several delay lines connected through a feedback matrix, producing a dense, controllable late tail.
Doppler effect. The change in perceived frequency caused by motion between source and listener: for a source approaching a stationary listener at speed .
Air absorption. The attenuation of sound by the air itself. It rises steeply with frequency, grows in proportion to distance, and depends on humidity and temperature (ISO 9613-1); it dulls distant sources and acts as a distance cue.
Systems and measurement
Calibration. Adjusting a playback system so that every loudspeaker reaches the listening area at the intended level, time and frequency response. Cinema, for example, sets each screen channel to 85 dB SPL, C-weighted, for wideband pink noise at −20 dBFS RMS, measured with slow response and spatially averaged.
Time alignment. Delaying the feeds of nearer loudspeakers so that sound from all of them arrives together at a reference position.
Latency. The delay between a signal entering a system and leaving it. In a spatial system every path must have matched latency, or time alignment is lost.
Subwoofer. A loudspeaker dedicated to the lowest frequencies, fed by the LFE channel and by bass management.
Crossover. A filter that splits the spectrum between loudspeakers, or between the drivers of one loudspeaker. A fourth-order Linkwitz–Riley crossover (LR24) sums to a flat amplitude response.
Polarity. The sign of a signal. Inverting it is equivalent to a 180° phase shift at every frequency, whereas the phase shift of a delay or a filter generally varies with frequency.
Group delay. The delay a system imposes on the envelope of a narrow band of frequencies: the negative derivative of its phase response with respect to angular frequency.
Minimum phase. A system whose phase is the least possible for its magnitude response, and which can therefore be inverted by a stable filter. Room responses are generally not minimum-phase, which limits what equalisation can correct.
Spatial averaging. Measuring at several microphone positions and averaging the results, so that equalisation corrects the room over an area rather than at a single point.
X-curve. The target frequency response for cinema playback (SMPTE ST 202 and ISO 2969): flat to about 2 kHz, then falling by about 3 dB per octave. It is defined for a medium-sized theatre and is modified for larger and smaller rooms.
Measurement signals. Signals used to measure impulse responses. A maximum-length sequence (MLS) is a pseudo-random binary sequence; an exponential sine sweep (Farina 2000) rises through the spectrum and separates harmonic distortion from the linear response.
FFT (fast Fourier transform). An efficient algorithm that computes the spectrum of a block of samples, at the heart of measurement, analysis and frequency-domain processing.
OSC (Open Sound Control). A network protocol for exchanging control messages, such as source positions, between applications and devices.
Timecode. A time reference used to synchronise systems: LTC is carried as an audio signal, MTC over MIDI.
AES67. A standard for exchanging audio over IP networks between systems from different makers, built on RTP streams and PTP clock synchronisation.
PTP (Precision Time Protocol). IEEE 1588: a protocol that synchronises clocks across a network to well under a microsecond. It is the timing basis of AES67, and AVB uses its profile gPTP (IEEE 802.1AS).
AVB (Audio Video Bridging). A set of IEEE 802.1 standards for time-synchronised, low-latency audio and video over Ethernet with reserved bandwidth.
FOH (front of house). The mixing position in the audience area, and by extension the main system that serves the audience.
Units and metrics
Decibel (dB). A logarithmic unit of ratio: for pressures and for powers or energies.
Sound pressure level (SPL). The level of a sound pressure relative to 20 µPa, in dB. It measures the physical level, not loudness, which also depends on frequency and duration.
LUFS and LKFS. Loudness units relative to full scale, measured according to ITU-R BS.1770 with K-weighting and gating; the two names denote the same unit, and 1 LU equals 1 dB. EBU R128 sets a broadcast target of −23 LUFS.
True peak. The peak level of the reconstructed analogue waveform, estimated by oversampling as specified in ITU-R BS.1770 and expressed in dBTP. It can exceed the highest sample value.
Inverse-distance law. In the far field of a point source in free space, pressure falls as , so the level drops by 6 dB for each doubling of distance.
Directivity. The variation of a source’s radiated level, or a microphone’s sensitivity, with direction, described by a polar pattern. It sets how much direct sound, compared with reflected sound, reaches a listener.
Directivity index (DI). How much a source concentrates its output on its axis compared with an omnidirectional source of the same power: dB, where is the directivity factor.
Wavelength. The spatial period of a sound wave, . At 1.5 kHz it is about 0.23 m, close to the longest path difference between the ears, which is why the duplex transition sits there.
Speed of sound. About 343 m/s in air at 20 °C, rising by about 0.6 m/s per degree. It sets the time scale of ITDs, reflections and array spacing.