Aller au contenu principal

Case Study — Immersive Music Production & Post

In brief
  • The hub for the music cases: the shared studio brief and the vocabulary of spatial creation.
  • Six axes, each with a technical handle: position, depth, size, trajectory, envelopment, density.
  • Cross-cutting realities: fold-down integrity, the object budget, loudness, remix-vs-re-render.
  • Genre rewrites the brief — captured versus constructed space — so each genre gets its own page.

The brief

A label wants an artist's new album released in immersive audio for the streaming platforms (Apple Music, Tidal), plus older singles from the back-catalogue reworked to match. Deliverables: a Dolby Atmos master (ADM BWF), a binaural render for headphones, and a stereo fold-down that must not sound worse than the original stereo. The mix happens in a small, treated studio on a 7.1.4 monitor layout. Budget is a few weeks, not a film schedule.

This is the studio end of the spectrum: no audience in the room, but every downstream listening situation — a phone, a soundbar, a car, a pair of earbuds — is out of your control. And it is not one job: the aesthetic contract differs so sharply by genre that each gets its own page. This hub covers what they share — the creative vocabulary and the delivery constraints — and hands off to the genre chapters for the rest.

The vocabulary of spatial creation — and its technical handles

Before any genre, there is a vocabulary. Working spatially with music means composing along a small set of perceptual axes, and each one has a concrete technical handle in this guide. Learning the pairing — the creative idea and the mechanism that realises it — is what stops immersive mixing from becoming knob-twiddling. This table is the spine of every genre page that follows; the artistic vocabulary develops the intention side, the linked chapters the mechanism.

Position — where a sound sits

The most basic decision: a location in the horizontal and (now) the vertical. Creatively it is figure-placement — the lead is here, the guitar there. Technically it is a panning law: amplitude panning (VBAP/DBAP) distributes a source's energy across the nearest speakers, and the ear fuses the result into a phantom image via inter-channel level and time differences (ICLD/ICTD). The creative caveat — a phantom image is only stable near the sweet spot — is a fact about the mechanism, and it is why placement that must survive off-centre (or on a phone) tends back toward the real speakers.

Distance and depth — how far away, how present

Depth is the axis amateurs ignore and professionals live in. A sound is pushed back not by lowering a fader alone but by three linked cues: falling level (roughly inverse-distance), a high-frequency roll-off from air absorption, and — decisively — a falling direct-to-reverberant ratio, more reverberant energy relative to direct (direct, diffuse & envelopment). Change these together and a source recedes convincingly; change level alone and it just gets quieter. Physically correct near-field depth (a source that feels in front of the speakers) additionally needs wavefront curvature, the domain of WFS.

Size and width — a point, or a body

A source is not always a dot. "How big is this sound" — a pinpoint hi-hat versus a pad that fills the room — is an expressive parameter, and its technical name is decorrelation. Feed near-identical signals to several speakers and the ear hears one narrow image; decorrelate them (independent phase/time relationships) and the image widens into an extended body. The perceptual measure of that width, Apparent Source Width, is governed by the interaural cross-correlation (IACC) of the sound reaching the two ears — low correlation, wide source. Practically this shows up as an object spread/size parameter, or RIPL's continuous Spread/Focus axis. The pitfall is also a fact about the mechanism: decorrelation that drifts toward anti-correlation collapses on mono fold-down.

Trajectory and motion — a path through time

Movement is a phrase, not a garnish (see gesture and trajectory). Technically a trajectory is position automation over time, and three mechanisms decide whether it sounds good: the interpolation of the panning gains and delays between keyframes (too coarse and you hear zipper noise; the renderer's interpolation order is a real quality lever); Doppler shift if the motion is physically modelled through a delay line; and the precedence effect, which for a fast move means the ear locks onto onsets, so rapid trajectories read as direction more than as precise position. Perception also sets a floor — below a minimum audible movement, automation satisfies the lane but not the ear.

Envelopment and the room — being inside a space

Distinct from placing sources is manufacturing the space they sit in. Envelopment — the sense of being surrounded and inside — is built from diffuse, decorrelated late energy, especially arriving laterally and from height, i.e. late reverberation spread around and above (direct, diffuse & envelopment). Crucially, room character is independent of source position: you can place a voice close and dry yet wrap it in a cathedral, which no natural space offers. The height layer, in most music, is this — air and hall — not a shelf for instruments.

Density — how much a listener can follow

The last axis is a budget. There is a hard perceptual limit on how many simultaneous, independently-placed spatial streams a listener can parse before they fuse into a mass (the density budget). Immersive tempts you to place everything; masking and this limit mean that past a handful of active spatial voices, more placement buys nothing. Deciding whether the listener should locate discrete sources or bathe in a field is a per-section choice.

The pairing, in one line

Position ↔ panning law · Depth ↔ direct/reverberant + air + level · Size ↔ decorrelation (IACC/ASW) · Trajectory ↔ automation + interpolation + Doppler + precedence · Envelopment ↔ diffuse late reverberation · Density ↔ the perceptual limit on simultaneous streams.

Genre rewrites the brief — a page each

"An immersive music mix" is really several jobs, separated by whether space is captured or constructed, and by where the voice may live. Each has its own chapter:

  • Classical & Orchestral — space is captured: fidelity to a real hall, perspective and envelopment from a microphone array, height as air. Objects barely used.
  • Jazz & Acoustic Small Ensemble — the ensemble in a plausible room: capture plus gentle placement, intimacy and depth over spectacle.
  • Pop, R&B & Hip-Hop — space is constructed: an anchored lead vocal, width by decorrelation, effects thrown into height, and the fold-down as the real master.
  • Rock & Band Music — the band as a believable stage: energy first, width and depth serving the song, live-capture realities.
  • Electronic & Experimental — space is the instrument: the full trajectory-and-size vocabulary, synthesis and spatialisation composed together.

Cross-cutting technical constraints

Whatever the genre, the same delivery realities bite:

  • The room is not the listener's room. A 7.1.4 studio is a reference; most listeners are on earbuds. Judge through the renders, and trust the binaural monitor path.
  • Fold-down integrity. Phantom-centre build-up, height collapsing into comb filtering, and the loudness relationship between immersive and stereo all surface only on the stereo and binaural fold-downs — check them, do not assume them.
  • The object budget and the bed. The Atmos music renderer works to a fixed budget of audio elements; deciding what is a discrete object (must localise or move) and what belongs in the bed (room, glue) is a real authoring choice.
  • Whose binaural renderer? The headphone result is not one thing — different platforms voice the same ADM differently. Check on the actual target chain.
  • Loudness and metadata. Streaming immersive normalises differently from stereo; reconcile the immersive and stereo levels, and treat ADM metadata as a deliverable (see Post-Production, Immersive Music Production).
  • Remix or re-render — and is it even deliverable? A true remix needs multitracks; with only a stereo master you are doing a principled re-render — a stereo decode, not channel invention. But mind the platform rules: Apple requires a Dolby Atmos music track to be built from multitracks or stems and prohibits upmixing a stereo mix as an Atmos deliverable. A stereo decode is the right tool where you control playback (live, installation) or as a mixing aid — not a shortcut around the stems requirement.
  • Capture decides your freedom. How the material was recorded sets how freely you can space it later: a captured field fixes the perspective, isolated sources let you place and move them. That upstream choice is a chapter of its own — see Capturing for Immersive.

Solutions — with RIPL

RIPL's relevance across every genre is its one unified source model: mono objects, native multichannel stems, and a decoded stereo field are all first-class sources in one scene, rendered by one engine to any output.

  • The vocabulary, as controls. The axes above map onto RIPL parameters directly: position and gain-and-delay placement, a continuous Spread/Focus for size, automation on the timeline for trajectory, and per-object interpolation that spends CPU on the moving foreground.
  • Captured or constructed, one tool. The classical record's captured perspective (a wide, distance-true stage plus a hall bed) and the pop record's constructed space (anchored vocal, placed and moving objects) are two uses of the same engine, not two products.
  • Catalogue reworks. Stereo masters enter through HSR, which separates localisable primaries from diffuse ambience before placement — so the vocal stays put and the ambience opens, with energy, coherence and timbre preserved.
  • One scene, every deliverable. Because authoring is separated from rendering, the Atmos feeds, the ambisonic archive, the binaural stream and the stereo fold-down are modes of one scene, not four mixes to keep in sync.

Pitfalls and checklist

Checklist
  • Did you pick the genre contract — captured fidelity vs constructed space — before placing anything?
  • Are you working the depth axis (level + air + direct/reverberant together), or only panning left–right and up?
  • Does any decorrelation for width survive mono fold-down, or does it collapse?
  • Is every trajectory above the audible-movement floor and free of zipper noise?
  • Is the density under what a listener can follow, or are you placing for placement's sake?
  • Did you verify binaural and stereo fold-downs on real devices and the platform's own renderer, and reconcile immersive vs stereo loudness?

Bibliography

  • Roginska, Agnieszka, and Paul Geluso, eds. Immersive Sound: The Art and Science of Binaural and Multi-Channel Audio. New York: Routledge, 2017.
  • Rumsey, Francis. Spatial Audio. Oxford: Focal Press, 2001.
  • Blauert, Jens. Spatial Hearing: The Psychophysics of Human Sound Localization. Rev. ed. Cambridge, MA: MIT Press, 1997.
  • Pulkki, Ville. "Virtual Sound Source Positioning Using Vector Base Amplitude Panning." Journal of the Audio Engineering Society 45, no. 6 (1997): 456–466.
  • Kendall, Gary S. "The Decorrelation of Audio Signals and Its Impact on Spatial Imagery." Computer Music Journal 19, no. 4 (1995): 71–87.

A note on sourcing: the reference list above gives well-established works; a few edition years are worth confirming against a primary source before print. Each genre page carries its own, more specific bibliography.

See also

In the technical guideImmersive Music Production and Post-Production (the full workflows); Object-based audio, Binaural, Ambisonics and Stereo is already spatial; Direct, diffuse & envelopment and Reverberation; Formats.

In artistic practiceThinking the Sound Space (the vocabulary of spatial creation and the anti-gadget tests); Diffusion Logics for the Audience.

The genre pagesClassical & Orchestral · Jazz & Acoustic · Pop, R&B & Hip-Hop · Rock & Band · Electronic & Experimental · and the capture chapter.


→ Start with a genre: Classical & Orchestral