Case Study — Capturing for Immersive
- Immersion is decided at the render, not the microphone — you do not need an "immersive" rig.
- A capture delivers either a coherent field (the space) or clean sources (to place later).
- In an object engine, any microphone — even a mono one — is a spatial object.
- Directivity is a distance choice; reaching spot mics must be time-aligned or they comb.
The question
Ask a room of engineers "how do I record in immersive?" and most will answer with hardware: a 3D microphone array, a tetrahedral Ambisonic mic, a Hamasaki square with height. Those are excellent tools, and the Recording & Capture part covers them in depth. But the premise deserves challenging, because it quietly confuses two different things — capturing and reproducing.
Immersion is decided at the render, not at the microphone. What a capture actually has to deliver is one of two things: a coherent field (the space itself, encoded in the inter-channel relationships between mics), or clean sources (individual sounds you will place later). Whether the result becomes stereo, 7.1.4 or a binaural stream is a rendering decision made downstream. In an object-based engine such as RIPL, the consequence is liberating: any microphone — even a single mono one — is a spatial object, a source with a position you assign, not a channel wired to a speaker. You do not need an "immersive" rig to work immersively; you need clean material and a renderer that treats space as something authored, not captured-and-frozen.
This chapter is about that reframing, and about the choices it opens up.
What a capture can deliver — and the two philosophies
Immersive capture splits along the same line as immersive delivery (see Capturing for Object/Scene Workflows and formats):
- Capture the field. A coherent main array (Decca-tree-plus-height, ORTF-3D, a Hamasaki square) or an Ambisonic microphone samples the space — its width, depth, height and reverberation — as one consistent perspective. What the array captures is essentially what plays back. This is fidelity: the room, transported.
- Capture the sources. Spot, close and accent microphones isolate individual sounds, each later authored as an object with a position (and elevation) you choose. The space is then reconstructed at the render rather than captured. This is freedom: the perspective is decided in the mix, not on the day.
Most real sessions do both — a coherent array for the bed and envelopment, spot mics for the sources — precisely so you can decide later how much of the space is captured and how much is authored (the guide's capture-once-decide-later principle). The object paradigm simply pushes that balance as far as it will go: capture clean sources, build the space downstream.
Different uses of capture
"Recording for immersive" is not one activity. At least four distinct uses coexist, often in the same session:
- Capture a space. You want the place itself — a hall, a church, a forest, a room tone. A coherent array or an Ambisonic mic is ideal, but so is a good stereo pair, or even a single well-placed omni in a room with character: as an object with a captured room around it, it can carry a convincing sense of place. The point is the space, and modest means can capture it.
- Go and get a source at the stage. You want a specific sound out of a crowded stage — one singer, one instrument, one effect — cleanly enough to place and move it. This is the world of spot, close and accent mics, and of directional reach: a more directional microphone effectively lets you sit farther from a source for the same direct-to-reverberant ratio (a cardioid pair reaches roughly √3 ≈ 1.7× farther than omnis — see the two-microphone reach of directivity). Reaching into the stage is how you get objects worth placing.
- Capture an ensemble on a stage. A whole band or ensemble at once: an array of spot mics plus overheads/room, each close mic an object, the room capture the bed. The rock case study is exactly this.
- Capture for movement and re-spacing. When the intent is to move a sound or re-place it per venue, you deliberately capture it isolated — because only an isolated source can become a freely-positioned object. Here the mic technique is chosen for separation, not for a captured image.
Arrays and reaching the stage — technical and artistic consequences
"Going to get the source at the stage" with directional mics and spot arrays has consequences that run in both directions.
Technical consequences.
- Directivity is also a distance choice. Choosing a more directional pattern to reach farther simultaneously sets the direct-to-reverberant ratio and therefore the captured distance — the two cannot be decoupled (principles of spatial capture). And you cannot fix distance later with a fader: the D/R ratio is baked in at the moment of capture by where the mic sits.
- Bleed and time-incoherence. Spot mics on a shared stage are at different distances from every source, so they are time-incoherent with each other and with the main array. Summed naively they comb-filter and smear the image — the first cancellation notch sits at f = c ⁄ (2·Δd) for a path difference Δd (mono compatibility). The fix is to time-align each spot to the array (delay by Δt = distance ⁄ c) and use it sparingly.
- Bleed as glue, not only as enemy. On an ensemble, leakage between mics is often the natural cohesion of the room; managed rather than eliminated, it helps the stage read as one place.
Artistic consequences.
- How you mic decides how free you are later. The more you capture isolated sources, the more you can place, move and re-space them — you are choosing construction. The more you capture the field, the more fidelity you get but the more the perspective is fixed at capture. Microphone technique is therefore already a compositional decision about how much spatial freedom you want downstream, not merely a technical one.
- Reaching changes intimacy. A close, reaching mic gives a present, high-D/R source with little room — intimate and placeable — where a distant array gives the source in its space. Choosing between them is choosing the character of the sound, not just its cleanliness.
Questioning the received paradigms
Seen from an object-based renderer, the two classic immersive-capture paradigms each carry assumptions worth questioning — not because they are wrong, but because they bind the result to decisions made at capture time.
- The main-array / channel-bed paradigm (Theile & Wittek; Williams; Hamasaki) maps microphones directly to a loudspeaker layout. Its strengths are real — physically-correct inter-channel relationships, one coherent perspective — but the guide's own account of it is candid about the cost: you cannot rebalance one instrument without moving the whole image, the direct-to-reverberant ratio is fixed by mic placement, and the array fixes the perspective at one listening point (immersive & 3D recording). It also tends to require many microphones in precise geometry, and ties the capture to a specific delivery format.
- The scene-based / Ambisonic paradigm (Gerzon; Zotter & Frank, 2019) captures the whole field at one point, which is elegant and rotatable — but assumes a single observation point and, at practical orders, offers limited spatial resolution and a real sweet spot. The perspective is again fixed at capture.
The common thread: both decide the space at the microphone. An object-based approach inverts the order — capture material, decide the space at the render — which is exactly what makes "any mic is an object" more than a slogan.
Parametric and object approaches decouple capture from reproduction
A substantial body of work formalises this inversion, analysing a capture into sound plus spatial parameters — or objects plus metadata — that are re-rendered to whatever system is present, rather than fixing the image at capture. A review tutorial makes the split explicit in its very title, treating scene acquisition, modification and reproduction as three decoupled stages (Kowalczyk et al., 2015).
- Directional Audio Coding (DirAC) estimates, at a single point and per time–frequency tile, the direction of arrival and the diffuseness of the field, splits the signal into non-diffuse and diffuse parts, and re-synthesises for whatever loudspeaker layout is present — analysis fixed, rendering adapted to the system (Pulkki, 2007).
- Object and scene coding — MPEG Spatial Audio Object Coding (SAOC) — transmits several objects at roughly the bit-rate of a two-channel signal and lets the decoder render the scene interactively, onto a target loudspeaker configuration or a portable device (Herre et al., 2012). The scene is built at the far end, not fixed at capture.
- Virtual Microphone Control (ViMiC) makes the microphone itself a render parameter: it synthesises the signals of virtual microphones with chosen positions and directivities, so "the perspective" is chosen at reproduction rather than set by a physical placement (Braasch, Peters & Valente, 2008; also in amplitude panning).
Two results bear directly on the number and placement of microphones. Parametric methods have been shown to record perceptually from a small microphone array and reproduce to arbitrary loudspeaker configurations (Politis, Vilkamo & Pulkki, 2015) — the clearest published support for "fewer microphones, less critical positions, layout decided downstream." The field-encoding idea itself runs back to Gerzon's founding Ambisonic proposition (1973), that a sound field can be encoded independently of the speakers that reproduce it, and to wave-field synthesis on the reproduction side (Berkhout, de Vries & Vogel, 1993). What object-based capture adds is to push that independence down to the individual source: capture clean material, and let the render decide the space.
These references establish two sourced claims: that parametric and object approaches decouple capture from reproduction, and that small arrays can feed arbitrary layouts. They do not all support a blanket "use fewer mics." ViMiC concerns virtual microphones synthesised at the render; Gerzon and wave-field synthesis are about field encoding and synthesis, not economical capture. The strong, sourced statement is the decoupling; the specific "fewer mics, looser positions" benefit is demonstrated by parametric small-array methods (Politis et al., 2015) and is the design position RIPL takes, below.
The thesis: any microphone is an object
In RIPL's unified source model, the format of what you captured stops being a routing problem and becomes a source type:
- a mono spot or room mic → an object with a position (and elevation) you assign;
- a stereo pair → a decoded field (HSR) whose primaries and ambience are separated before placement, not two mono channels to scatter;
- an Ambisonic mic → a rotatable scene;
- a native multichannel stem → a field carried as itself.
All four are first-class sources in one scene, rendered together to any layout. Immersive capture, in this light, is not a gear tier you must buy into; it is a way of treating whatever you recorded.
Fewer microphones, looser positions. Because RIPL reconstructs the space at the render rather than reading it off a fixed channel map, a microphone's position informs an object instead of hard-wiring a channel. RIPL's array approach leans on exactly this: you can go further with fewer microphones and less strict positioning than a channel-bed array demands, because the exact geometry no longer has to be the playback format — it feeds a model that is re-rendered to whatever system is present. A pragmatic capture of clean sources plus a sense of the room can become a full immersive scene, where the classical channel-bed approach would have required a larger, more precisely-placed array to capture the same result directly.
(This is RIPL's design position, and the general entailment of object-based capture — capture material, author space downstream. The underlying principle has published support in parametric spatial-audio research, where small microphone arrays are shown to feed arbitrary loudspeaker layouts (Politis, Vilkamo & Pulkki, 2015); the specific RIPL claims are offered as the tool's approach, distinct from that sourced work.)
Solutions — with RIPL
- Treat every input as a source, not a channel. Mono spots, stereo pairs, Ambisonic mics and multichannel stems all enter the one model and are placed, decoded or carried as appropriate — then rendered to Atmos, ambisonics, WFS or binaural from one scene.
- Alignment is placement. Because RIPL renders with gain and delay, time-aligning a reaching spot to the room capture — the cure for comb filtering — is part of positioning it, not a separate patch.
- Decide the space in the mix. With sources captured clean, perspective, distance and width are authored downstream, and re-authored per venue, instead of being frozen on the day.
- Fewer mics, re-rendered anywhere. The captured geometry feeds a model, so a lean capture becomes a scene that renders to whatever system the piece meets next.
Pitfalls and checklist
- Are you buying "immersive" hardware when clean sources plus a renderer would do?
- Did you decide, consciously, how much to capture the field (fidelity, fixed perspective) versus capture sources (freedom to re-space)?
- Are reaching spot mics time-aligned to the array, or are they combing the sum?
- Did you get the array distance right on the day (D/R is baked in — no fixing it with faders)?
- If you intend to move a sound, did you capture it isolated enough to become a real object?
- Does your capture survive being re-rendered to a different layout, or is it wedded to one channel bed?
Bibliography
- Theile, Günther, and Helmut Wittek. "Principles in Surround Recordings with Height." AES 130th Convention, London, 2011, preprint 8403.
- Wittek, Helmut. "Development and Application of a Stereophonic Multichannel Recording Technique for 3D Audio and VR (ORTF-3D)." AES 143rd Convention, New York, 2017.
- Hamasaki, Kimio, Koichiro Hiyama, and Reiko Okumura. "The 22.2 Multichannel Sound System and Its Application." AES 118th Convention, 2005, paper 6406.
- Rumsey, Francis. Spatial Audio. Oxford: Focal Press, 2001.
- Zotter, Franz, and Matthias Frank. Ambisonics: A Practical 3D Audio Theory for Recording, Studio Production, Sound Reinforcement, and Virtual Reality. Springer Topics in Signal Processing 19. Cham: Springer, 2019. Open access (CC BY 4.0). DOI: 10.1007/978-3-030-17207-7.
- Gerzon, Michael A. "Periphony: With-Height Sound Reproduction." Journal of the Audio Engineering Society 21, no. 1 (1973): 2–10.
- Berkhout, Augustinus J., Diemer de Vries, and Peter Vogel. "Acoustic Control by Wave Field Synthesis." Journal of the Acoustical Society of America 93, no. 5 (1993): 2764–2778. DOI: 10.1121/1.405852.
- Pulkki, Ville. "Spatial Sound Reproduction with Directional Audio Coding." Journal of the Audio Engineering Society 55, no. 6 (2007): 503–516.
- Braasch, Jonas, Nils Peters, and Daniel L. Valente. "A Loudspeaker-Based Projection Technique for Spatial Music Applications Using Virtual Microphone Control." Computer Music Journal 32, no. 3 (2008): 55–71. DOI: 10.1162/comj.2008.32.3.55.
- Herre, Jürgen, et al. "MPEG Spatial Audio Object Coding — The ISO/MPEG Standard for Efficient Coding of Interactive Audio Scenes." Journal of the Audio Engineering Society 60, no. 9 (2012): 655–673. (Standard: ISO/IEC 23003-2.)
- Kowalczyk, Konrad, Oliver Thiergart, Maja Taseska, Giovanni Del Galdo, Ville Pulkki, and Emanuël A. P. Habets. "Parametric Spatial Sound Processing: A Flexible and Efficient Solution to Sound Scene Acquisition, Modification, and Reproduction." IEEE Signal Processing Magazine 32, no. 2 (2015): 31–42. DOI: 10.1109/MSP.2014.2369531.
- Politis, Archontis, Juha Vilkamo, and Ville Pulkki. "Sector-Based Parametric Sound Field Reproduction in the Spherical Harmonic Domain." IEEE Journal of Selected Topics in Signal Processing 9, no. 5 (2015): 852–866. DOI: 10.1109/JSTSP.2015.2415762.
A note on sourcing: the AES-convention entries elsewhere in this chapter are convention papers/preprints; the parametric and object-coding references in this section are peer-reviewed journal articles with verified metadata. The exact edition year of the ISO/IEC 23003-2 standard is worth a final check against iso.org if cited by date. The claims attributed to RIPL are the tool's design position — the decoupling of capture from reproduction is sourced; the "fewer mics, looser positions" benefit is supported specifically by parametric small-array work (Politis et al., 2015) and otherwise offered as RIPL's approach, flagged as such and kept distinct from the sourced material above.
See also
In the technical guide — Recording & Capture and Principles of spatial capture (the reach of directivity, mono compatibility); Immersive & 3D recording and Surround recording (arrays and beds); Object-based audio, Ambisonics and Stereo is already spatial (the source types a capture becomes); Formats.
In artistic practice — Thinking the Sound Space (composed versus listening space; site-specific versus portable); Diffusion Logics for the Audience (deciding the space at render time).