Aller au contenu principal

Music Case Study — Classical & Orchestral

In brief
  • Space is captured, not constructed — fidelity to a real hall from a 3D main array.
  • Depth comes free from the hall; spot mics must be time-aligned or soloists jump forward.
  • Width is the room's early lateral energy; envelopment is late ambience (beyond ~80 ms).
  • Height is air and hall, not a shelf for instruments; objects barely used.

Part of Immersive Music Production & Post. Read the shared vocabulary first.

The aesthetic contract

Fidelity to a real event in a real hall. The listener wants the perspective of an excellent seat: the orchestra ahead in believable width and depth, the hall's reverberation arriving from around and above, nothing teleported overhead for effect. This is the genre where space is captured, not authored — the microphone array does most of the composing, and the mix balances perspective and tames the room rather than placing objects. The reflexes of pop are actively harmful here: a spot-mic'd soloist pulled into the surrounds, or width pushed past what the hall gives, reads instantly as fake to this audience.

How the vocabulary applies

The six axes all bend toward capture here.

  • Space — the composed space is the hall. The composed/listening-space negotiation is unusually direct: you are trying to transport one real acoustic to the listener intact. Your primary instrument is not a panner but a 3D main microphone array, which encodes width, depth and height in one coherent capture.
  • Distance and depth — mostly free, easily broken. The hall hands you depth for nothing: the natural direct-to-reverberant ratio, air-absorption roll-off and level differences between near and far desks are already in the main array. The danger is the spot mics: a close mic on a soloist carries a high direct-to-reverberant ratio and jumps forward of its true position unless it is level-matched, time-aligned to the main array (precedence), and given a reverb send that restores its real distance. Distance perception rides on level and D/R and spectrum together (Kolarik et al., 2016); fix only one and the image lies.
  • Size and width — from the room, not a plug-in. Section width should come from the array's spaced, weakly-correlated pickup and the hall's early lateral reflections, which is where Apparent Source Width is built (Griesinger, 1997; ASW rises as interaural cross-correlation falls, Beranek, 2004). Artificially decorrelating a section to "widen" it usually reads as a phasey smear and threatens the stereo fold-down — the opposite of the natural width the array already captured.
  • Trajectory — essentially none. An orchestra does not move; composed motion is out of place. The height and surround content is static air, not travelling objects. This is the one genre where the trajectory vocabulary is deliberately unused.
  • Envelopment — the prize immersive actually delivers. What height and surround buy over stereo is Listener Envelopment: the sense of sitting inside the hall. LEV is built from late, lateral and overhead reverberant energy (arriving beyond roughly 80 ms — Bradley & Soulodre, 1995), so it is captured by dedicated height and surround ambience mics (Hamasaki-style ambience arrays), not synthesised. Getting this right is most of the immersive-classical job.
  • Density — high count, low placement. Dozens of instruments, but organised by the ear into sections and a hall-ground rather than tracked individually — the density budget is met by the acoustic, not by restraint in placement.

Technical specifics

  • Main array plus height. A coherent 3D main system — for example ORTF-3D (Wittek, 2017: two stacked ORTF-Surround planes, four supercardioids each, 10 × 20 cm rectangles) or a Decca-tree-plus-height variant — carries the frontal image and its natural depth. The principles of surround-with-height capture are set out in Theile & Wittek (2011).
  • Ambience for envelopment. Separate, well-spaced height/surround ambience mics feed the late lateral energy that makes LEV; their level against the main array is the master "how much hall" control.
  • Spot mics, blended not placed. Supports for balance only, delayed to the main array and reverb-returned to their true distance so they never step forward.
  • Minimal objects. Most of the mix lives in the bed (the captured field); objects, if any, are a discreet soloist support, never motion. Evaluate against the scene-based attributes — width, distance, depth, envelopment, spaciousness (Rumsey, 2002).

Solutions — with RIPL

  • A captured field as a first-class source. The 3D array and ambisonic room feed enter RIPL's unified source model as native multichannel / scene sources — the captured hall is not flattened into "channels to route" but carried as the field it is, and rendered to the Atmos bed, an ambisonic archive or a binaural stream from one scene.
  • Distance-true supports. A spot mic placed as an object is rendered with RIPL's gain-and-delay model, so its time alignment to the main array — the cue that keeps it from jumping forward — is part of the placement, not a separate patch.
  • No forced motion. The tool does not push spatialisation onto material that wants to sit still; the height layer stays air.

Pitfalls and checklist

Checklist
  • Do the spot mics sit at their true distance (level + delay + reverb), or do soloists jump forward of the orchestra?
  • Is envelopment coming from late lateral/overhead ambience, or are you trying to fake it with early reflections?
  • Did you resist artificially widening sections — does the width survive the stereo fold-down?
  • Is the height layer hall and air, with nothing placed up there for effect?
  • Does the immersive perspective still read as "a great seat," not "a gimmick"?

Bibliography

  • Theile, Günther, and Helmut Wittek. "Principles in Surround Recordings with Height." AES 130th Convention, London, 2011, preprint 8403.
  • Wittek, Helmut. "Development and Application of a Stereophonic Multichannel Recording Technique for 3D Audio and VR (ORTF-3D)." AES 143rd Convention, New York, 2017.
  • Hamasaki, Kimio, Koichiro Hiyama, and Reiko Okumura. "The 22.2 Multichannel Sound System and Its Application." AES 118th Convention, 2005, paper 6406.
  • Rumsey, Francis. "Spatial Quality Evaluation for Reproduced Sound: Terminology, Meaning, and a Scene-Based Paradigm." Journal of the Audio Engineering Society 50, no. 9 (2002): 651–666.
  • Griesinger, David. "The Psychoacoustics of Apparent Source Width, Spaciousness and Envelopment in Performance Spaces." Acta Acustica united with Acustica 83 (1997): 721–731.
  • Bradley, John S., and Gilbert A. Soulodre. "The Influence of Late Arriving Energy on Spatial Impression." Journal of the Acoustical Society of America 97, no. 4 (1995): 2263–2271.
  • Beranek, Leo L. Concert Halls and Opera Houses: Music, Acoustics, and Architecture. 2nd ed. New York: Springer, 2004.
  • Kolarik, Andrew J., Brian C. J. Moore, Pavel Zahorik, Silvia Cirstea, and Shahina Pardhan. "Auditory Distance Perception in Humans: A Review of Cues, Development, Neuronal Bases, and Effects of Sensory Loss." Attention, Perception, & Psychophysics 78, no. 2 (2016): 373–395.

A note on sourcing: the AES-convention and Tonmeistertagung references above appear as convention papers/preprints rather than peer-reviewed journal articles. The Griesinger 1997 volume/page detail and the Rumsey 2002 pagination are corroborated across sources but worth a final check against the publisher before print.

See also

In the technical guideImmersive & 3D recording, Surround recording and Principles of spatial capture (the arrays that do the composing here); Direct, diffuse & envelopment (LEV and ASW), Distance & air and Reverberation.

In artistic practiceThinking the Sound Space (composed versus listening space). See also the shared music hub and the capture chapter.


→ Next genre: Jazz & Acoustic Small Ensemble