Skip to main content

Case Study — Conference, Keynote & Corporate

In brief
  • Priorities in order: intelligible speech first, a scripted immersive reveal second, a clean stream third.
  • Most of the job is not spatial; the reveal is one earned gesture, working by contrast.
  • Speech is a coverage / time-alignment / precedence problem; the stream is a separate fold-down.
  • The speech path must be structurally independent of the immersive layer — a built-in fallback.

The brief

A product launch in a 1,200-seat conference hall. The requirements arrive in a strict priority order, and the order is the whole point of the job: (1) the presenter must be understood, perfectly, in every seat, from the front row to the last row of a raked balcony; (2) at the reveal, a scripted immersive moment — a soundscape that blooms around the audience, an effect that flies from the stage to the back of the room — locked to video and lighting; (3) a clean, self-contained audio feed for the live stream and the archive recording. It runs once, to timecode, with rehearsal time measured in hours rather than days, and it cannot fail in front of press, executives and a camera.

This case is a deliberate corrective to the rest of this part. Most of the job is not spatial at all, and the spatial part earns its place only if it never once compromises the speech. Where the acousmatic concert treats space as the primary material and the live concert treats it as a signature, here space is a single seasoning applied to a few scripted seconds of an otherwise transparent, intelligibility-first show.

Creative problematics and reflections

  • Clarity is the creative brief, not a constraint on it. It is tempting to treat "make the speech clear" as the engineer's problem and "make it immersive" as the creative one. That framing is backwards and it is how these shows go wrong. The creative goal is comprehension: the audience leaving able to repeat the three things the presenter said. Everything spatial is subordinate to that, in the same way that in theatre the words come first and the atmosphere serves them. Treat the transparent, effortless-to-follow mix as the design, and the reveal as a single punctuation mark inside it.
  • One or two reveal gestures, no more. The reveal earns its impact entirely by contrast. If the room has been sitting in a clean frontal mix for twenty minutes, then a soundscape opening up behind them, or a single object flying overhead on cue, lands as a genuine event. If the whole show has been "immersive," the reveal has nothing to push against and reads as more of the same. Decide the one moment the room should open up — occasionally two — and spend the design budget there. This is the material-not-gadget discipline in its purest commercial form: a spatial move that is necessary to the message beats a spatial move that merely demonstrates the rig.
  • Restraint and taste under pressure. Corporate spatial audio fails in a recognisable way: it becomes a demo of the technology rather than a support for the message. The brief from the client may even ask for "immersive" as a shopping-list feature, and part of the creative work is translating that into a single earned gesture rather than a channel-count showcase. The test from diffusion logics applies directly — if a listener could not say why the sound just moved, in terms of the product or the story, the move is decoration and it should be cut. The most sophisticated version of this job often uses less spatial content than the client first imagined, deployed with more precision.
  • The reveal has a narrative reason. Anchor the one immersive moment to something concrete in the presentation — the product appearing, a shift in the story, a change of place. A fly-across that coincides with the logo resolving on screen is a gesture the audience reads as meaning; the same fly-across at a random moment is noise. Coherence with the visible event on screen does the interpretive work for you, exactly as audio-visual coherence does on a stage.

The cases within the case: it is rarely one presenter

Real corporate audio is a set of speech sub-problems, each of which shapes the system before any immersion is added:

  • Single keynote presenter, walking. A lav or headset mic on a moving speaker is the theatre radio-mic problem in miniature — off-axis HF loss, gain before feedback into the same PA, and RF coordination if there are several packs. The presenter often wants to roam the stage, so even frontal coverage and precedence to the stage matter.
  • Panels. Several open mics at once invite feedback and comb filtering; the fix is an automatic mixer (gain-sharing or gating) that keeps only the talking mics open, holding system gain constant as speakers hand off.
  • Audience Q&A. Roving handheld or aisle mics, unpredictable positions, non-expert users — a coverage and feedback problem, and a spatial afterthought at best.
  • Simultaneous interpretation. Interpreter booths and per-language distribution (headsets, or an app) run as parallel deliverables alongside the floor sound.
  • Hybrid / remote presenters. A speaker joining by video needs their audio placed intelligibly in the room and mix-minus routing so remote and in-room participants do not hear themselves delayed.

Technical problematics and reflections

  • Speech intelligibility across a big room is the entire foundation. This is not a spatial problem; it is a coverage, time-alignment and gain-before-feedback problem, and it is where the bulk of the engineering effort belongs. Every seat needs even, direct sound at a consistent level and a consistent arrival time. That means main arrays sized and aimed for the room's geometry (see speaker layouts & topologies), delay fills for under-balcony and deep-room coverage, and the precedence effect used deliberately so the image stays on the presenter rather than pulling to the nearest local speaker. Comb filtering between mains and delays — the classic destroyer of intelligibility — has to be designed out through correct delay times and level tapering, and verified by measurement, not by ear alone under showtime pressure.
  • The stream is a separate deliverable, not a tap off the room. A mix optimised for 1,200 people in a live acoustic is not a broadcast mix, and the immersive layer complicates this further: whatever surrounds the room has to fold down cleanly to the stream's stereo or binaural path without the reveal collapsing into phantom-centre build-up or a smeared, incoherent wash. Plan the stream fold-down as its own render with its own monitoring from the start, rather than discovering on the day that the flying effect that thrilled the room is inaudible or ugly online.
  • Rock-solid show control. This is a one-shot, timecode-locked event running alongside video and lighting, so every cue must fire on time and in the correct order, and there must be a defined, tested behaviour if the spatial engine hiccups mid-show. The reveal cue, the bed entrance and the return to the clean speech mix are all events on a timeline that someone hits "GO" on — or that chases timecode — and the failure modes have to be rehearsed, not hoped away.
  • Minimal rehearsal means everything is pre-built. With hours rather than days in the room, there is no time to mix the reveal live. The spatial content must be authored, stored and recallable in advance, so that rehearsal is spent verifying coverage, timecode sync and the fallback — not building the effect. Anything that depends on live operator dexterity under one-shot pressure is a risk; anything recallable and deterministic is an asset.
  • Redundancy and a defined fallback. "It cannot fail" is a specification. The reinforcement path carrying the presenter's voice should be structurally independent of the immersive layer, run over a redundant audio-over-IP backbone, and have a rehearsed bypass so that if the spatial engine drops, the speech is completely unaffected and the show continues. The immersive moment is expendable; the words are not.

Solutions — the general approach

Build a frontal, intelligibility-first system as the load-bearing structure — main arrays plus time-aligned delay fills covering the whole room, tuned and measured to standard live-sound practice — and treat the immersive reveal strictly as a layer sitting on top of it. That layer is a handful of surround and height speakers carrying object-based effects and an enveloping bed, fired as timecoded cues at the scripted moment and then withdrawn. Author the spatial content abstractly so it recalls identically at every rehearsal and on the night, and derive the stream feed as a separate, clean fold-down render rather than a copy of the room mix. Crucially, keep the speech reinforcement on a signal path that is unaffected if the immersive layer is bypassed — the spectacle is additive, and its absence must be a non-event for the audience listening to the presenter.

The order of operations follows the priority order of the brief: get every seat intelligible and measured first, prove the stream fold-down second, and only then add and rehearse the one reveal gesture on top.

Solutions — with RIPL

  • Speech untouched, immersion layered on. In a RIPL scene the presenter's voice is a simple frontal source rendered for even, consistent coverage across the room, while the reveal's flying effect and its surround bed are additional objects and fields in the same scene. Because the immersive content is separate objects rather than a processing stage inserted into the voice path, the spectacle never routes through — and therefore can never endanger — the speech. The structural fallback (immersive bypassed, speech intact) is a property of how the scene is built, not a patch improvised on the day.
  • Reveal authored once, fired to timecode. The reveal is drawn a single time on the timeline: the fly-across trajectory and the bed's entrance are automation that fires to the show's timecode alongside video and lighting, reproducing identically at every rehearsal and at the performance. There is no live-mixing dependency to go wrong under one-shot pressure — the gesture is deterministic and recallable.
  • Room and stream from one scene. Author-once/render-anywhere means the hall's speaker feeds and the stream's binaural or stereo fold-down are two renders of the same RIPL scene, so the immersive moment reaches online viewers as a coherent fold-down rather than requiring a second, separately-built broadcast mix.
  • Predictable and recoverable. The scene runs over a redundant audio-over-IP backbone, and because the speech reinforcement is structurally independent of the immersive layer, the "safe mode" — speech fully intact, immersive layer dropped — is built into the architecture rather than assembled in a panic if something fails mid-keynote.

Pitfalls and checklist

Checklist
  • Is the presenter intelligible in the worst seat — back of the balcony, extreme side — before any immersive layer exists at all? Verify by measurement, not just a walk.
  • Does the immersive content fold down cleanly to the stream's stereo/binaural path, checked on headphones and a laptop speaker, not assumed?
  • If the spatial engine drops mid-show, does the speech survive completely untouched, and has that bypass been tested — not just designed?
  • Are all reveal cues locked to timecode and rehearsed at least once end-to-end with video and lighting?
  • Is the reveal a single earned gesture tied to the message, or has it quietly become a channel-count demo?
  • Is comb filtering between mains and delays designed out, so intelligibility does not degrade in the fill zones?

See also

In the technical guideSpatial Audio for Live Sound for the reinforcement backbone; time alignment & phase and speaker layouts & topologies for coverage; psychoacoustics for the precedence effect that keeps the image on the presenter; measurement & calibration for verifying intelligibility; object-based audio and direct, diffuse & envelopment for the reveal layer; binaural for the stream fold-down; networking & integration for the redundant backbone.

In artistic practiceThinking the Sound Space for the material-not-gadget test; Diffusion Logics for the Audience for restraint and the single earned gesture; Theatre & Sound Dramaturgy for sound in service of a message.


→ Next: Acousmatic & Fixed-Media Concert