Case Study — Theatre & Live Performance
- Spatial intent is dramaturgical — sound belongs to the stage picture.
- Source-oriented reinforcement makes a voice localise to a moving actor via delay + level.
- The hard problem is tracking (a body mic carries no direction); a tablet operator is today's baseline.
- Intelligibility and repeatable cues override every effect.
The brief
A drama production wants sound to belong to the stage picture rather than sit in front of it. Concretely: a character's amplified voice should appear to come from where the actor is standing and to travel with them as they cross the stage; off-stage voices, footsteps and effects should arrive from specific, believable directions beyond the set; one scene calls for an enveloping storm that surrounds the whole house; and the dialogue throughout must stay pin-sharp and completely intelligible. The venue is a 600-seat proscenium theatre with a steeply raked stalls and a balcony. The show is fully cued — run by a stage manager calling "GO" — and it must be identical every night and safe to hand to a touring operator who has never seen the build.
The distinguishing feature of this case is that the spatial intent is dramaturgical before it is technical. Sound here is in service of the narrative and the visible scene, not an end in itself; the whole approach is framed in Theatre & Sound Dramaturgy, and the discipline of subordinating spatial ambition to the story is the through-line of the artistic side. Read this case as the engineering realisation of that chapter.
Creative problematics and reflections
-
Diegetic geometry — mapping the fiction before the patch. The first document is not a channel plan but a map of where each sound lives in the story: on stage, just off it in the wings, far off in another room, overhead, or all around. That map is a dramaturgical decision — it says what world the audience is in — and it is drawn from the text and the staging, not from the speaker positions available. This is the figure/ground and off-stage (hors-champ) thinking applied scene by scene: the visible stage is the figure, and the composed unseen world is the ground that makes the fiction feel larger than the set. Only once that map exists does it become a patch.
-
Reinforcement that stays put. The core creative-technical hinge of theatre sound is that an amplified voice must appear to come from the actor, not from a loudspeaker stack at the edge of the proscenium. This is an image problem, not a level problem — the audience should never be able to point at the speaker. When the actor moves, the image must move with them, or the illusion that the voice belongs to the body collapses. The artistic stakes of following, extending and — occasionally, deliberately — contradicting the seen performer are laid out in Stage & Live Performance; theatre almost always wants following for dialogue, and reserves contradiction for a conscious dramaturgical shock.
-
Foreground versus world — two logics at once. Dialogue must be located and intelligible; the storm must envelop and have no single source. These are opposite spatial logics, and theatre routinely runs them simultaneously: a tightly anchored voice in the front-centre of the stage while a decorrelated tempest wraps the room. The way to hold both without contradiction is the anchor-plus-bed hybrid described in Thinking the Sound Space — a stable, located figure read against a diffuse ground. Deciding, per cue, which sounds are figures (placed, dry, intelligible) and which are ground (enveloping, decorrelated) is the central compositional act.
-
Reverberation as constructed place. A line delivered dry is in the theatre; the same line given a long stone tail is suddenly in a cathedral, and a tight dead ambience puts it in a cell. Reverberation here is not describing the real room the audience sits in — it builds a fictional acoustic and drops the audience inside it, and a change of reverberation reads as a change of location as surely as a change of set or light. This is developed as invisible scenery in Theatre & Sound Dramaturgy, and the mechanism — and the crucial independence of a room's character from a source's position — is in reverberation. The creative caution is restraint: because a reverberation change signals a change of world, it must be motivated, or it just muddies both meaning and speech.
-
Repeatability is the art form. In theatre the design is the cue stack. Expression does not live in live improvisation but in a sequence of precise, recallable spatial states that land identically on the hundredth performance. Composing for theatre means composing states and transitions, not gestures played by hand — the opposite pole from the acousmatic diffusion performance, and worth being clear-eyed about from the start.
The problem that dominates the job: body-worn radio mics
Everything above assumes you have a clean, controllable voice signal to place. In practice the hardest, most time-consuming part of theatre sound is getting that signal at all, from a moving human body, and it deserves its own treatment because it constrains the spatial design.
- RF coordination is the real bottleneck. A cast of twenty means twenty-plus simultaneous UHF wireless channels that must coexist without interference. Frequencies have to be coordinated to avoid intermodulation products (sum-and-difference frequencies generated by transmitters that land on another channel), and the usable spectrum keeps shrinking — the "digital dividend" has repeatedly sold off UHF TV bands that theatre relied on. This pushes productions toward digital wireless (better spectral efficiency, encryption, but its own latency and coding trade-offs) and toward careful scanning and licensing of the local band before every run.
- Capsule placement fights you. A lavalier or headset capsule sits on the forehead, in the hairline or at the cheek — off the mouth's axis, so it loses high-frequency energy and presence and needs corrective EQ (a presence lift, careful de-essing) just to sound natural. Headset booms sound better and more consistent but are visible; hairline lavs are hidden but dull and position-dependent. Sweat, makeup, wig changes and costume rustle all attack reliability during a run.
- A body mic carries no directional information. This is the crucial spatial consequence: a close body mic captures the voice essentially mono, with no cue about where the actor is on stage. So localisation to the performer cannot be captured — it must be imposed by the reinforcement system, from a known or tracked position. The spatial design is therefore only as good as your knowledge of where each actor is.
- Tracking closes the loop. To make the image genuinely follow a moving actor, productions use performer-tracking systems (e.g. BlackTrax, TTA Stagetracker) that report each actor's stage position in real time, which drives the pan position of that actor's voice object. Without tracking, the alternative is hand-cued position changes timed to blocking — workable for a few marked positions, impractical for free movement.
The cases within the case: play vs musical vs dance
- Straight play. Few mics, dialogue-led, intelligibility is everything; spatial reinforcement is subtle and naturalistic, and the risk is over-designing.
- Musical theatre. The hardest case: a pit or band, a large mic'd ensemble, big numbers where a dozen voices must stay intelligible and stay placed, plus click/timecode for the band. Vocal-to-band balance and mono-cluster intelligibility often compete with spatial placement — you spatialise within a system that must first be intelligible.
- Dance & physical theatre. Voices may be minimal, but bodies move fast and unpredictably, so tracking and enveloping design matter more than dialogue reinforcement; sound often carries the space itself.
- Immersive/promenade theatre. The audience moves through the set — this shades into the installation problem, with no fixed seat and localisation that must hold everywhere.
Technical problematics and reflections
-
Source-oriented reinforcement. Making a voice localise to a moving actor across a raked auditorium and a balcony is fundamentally a precedence-effect and delay problem. The technique is to distribute many modestly-powered loudspeakers around the stage front, proscenium and house, and to time-align them so that the first wavefront the ear receives comes from the direction of the actor — the ear then localises to the stage even though most of the energy may come from a nearer box. Getting this right depends on the material in time alignment & phase and on a sensible speaker layout; getting it wrong produces a voice that sticks to the nearest cabinet and destroys the illusion.
-
Every seat, including the balcony. A theatre has no sweet spot — the design must hold from the front corners of the stalls to the back of the balcony, seats with wildly different distances and angles to both stage and speakers. This rules out a simple stereo pair (whose phantom image only forms on the centre line) and pushes toward the wide-area, source-oriented, delay-based methods above, closely related to the object and WFS families and to standard live-sound coverage practice.
-
Actor tracking — the defining hard problem. If voices are to follow performers, something must know where each performer is, continuously, accurately and with low latency, so the rendered image stays attached to the moving body. This is genuinely difficult — actors turn, are occluded by scenery and each other, and move unpredictably — and it is the same problem examined in depth in Stage & Live Performance. The dependable baseline today is a human in the loop: an operator following the actor by eye and driving the source position live from a tablet or touch surface, which has the virtue of anticipation and musical judgement a sensor lacks, but does not scale past one or two moving sources and cannot be perfectly repeatable. The direction of travel is complementary automatic tracking — optical/markerless vision, ultra-wideband (UWB) tags, inertial sensors, or a fusion of these — feeding position to the renderer while the operator supervises and overrides. Neither extreme is the answer: plan for a hybrid, and never build a show that cannot fall back to a human on a tablet.
-
Cue integrity and handover. Spatial states must recall on "GO", in a fixed and legible order, robustly enough that a relief operator on tour can run the show cold. That means tight integration with the show-control ecosystem (QLab, timecode, MIDI) and a cue list that is a genuine deliverable — human-readable, not a private patch. The design has to survive being handed to a stranger.
-
Intelligibility above all. Every spatial and reverberant idea is subordinate to the audience understanding the words. The diffuse and heavily-reverberant treatments that create envelopment are exactly the ones that smear consonants, so the practical order of priority is fixed: even, intelligible speech coverage first — a measurement-and-calibration and gain-before-feedback discipline — and spatial dramaturgy layered in and around it, yielding to the text wherever the two conflict.
Solutions — the general approach
Build the system as source-oriented reinforcement: distribute many loudspeakers around the stage and house, and treat each amplified source as an object whose position the renderer turns into per-speaker level and delay, so the ear is pulled to the source location rather than to the nearest cabinet. This is the WFS/object family used for narrative ends, not a stereo pair — the same wide-area logic as live sound, aimed at localisation and intelligibility instead of impact. Author the whole show as a cue stack of spatial states, each recalled on "GO" and bound to the show-control system, so the performance is deterministic. Run the two logics together per scene: dialogue as tight, delay-anchored placement; wrap-around moments as a diffuse, decorrelated, enveloping bed; and fictional locations built with deliberately chosen reverberation. For following voices, assume the tablet-driven operator as the baseline and design any automatic tracking as a supervised complement, per Stage & Live Performance.
Solutions — with RIPL
-
Image locked to the stage. RIPL's unified gain-and-delay rendering is precisely what source-oriented reinforcement needs: each voice object is rendered with the per-speaker delays that make the ear localise to the actor's position via the precedence effect, plus the level that keeps the line audible in the back of the balcony — both computed from a single source placement rather than patched by hand. Placement and time-alignment stop being two separate engineering tasks and become one property of the source.
-
Moving actors, stored once. Position is an automatable parameter on the timeline: a voice can be keyframed to follow the blocking, or driven live from a tablet during rehearsal, and either way the resulting move is captured and recalled identically every night — the follow becomes part of the repeatable cue stack rather than a nightly manual chase. Where automatic tracking is available it can feed the same position parameter, with the operator supervising.
-
Anchor and world in one scene. Because an entire scene is a single object model, placed dialogue (located, delay-anchored, dry) and an enveloping storm (a decorrelated field with no source) coexist without contradiction — the frontal-anchor-plus-diffuse-bed hybrid expressed directly, with each element assigned its own logic inside the same authored scene.
-
Cues that tour. Spatial states are stored as recallable cues fired from the show-control system on "GO", so the design reproduces exactly night after night and survives handover to a relief operator; the piece rides a redundant audio-over-IP backbone from venue to venue.
-
Render to the house you have. Author-once/render-anywhere means the same cue stack maps onto each touring theatre's specific speaker plot without re-designing the show — the geometry of tonight's house is simply the current render target, so the dramaturgy is authored once and the rendering adapts.
Pitfalls and checklist
- Do amplified voices localise to the actors — from the balcony and the stalls corners — or do they give themselves away by sticking to the proscenium stacks?
- Is the dialogue fully intelligible in the worst seat before any spatial effect or reverberation is added on top?
- Does every cue recall silently, in the right order, on "GO", with no zipper noise on moving sources?
- If voices follow performers, is there a dependable fallback (a human on a tablet, and a sensible fixed position if a track is lost) so the image never flies off the body?
- Is each reverberation change motivated by the fiction, or is it fogging the words for its own sake?
- Have you walked the extreme seats — front corners, back balcony — not just the mix position?
- Is the cue list legible enough for a stranger to run the show cold on tour?
See also
In the technical guide
- Time alignment & phase and Psychoacoustics — the precedence effect and delay design that make a voice localise to the stage.
- Object-based audio and Wave field synthesis — source-oriented reinforcement over a whole auditorium.
- Speaker layouts & topologies, Measurement & calibration and Networking & integration — coverage, intelligibility and a touring AoIP backbone.
- Reverberation and Direct, diffuse & envelopment — constructed acoustics and the enveloping storm.
- Spatial Audio for Live Sound — the wide-area reinforcement toolkit these techniques draw on.
In artistic practice
- Theatre & Sound Dramaturgy — sound in service of the fiction, the off-stage world, and reverberation as invisible scenery.
- Stage & Live Performance — following/extending/contradicting the visible performer, and the tracking problem in depth.
- Thinking the Sound Space — figure/ground, anchors and travellers, reverberation as material.
- Diffusion Logics for the Audience — audience geometry and the every-seat constraint.