Skip to main content

Stage & Live Performance

In brief
  • With visible performers, the eye leads the ear — spatialise around what is seen.
  • Choose per element: sound that follows, extends, or deliberately contradicts the performer.
  • The defining hard problem is tracking a moving performer — a tablet operator today, automatic tracking as a complement to come.
  • Keep a fallback so the image never flies off the body.

When there are performers on a stage, in view, spatialisation gains a partner and a constraint it does not have in acousmatic or purely electronic work: the eye. The audience can see where the sound "should" come from, and that changes everything. This chapter is about spatialising around visible performers — reinforcing them, extending them, and sometimes deliberately contradicting them — and about the practical problem that dominates the domain: keeping the sound attached to a body that moves.

The eye leads the ear

The founding fact of staged spatial audio is audio-visual coherence. When a listener can see a source, vision generally wins: the sound is perceived as coming from the seen performer even if the loudspeakers placing it are elsewhere (the ventriloquism effect). This is a powerful ally and a strict limit at once.

  • As an ally, it lets a frontal or extended-frontal system (see Diffusion Logics) carry a whole ensemble convincingly: the audience sees the players spread across the stage and hears them there, even though the reinforcement may be doing something looser underneath. It also means you can widen and deepen the stage image without the audience noticing the machinery — the eye papers over the seams.
  • As a limit, it punishes incoherence. Placing a visible singer's voice hard behind the audience reads as a gimmick or an error, because it fights what the eyes report. Spatialisation that contradicts the stage picture has to be a deliberate dramaturgical choice, not a default.

The base logic for staged music is therefore frontal or hybrid: anchor the visible and the intelligible to the stage, and use surround, height and movement for what is not tied to a visible body — ambience, electronics, effects, disembodied voices.

Following, extending, contradicting

Within that frame there are three distinct relationships between sound and the seen gesture, and choosing among them is the artistic work:

  • Following. The sound tracks the performer — a moving musician's amplified sound moves with them across the stage, an instrument's spatial position matching its player's. This deepens realism and can be quietly effective, but it is also the hardest to achieve technically (see below), and, done too subtly, the least noticed.
  • Extending. The visible performer is the anchor, but the sound blooms beyond them — a solo instrument that flowers into an enveloping halo, a voice that is dry and present on stage but throws a reverberant double into the room. The body stays the source; the space around it is composed. This is often the richest option, because it uses the eye's anchoring to license spatial play that would read as incoherent without it.
  • Contradicting. The sound deliberately detaches from the body — a performer stands still while their sound circles the room, or a whispering voice comes from everywhere but the mouth producing it. Used sparingly and intentionally, this is a strong dramaturgical device precisely because it violates the expected coherence; used carelessly it just reads as broken.

The through-line to Thinking the Sound Space: the visible performer is your most powerful anchor, and anchors are what make travellers legible. Staged spatialisation is often most effective when a stable, seen source lets one clear spatial gesture stand out against it.

The tracking problem

The moment you want sound to follow a moving performer, you hit the defining practical difficulty of this domain: something has to know where the performer is, in real time, accurately enough and fast enough that the sound stays attached to the body.

This is genuinely hard. A performer walks, turns, is occluded by scenery or other performers, moves fast and unpredictably, and the spatial rendering needs their position continuously and with low latency. Get it wrong and the sound lags, jumps, or drifts off the body — worse than not tracking at all, because the incoherence is now moving.

The tablet: the practical solution today

In current practice, the reliable approach is usually a human in the loop. An operator watches the stage and drives the performer's spatial position live from a tablet (or a touch surface / control desk), dragging the source to follow the performer, correcting by eye and ear. This works, and it has real virtues: a skilled operator anticipates, smooths, and makes musical decisions a sensor cannot — nudging a position for effect, holding it steady through an occlusion, treating the follow itself as a small performance. It is the direct descendant of the acousmatic diffuser at the desk (see Electroacoustic & Acousmatic Music): a person interpreting space in real time.

Its limits are equally real. It needs a dedicated, skilled operator per show; it does not scale to many simultaneously moving performers; precision and latency depend on human reaction time; and it cannot be perfectly repeatable night to night. For a single soloist it is excellent; for six performers moving at once it breaks down.

Toward complementary automatic tracking

The tablet should therefore be understood as the dependable baseline, not the ceiling. The clear direction of travel is automatic performer tracking as a complement to — not a replacement for — the human operator: sensing a performer's position directly and feeding it to the spatial renderer, so the operator supervises and shapes rather than manually chasing every step.

Several families of technology point this way, each with trade-offs to weigh rather than a settled winner:

  • Optical / camera-based tracking (markers, or increasingly markerless computer vision): good spatial accuracy, but vulnerable to occlusion, stage lighting, and costumes, and it raises latency and calibration questions.
  • Radio-frequency positioning, especially ultra-wideband (UWB) tags worn by performers: robust to line-of-sight problems and low-latency, at the cost of putting a device on each performer and installing anchors around the stage.
  • Inertial sensors (IMUs) on the performer, often fused with one of the above to smooth and de-jitter the estimate.
  • Sensor fusion combining these, since no single modality is reliable alone under theatrical conditions.

None of these is yet a plug-and-play standard for live performance, and each imposes its own setup and failure modes — which is exactly why the tablet-driven human operator remains the pragmatic default today. The productive stance is to plan for a hybrid: build shows so that a human can always drive the position from a tablet (the fallback that always works), while designing in the capacity to accept an automatic position feed where the venue, budget and reliability allow it, with the operator supervising and overriding. Treating tracking as "solved by the tablet" forecloses the better systems coming; treating it as "solved by sensors" ignores how unreliable they still are on a real stage. The honest position is that this is an open problem worth investing in, and that complementary automatic tracking is where the practice should be heading.

Working method

  • Anchor to the stage first. Establish the frontal/hybrid image that matches the visible performers before adding any detached spatial play.
  • Choose the relationship per element. Decide, for each sound, whether it follows, extends or contradicts the body — and make contradiction a deliberate, sparing choice.
  • Make following legible or don't bother. If a follow is too subtle to notice, it is cost without benefit; either commit to a follow the audience can perceive, or leave the source anchored.
  • Plan the tracking honestly. Assume a tablet-driven operator as the baseline; add automatic tracking only where you can trust it, and keep the human able to take over instantly.
  • Protect coherence under failure. If a tracked position is lost, the sound should fall back to a sensible fixed stage position, not fly off — the same fallback discipline as everywhere in live work (live sound).

Bibliography and references

  • Chion, Michel. Audio-Vision: Sound on Screen. Trans. Claudia Gorbman. New York: Columbia University Press, 1994. (On audio-visual coherence and the ear's deference to the eye.)
  • Smalley, Denis. "Space-form and the acousmatic image." Organised Sound 12, no. 1 (2007): 35–58. (On performed/gestural space versus acousmatic space — see the discussion of Clarinet Threads.)
  • See also the guide's technical treatment of stage systems and delay in Spatial Audio for Live Sound and speaker layouts & topologies.

A note on sourcing: the survey of tracking technologies above describes the current state of a fast-moving field; specific systems and their performance should be evaluated directly for any given production.


→ Next: Theatre & Sound Dramaturgy