Aller au contenu principal

Diffusion Logics for the Audience

In brief
  • The core trade-off: one perfect sweet spot versus fidelity spread across the whole audience.
  • Choose a logic per layer — frontal, point-source, field/ambisonic, or diffuse.
  • Audience geometry (seated, in-the-round, standing-mobile, individual) constrains the choice.
  • Diffusion can be a live interpretive gesture — the lineage from the 1951 pupitre d'espace to tablet control.

The previous chapter was about the composed space — the scene you imagine. This one is about getting that scene to a room full of people. It is where the imagined space meets the unforgiving facts of a real audience: they are not all sitting in the ideal seat, they may be standing or moving, and there may be hundreds of them. The question this chapter answers is not "which product" but "which logic" — what strategy of spatialisation actually serves your intention, given who is listening and from where.

The central question: one perfect seat, or fidelity for everyone?

Almost every diffusion decision reduces to a single trade-off:

Do you optimise for one ideal listening point (a sweet spot, sharply defined, spectacular for the person in it and progressively wrong for everyone else), or do you distribute a good-enough experience across the whole audience (no single perfect seat, but no bad ones either)?

Stereo, and most channel-based immersive formats, are built around a sweet spot: they encode a scene that decodes correctly at one point and collapses toward the nearer loudspeaker everywhere else — the "everyone but the centre gets mono" problem discussed for live sound in Spatial Audio for Live Sound. Other approaches — wavefield synthesis, dense diffuse fields, careful object-based rendering — trade peak precision for coverage, so that a large area gets a coherent, if less pinpoint, image.

There is no correct answer; there is only a decision that follows from your intention and your audience. A twelve-person listening-room premiere can indulge a razor sweet spot. A thousand-seat show cannot, and pretending otherwise just means most of the room hears a broken version of your careful work. Decide, consciously, which side of this trade-off your piece lives on — everything below follows from it.

This is also where Smalley's containment / transcendence returns: the same composed space, delivered into a domestic room, a concert hall, or a cathedral, is contained, matched, or overflowed. The diffusion logic is how you manage that fit.

A decision grid of diffusion logics

Think of the following not as products but as logics — families of strategy, each with a native intention. Real works usually combine them.

Frontal / proscenium

Sound comes from the front, aligned with a stage or screen. This is the logic of audio-visual coherence: when there is something to look at, the ear generally defers to the eye, and fighting that with sound from behind can read as a gimmick. Extended-frontal designs widen and deepen the front image without abandoning the stage relationship. Native to most staged music and to film. See Stage & Live Performance.

Object / point-source

Individual sounds are placed as discrete points and can be moved precisely, ideally with an image that holds over a wide area rather than at one seat. This is the logic of placement and trajectory — of the gesture-as-phrase from the previous chapter realised precisely. It leans on the techniques in object-based audio, amplitude panning (VBAP/DBAP) and, for wide-area precision and sources that seem to stand inside the room, wavefield synthesis.

Field / ambisonic

The whole surrounding sound field is treated as one object that can be rotated, tilted, zoomed. This is the logic of enveloping texture and coherent whole-scene motion — less about pinpointing one source than about the space itself moving. It rests on ambisonics. Native to dome work and to pieces conceived as a single evolving atmosphere.

Diffuse / decorrelated

Many loudspeakers carry a mutually incoherent field so that no single one localises and the listener is wrapped in sound with no clear direction. This is the logic of immersion without a source — the manufactured envelopment of late reverberation, spread across a room. It depends on decorrelation, and on avoiding the comb-filtering trap of simply duplicating one signal, both covered in direct, diffuse & envelopment. It is also the historical heart of the acousmatic diffusion tradition, below.

Hybrids

The honest common case. A frontal anchor for the visible or the intelligible, point-sources for the figures that must be located, a diffuse or ambisonic bed for envelopment. Most real pieces are a considered blend; the value of naming the pure logics is that it lets you be deliberate about which logic each layer of your piece is using, and why.

Diffusion as a live, interpretive gesture

One logic deserves its own history, because it reframes diffusion from a setting into a performance.

In the acousmatic tradition, playing a fixed stereo (or few-channel) work over a large array of loudspeakers is itself an act of interpretation. The performer — the diffuser — sits at a mixing desk and, in real time, moves the sound around a heterogeneous "orchestra of loudspeakers," re-orchestrating the piece for the particular room, much as a conductor shapes a score. The instrument is the Acousmonium, created by François Bayle at the GRM in 1974: dozens of loudspeakers of different sizes, colours and positions, played live.

The lineage is older than the Acousmonium, though. Its direct ancestor is the pupitre d'espace ("space desk") built by the engineer Jacques Poullin with Pierre Schaeffer and first used publicly on 6 July 1951, at the premiere of the Symphonie pour un homme seul: a ring wired to potentiometers that let an operator distribute a mono sound, live and by hand, across five loudspeakers (four around the room plus one overhead) — spatialisation as a real-time gesture, played "like a conductor." The British tradition carried the idea forward from 1982 with BEAST (Birmingham ElectroAcoustic Sound Theatre, founded by Jonty Harrison), which builds arrays of up to a hundred loudspeakers and treats diffusion as genuinely interpretive, near-improvised performance rather than a static spreading of a stereo file. Harrison's own writing ("Sound, Space, Sculpture," Organised Sound, 1998) and Annette Vande Gorne's ("L'interprétation spatiale," 2002) theorise this gesture directly.

A parallel French lineage is easy to miss, and it complicates the picture usefully. Eight months before the Acousmonium, in June 1973, Christian Clozier premiered the Gmebaphone at the Groupe de Musique Expérimentale de Bourges (GMEB, later the IMEB, closed in 2011). It rests on a different principle from Bayle's. Where the Acousmonium is played by controlling the level sent to a heterogeneous orchestra of loudspeakers, the Gmebaphone — renamed the Cybernéphone from 1997, and reaching seventy-six diffusion channels in its 2005 version — worked by spectral division: the signal is split into frequency bands, each routed to loudspeakers specialised for that register, and the performer interprets by shaping the spectrum through filtering, not by amplitude alone. Two philosophies of the loudspeaker orchestra developing in parallel, then — Paris diffusing by loudness, Bourges by filtered spectrum.

The through-line — 1951 space-desk (one operator, five speakers) → 1973 Gmebaphone / 1974 Acousmonium (loudspeaker orchestras, by spectrum and by level) → BEAST (up to a hundred) → today's tablet control and timeline automation — is a single idea: diffusion is not a frozen setting but an interpretive act, adapted to the room. Modern tools continue it in two directions: live control from a tablet during the show, and precise, recallable automation of every spatial parameter along a timeline, so the "performance" can be either played live or composed in advance (the RIPL timeline treats spatial position like any other automatable parameter, a fader you can draw). The Electroacoustic & Acousmatic Music chapter details the diffusion performance in practice.

Audience geometry as a hard constraint

Where the audience is, and whether it moves, constrains the whole decision — often more than the artistic intention does.

  • Frontal, seated (theatre, cinema, concert). Sightlines to a stage; sound generally deferential to the front. A sweet spot is broad and forgiving in the stalls, poor at the edges. Favours frontal and hybrid logics.
  • In the round / surround, seated (dome, immersive concert). No privileged front; the field and diffuse logics come into their own because there is no stage to defer to. But the "one perfect seat" problem is now acute — the centre is ideal and the edges are compromised — which pushes toward wide-area techniques and away from razor-sharp sweet spots.
  • Standing and mobile (club, festival, installation). The audience has no fixed listening point at all; they move through the field. This changes everything: a composed trajectory that only reads from one seat is pointless, and the logic must be robust to a listener who is somewhere different every few seconds. See Electronic & Live Music.
  • Individual and successive (headphone works, walks). The audience is not a simultaneous mass but a sequence of individuals, each having their own private, time-shifted version of the piece — the logic of binaural walks like Janet Cardiff's Her Long Black Hair (Central Park, 2004), where the work is laid over the real place for one person at a time.

The move from a simultaneous audience to a distributed or successive one is one of the deepest choices you can make, and it has clear precedents. Max Neuhaus spent a career insisting on the distinction between the concert — a bounded frame, an audience gathered facing a source — and the sound installation, unbounded and discovered. His Times Square (1977), an unmarked continuous drone rising from a subway ventilation grating in Manhattan, has no announcement, no plaque, no beginning or end; you find it for yourself, by ear, and it becomes, in his words, not an event but a place. (The critic Liz Kotz frames Neuhaus's project as a deliberate refusal of the concert-hall, "narrative" model of spatialised sound — a useful contrast to draw, though it is her framing more than Neuhaus's own words.) At the collective-but-fixed extreme sits Cardiff & George Bures Miller's The Forty Part Motet (2001), perhaps the clearest demonstration in all of sound art that spatial arrangement alone can be the composition.

It takes Thomas Tallis's Spem in Alium (c. 1570) — a motet written for eight five-part choirs, forty independent voices — and gives every voice its own loudspeaker. The Salisbury Cathedral Choir was recorded all together in a single session — the singers standing a few feet apart to keep the parts separate — but each vocal line on its own microphone and channel, and the forty channels are replayed through forty speakers standing on poles at about head height, arranged as an oval in eight groups of five, so the physical layout reproduces the choir's own geometry. There is no live diffuser, no mix and no processing: the spatial "score" is fixed in the arrangement of the room, and the audience composes its own version by where it walks — pressing an ear to a single voice, or standing back in the middle for the full forty-part tutti. Cardiff states the intent exactly: "While listening to a concert you are normally seated in front of the choir, in traditional audience position. With this piece I want the audience to be able to experience a piece of music from the viewpoint of the singers." The loop runs fourteen minutes — eleven of music and a three-minute intermission that keeps the singers' recorded chatter, coughs and warm-up, so the room breathes between repetitions. It is the limit case of discrete, one-source-per-loudspeaker diffusion — a conceptual ancestor of object-based audio in which each "object" is not a virtual point conjured by panning but a real physical source you can walk up to — and it has no sweet spot, because there is no illusion to break.

Unique work versus touring work

As with the composed space, ask whether you are making a one-off or something that must travel.

  • A bespoke, site-specific diffusion studies one room and one audience layout and exploits them completely — the ideal of the acousmatic concert and the installation alike.
  • A portable spatial script encodes the piece abstractly (as objects, as an ambisonic scene, as automation) and is re-rendered per venue for whatever system is present. This is the logic behind object-based and scene-based delivery; it trades the perfect fit of a bespoke design for the ability to play the piece anywhere without rewriting it. The technical machinery for re-rendering a scene to different layouts is in object-based audio and ambisonics.

Deciding this early tells you whether to invest your effort in a room or in a portable representation — and prevents the common trap of building a piece so tied to one system that it dies the moment it leaves.

A short decision checklist

Run a piece through these, in order:

  1. Sweet spot or distributed? Given the audience size and layout, which side of the central trade-off are you on? (Everything else follows.)
  2. Audience geometry? Frontal-seated, in-the-round-seated, standing-mobile, or individual-successive?
  3. Which logic per layer? Assign each layer of the piece a logic — frontal anchor, point-source figures, ambisonic/diffuse bed — deliberately, not by default.
  4. Live or fixed diffusion? Is spatialisation performed in the room, composed as automation, or a mix of both?
  5. Bespoke or portable? One room, or many?
  6. Does the motion survive the geometry? Re-check every composed trajectory against where the audience actually is. A gesture that only reads from the centre seat is not a gesture for a standing crowd.

Common mistakes artists make

These are the failures specific to artistic spatialisation, distinct from the engineering pitfalls elsewhere in the guide:

  • Imperceptible motion. Movement so slow or subtle that no listener notices it — spatialisation that satisfies the automation lane but not the ear. If a gesture matters, make it legible.
  • Confusing "more speakers" with "more spatial." Envelopment and drama come from how sound is placed and moved, not from channel count — Robert Normandeau flatly says the number of loudspeakers "doesn't really matter… it's virtual." A thoughtfully diffused stereo piece can be more spatial than a careless 32-channel one. The Dolby Atmos mixer Matthias Stalter puts the failure mode sharply: movements must "have a purpose, instead of the mix just being a means to demonstrate what a cool new technological gadget you've just discovered."
  • Composing only for the centre seat. Designing on headphones or at the sweet spot and forgetting that most of the audience is off-axis — the single most common cause of work that "doesn't translate."
  • Motion without stillness. Everything moving at once, so nothing reads as movement. Anchors make travellers legible.
  • Diffuse where you meant to locate (and vice-versa). Wanting a precise figure but decorrelating it into a fog, or wanting immersion but leaving a hard point-source that pins the ear to one wall.

The next four chapters take these two frameworks — thinking the space, and choosing a diffusion logic — into the specific realities of four different practices, where the same principles produce very different answers.


Bibliography

  • Austin, Larry. "Sound Diffusion in Composition and Performance: An Interview with Denis Smalley." Computer Music Journal 24, no. 2 (2000): 10–21.
  • Baalman, Marije A. J. "Spatial Composition Techniques and Sound Spatialisation Technologies." Organised Sound 15, no. 3 (2010): 209–218.
  • Brodie, Susan. "Janet Cardiff's The Forty Voice Motet: Sound as Matter." Classical Voice North America, 18 September 2013.
  • Cardiff, Janet, and George Bures Miller. "The Forty Part Motet (2001)." Artist's statement, cardiffmiller.com.
  • Clozier, Christian. "The Gmebaphone Concept and the Cybernéphone Instrument." Computer Music Journal 25, no. 4 (2001): 81–90.
  • Normandeau, Robert. "Timbre Spatialisation: The Medium is the Space." Organised Sound 14, no. 3 (2009): 277–285.
  • Harrison, Jonty. "Sound, Space, Sculpture: Some Thoughts on the 'What,' 'How' and 'Why' of Sound Diffusion." Organised Sound 3, no. 2 (1998): 117–127.
  • Kotz, Liz. "Max Neuhaus: Sound into Space." In Max Neuhaus. New York: Dia Art Foundation / New Haven: Yale University Press, 2009.
  • Neuhaus, Max. "Sound Art?" In Volume: Bed of Sound. New York: P.S.1 Contemporary Art Center, 2000.
  • Smalley, Denis. "Space-form and the acousmatic image." Organised Sound 12, no. 1 (2007): 35–58.
  • Teruggi, Daniel. "Technology and musique concrète: the technical developments of the Groupe de Recherches Musicales and their implication in musical composition." Organised Sound 12, no. 3 (2007): 213–231. (On the pupitre d'espace and the GRM's spatial tools.)
  • Vande Gorne, Annette. "L'interprétation spatiale. Essai de formalisation méthodologique." Revue DÉMéter, 2002.

A note on sourcing: work titles, dates and attributions above follow well-established accounts. A few precise details vary between sources and are worth confirming before print — for instance the exact loudspeaker count of a given Acousmonium or Cybernéphone configuration. The Forty Part Motet figures follow the artist's own statement (fourteen-minute loop, eleven of music plus a three-minute intermission; forty channels mixed down from fifty-nine singers, the soprano lines sung by grouped children). Reading the Motet as a "conceptual ancestor of object-based audio" is our editorial framing, not a claim we found in the literature: no manufacturer or published history of spatial audio we consulted cites the work directly.


→ Next: Practice by Practice