Skip to content
AI.info

Speech and audio

Microphone Arrays, Beamforming, and Spatial Audio

Explain array geometry, steering, beamforming, localization cues, spatial aliasing, and evaluation for multi-channel audio.

By the end you can

Example

Two devices that changed speed mid-session

Twenty dinner parties were recorded, and the recordings were not synchronous. That is why CHiME-6 exists: it is CHiME-5's audio, re-released with array synchronization. The baseline tool published with it says plainly what it is for. "It compensates for two separate issues: audio frame-dropping (which affects the Kinect devices only) and clock-drift (which affects all devices)." Drift was estimated by cross-correlating each device against the session's reference binaural recording, then removed by resampling with sox. The adjustment is typically smaller than 100 ms over a 2.5 h session.

Nothing in that is audible. A hundred milliseconds spread across two and a half hours changes no channel's sound. Every channel played on its own would pass. What it changes is the estimated direction, because direction is computed from the differences between channels.

And the correction did not always work. It failed outright for devices S01_U02 and S01_U05. Those two changed speed mid-session and had to be fitted piecewise linearly instead.

  • Those two devices set the agenda for everything below: array geometry, steering, beamforming, localization cues, spatial aliasing, and how any of it gets evaluated.
  • The fault that does the damage is the pair CHiME-6's tool had to remove by hand — dropped audio frames on the Kinect devices, clock drift on every device. Ordinary audio QA is not built to detect either.
  • The evidence that catches it is direction or position error, broken out by angle, by distance and by reverberation, plus a reference recording to cross-correlate against. That is exactly what the binaural channel gave CHiME-6.
  • The response is cheap enough to be embarrassing. Inject a known signal at each microphone in turn, confirm the order and the gain, then keep measuring drift after the test has passed. S01_U02 and S01_U05 changed speed while the session was running.

Case

Twenty dinner parties, one hundred hours, three instrumented rooms

A published array recording tells you what the geometry actually is. CHiME-6 records twenty dinner parties, more than forty hours in total, four people to a session. It captures them through six far-field Kinect arrays of four microphones each, plus binaural microphones on every participant, at 16 kHz.

The AMI Meeting Corpus had made the same choice earlier, and documented it down to the part number. It is 100 hours of meetings in three instrumented rooms — Edinburgh, Idiap and TNO — with close-talking and far-field microphones running together. Of the Edinburgh table, the corpus documentation says "16 Sennheiser MK2E-P-C miniature omni-directional electret microphones are arranged in two 10cm radius circular arrays of eight." That table was captured at 48 kHz/16-bit. Idiap used an 8-element 10 cm-radius table array plus a 4-element ceiling array. TNO used an 8-element circular array plus a 10-element linear array above the screen. The corpus paper appeared in 2005.

What both corpora are protecting is the close-talking reference. Without one you cannot tell a beamforming gain from a lucky seat. And, as CHiME-6 showed, without a reference channel you cannot even measure your own drift. Keep those rooms in mind for the rest of the lesson: four talkers, six arrays watching them from across the room, two circular arrays of eight on a table in Edinburgh, and a per-speaker recording to check any claim against.

Analogy

A telescope made from synchronized ears

What those six Kinect arrays are doing is closer to astronomy than to recording. Several observers time the same flash. They compare their stopwatches to infer where it came from, then point their attention at one region together.

The comparison of stopwatches is the entire mechanism. It is also precisely what drift corrupts, silently, at well under a tenth of a second across a two-and-a-half-hour dinner.

The picture marks its own limit too. Light arrives once and travels straight. Sound reflects, bends, overlaps, and behaves differently at every frequency. That is why a dinner party is a harder object than a star.

Array performance depends on geometry and synchronization as much as on model capacity.

A narrow beam proves nothing about sources

Because sound behaves that way, an array cannot simply be aimed. It infers spatial structure from relative delay, phase, level and coherence across synchronized channels. Those quantities mean something only if you know where the microphones are and when each sample was taken. Beamforming combines the channels to emphasize selected directions. Localization estimates a source's direction or position under geometric and acoustic assumptions.

Those assumptions are the operational boundary, and the price of crossing it has been measured. A 2023 study simulated 100 two-minute meetings recorded by three independent devices. One carried a 4-microphone array in a 5 cm square; the other two had a single microphone each. The sampling rate offsets varied over time, their mean drawn from a uniform distribution over [−100 ppm, 100 ppm]. Synchronous, the two extra microphones cut cpWER from the single array's 18.57% to 11.81%. Asynchronous, the same microphones bought only 16.08%. Compensating the sampling rate offset restored 11.88% (DWACD-based) and 11.97% (spatial-covariance-matrix-based).

Gburrek and colleagues are blunt about what that means: "However, it becomes obvious that the sampling rate of the devices have to be synchronized to make use of the whole potential of the additional microphones." Most of what the extra microphones were bought for is spent on the clocks. Compensation buys it back.

So spatial processing cannot overcome unknown geometry, unsynchronized clocks, severe reverberation, spatial aliasing, or multiple sources that violate the chosen model. A narrow beam is not proof that only one source contributed. Defensible work runs from calibrating geometry and clocks through to stressing that geometry, with somebody named at both ends.

Ship the array without measuring geometry and clocks and you keep 16.08% where 11.81% was available: the extra microphones are paid for, and the offsets spend them.

Visual

Measure the microphone positions yourself

Calibration and the stress test are the two ends of that run. Between them sit measurements that are easy to blur together. Geometry, clocks, beams and downstream quality fail for different reasons and are reported in different units. The path below keeps them separate rather than folding independent kinds of error into one score.

The first step is the one most often skipped, and the LOCATA corpus shows what it costs and what it yields. Source and array positions came from a 10-camera OptiTrack Flex 13 system, giving roughly 1 mm marker accuracy at 120 Hz. The room was 7.1 × 9.8 × 3 m, with T60 ≈ 0.55 s. The microphones themselves were not tracked. The corpus documentation is candid about the gap: "The positions of the microphones in relation to the marker positions were derived by means of technical drawings of the microphone positions and caliper measurements with an estimated accuracy of a few millimeters." Reconstructed trajectories deviated by less than 12 mm for all objects, and below 6 mm for most.

Even that budget has a hole in it that only measurement could find. The DICIT array was slightly bending back and forth while being moved for Tasks 5 and 6. That violates the rigid-array assumption used to compute its microphone positions. Drawings, calipers, an optical tracker, a millimetre-scale residual — and an array that bends.

FigureTimeline · 4 stops
  1. 1. Calibrate geometry and clocks

    Measure microphone positions, channel order, delay, gain, and frequency response.

  2. 2. Define the spatial objective

    Choose source enhancement, direction estimation, separation, or scene rendering.

  3. 3. Select an array method

    Use delay-and-sum, adaptive beamforming, neural spatial filters, or localization according to assumptions.

  4. 4. Stress the geometry

    Test movement, reverberation, multiple speakers, reflections, and partial channel failure.

Geometry and clock calibration hand their numbers forward as fact, and the stress test on the geometry inherits those numbers instead of questioning them.

Key idea

Ordinary audio QA never sees a channel swap

Fold those measurements into one good angular-error number and it will hide four distinct problems at once.

The first is the pair CHiME-6 had to remove by hand: dropped audio frames and clock drift. Ordinary audio QA never catches them, because listening channel by channel is exactly the wrong test. A drift typically smaller than 100 ms over a 2.5 h session is inaudible on any single channel and decisive across four.

The second is spatial aliasing. It comes out of the microphone spacing itself, and it bites at high frequencies. LOCATA makes that trade-off physical rather than abstract. Its planar DICIT array puts 15 microphones across a 2.24 m horizontal aperture, and the documentation records why: "It contains four linear uniform sub-arrays with microphone spacings of 4, 8, 16 and 32 cm." The nested spacings were chosen expressly to cover arrays with large microphone spacings. The same corpus carries a 32-microphone Eigenmike on an 84 mm rigid baffle, a 12-microphone pseudo-spherical NAO robot head, and hearing-aid dummies whose two microphones sit 9 mm apart, 157 mm ear to ear. Spacing is a design choice with a frequency written into it.

The third problem is overfitting to one fixed room or one array geometry, which a system trained and evaluated in the same room will never reveal. The fourth is reporting angular error and stopping there.

And steering more sharply is not a remedy for any of them. A tighter beam does not empty itself. It changes which mixture arrives at the output, and these four decide what else stays inside it.

A drift smaller than 100 ms across 2.5 hours is inaudible on every channel and still moves the angle — ordinary audio QA passes the file without hearing a thing.

Example

The array answers four different questions

Where those four problems hurt depends on what the array is being asked for, and the LOCATA results show how much of the answer is geometry rather than method. Across the submissions, azimuth error was about 1.0° for the static single-source Task 1 with the robot head, Eigenmike and DICIT arrays. In the same room, hearing aids gave 8.5°.

Motion costs again on top of that. The challenge report states it directly: "The mean azimuth accuracy over S135, averaged over the corresponding submissions and arrays, decreases from 5.5° for Task 3, using static arrays, to 9.7° for Task 5, using moving arrays." Tracking recovered accuracies of up to 1.8° for DICIT, 3.1° for the robot head and 7.2° for hearing aids in the dynamic tasks. The results appeared in IEEE/ACM TASLP in 2020. Ask each application for its own units and its own evidence.

  • Far-field ASR treats the beam as a cleaning stage: beamforming can improve the target speech before recognition starts. On CHiME-6 that is measurable as a change in word error rate with everything else held fixed.
  • Meeting diarization can let location support the assignment of speech to speakers, so long as location is never allowed to stand in as identity evidence.
  • In robotics, localization and scene awareness have to hold up against the platform's own motion and self-noise. LOCATA's averages go from 5.5° with static arrays to 9.7° once the arrays move, and the hearing-aid geometry sits at 8.5° before anything moves at all.
  • Immersive media pulls the other way. Spatial rendering rests on assumptions about the channels and about head-related transfer — the 84 mm Eigenmike baffle and the 157 mm ear-to-ear spacing are those assumptions in hardware.

Example

A steering vector is a model

The work in those paragraphs was done by four terms. Near-synonyms in this field carry different proof, different units and different rights over a decision. The distinction worth holding on to is between what you assume and what you observe. A steering vector is a model. A time difference of arrival is a measurement. Beamforming and spatial aliasing sit either side of the pair. LOCATA's DICIT array is the case in point: the microphone positions fed to the model came from technical drawings and calipers, and the array bent anyway.

  • Beamforming is the combining of microphone channels to emphasize or suppress spatial directions.
  • A time difference of arrival is the relative arrival delay of one source across the sensors — the quantity CHiME-6 corrupted by drift and recovered by cross-correlating against the binaural reference.
  • Spatial aliasing is the ambiguity that appears when sensors are spaced or sampled too coarsely to tell one spatial pattern from another. DICIT's 4, 8, 16 and 32 cm sub-arrays are that ambiguity traded deliberately against aperture.
  • A steering vector is a model of the array response for a source direction or position — an assumption written down, not something observed, and only as good as the drawing and the caliper behind it.

Example

Angular error is not speech quality

On identical CHiME-6 audio, with oracle segmentation, only the multi-channel front end changed. Track 1 word error rate went from 69.8% dev / 61.2% eval with BeamformIt to 51.8% / 51.3% with guided source separation. The models were not the variable. The challenge paper's own table caption settles it: "We used the same acoustic and language models for both tracks." Track 2 ran system diarization instead of oracle segments and sat at 84.3% / 77.9%. That baseline diarization scored 63.4% DER and 70.8% JER on dev, 68.2% / 72.5% on eval. No angular number reports any of that.

Angular error also moves for reasons that have nothing to do with the source. DCASE 2020 Task 3 released TAU-NIGENS Spatial Sound Events 2020: 600 one-minute development recordings plus 200 evaluation recordings at 24 kHz, 14 classes. The same scenes came in two formats, first-order Ambisonics and a four-channel tetrahedral array. The baseline SELDnet scored localization error 22.8° with recall 60.7% and F20° 37.4% on FOA, against 27.3° / 59.0% / 31.4% on the microphone-array format. Within FOA it scored 18.1° on single-source segments against 26.3° on overlapping ones. The authors flag the format gap themselves: "In Table 2, although both the FOA and MIC datasets are synthesized from the same microphone array, the SELDnet is observed to perform better for FOA than the MIC dataset." Same scenes, same array, two channel formats, several degrees of difference.

So name the unit. Say which arrays and which rooms it was measured over, how uncertain it is, and what the system was running under.

  • The core task number is direction or position error, reported by angle, by distance and by reverberation — 18.1° against 26.3° once sources overlap, on the very same DCASE 2020 scenes.
  • Array gain and speech distortion describe how the system behaved: what the beam did to the audio, not to the estimate.
  • The robustness slice is behaviour under channel dropout and calibration error. That is the slice CHiME-6's frame-dropping and drift correction lives in, and a single angular figure cannot expose it.
  • Lifecycle evidence is the change in downstream ASR, separation or detection once the beam has done its work: 69.8% / 61.2% to 51.8% / 51.3% on CHiME-6 Track 1, with the acoustic and language models untouched.

Report direction or position error by angle, distance and reverberation together with the downstream number — on CHiME-6 the front end alone was worth 69.8% against 51.8% on dev.

Steps

Run an array commissioning test

Commissioning is where an array stops being a diagram and becomes a room with chairs in it. LOCATA's team did it with technical drawings, calipers accurate to a few millimetres and a 10-camera optical tracker. They still found the DICIT array bending back and forth while it was moved. CHiME-6's team did it by cross-correlating every device against a reference binaural recording. They still had two devices, S01_U02 and S01_U05, change speed mid-session. Design the test so that another team can argue with your evidence that the channels are in the right order and that the clocks agree.

Then write three things down: what your calibration of geometry and clocks assumes, one case that would break it, and what a failed stress test on the geometry obliges you to do next.

FigureProcess · 4 steps
  1. 1. Verify channel identity

    Inject a known signal at each microphone and confirm order and gain.

  2. 2. Measure delays

    Use impulses or synchronized sources to estimate relative timing.

  3. 3. Sweep source positions

    Test azimuth, elevation, distance, movement, and reflections.

  4. 4. Create failure alarms

    Detect channel drift, dropouts, clipping, and geometry changes in production.

Commissioning is worth running only if the same checks keep running afterwards: S01_U02 and S01_U05 changed speed after the recording had started, and no pre-flight test could have caught that.

Key takeaways