Speech and audio
Bioacoustics and Environmental Monitoring
Design bioacoustic and environmental monitoring around sensor placement, species calls, occupancy, soundscapes, domain shift, and ecological inference.
By the end you can
- Define bioacoustics and environmental monitoring as an operational problem with explicit inputs, outputs, and boundaries
- Distinguish call detection, occupancy inference, and soundscape index without treating them as interchangeable
- Trace the workflow from define the ecological estimand through model observation bias
- Evaluate bioacoustics and environmental monitoring using event precision and recall under independent sites and seasons and evidence from difficult deployment slices
Example
More detections did not mean more animals
Twelve microphone models went into the field together. The point was to measure what the microphone alone does to a wildlife survey. The 2020 result is the whole lesson in two sentences: “Microphone signal-to-noise ratio positively affects the sound detection space areas, which increased by a factor of 1.7 for audible sound, and 10 for ultrasound, from the lowest to the highest signal-to-noise ratio microphone. Consequently, the sampled vocalisation activity increased by a factor of 1.6 for birds, and 9.7 for bats.”
Nothing in those factors is about animals. Between the lowest and the highest signal-to-noise ratio microphone the birds did not become more numerous, more talkative or more active. The recorded bird activity rose by 1.6 and the bat activity by 9.7 because the sound detection space around the sensor had grown — by 1.7 for audible sound, by 10 for ultrasound. A project that swaps its recorders between seasons and then reports a rise has bought that rise from a supplier.
The obvious safeguard would not catch it. A review that only checked detection scores across different sites and seasons would see the same rise wherever the new hardware went. The extra detections come from the equipment. The equipment travels with the survey.
- The decision here is a wide one. Bioacoustic and environmental monitoring has to be designed around sensor placement, species calls, occupancy, soundscapes, domain shift, and ecological inference. A change of microphones touches most of that list at once, at a measured 1.6× for birds and 9.7× for bats.
- The failure that recurs most often across that list is site leakage, where a detector keys on the background or the recorder signature instead of the animal.
- The evidence that settles it is event precision and recall measured on sites and seasons the detector has not seen. The Bird Audio Detection challenge added one condition: score it site by site, not pooled.
- The practical response costs one sentence: state the population, the site, the season, and the unit you are drawing an inference about.
Case
Detections are not animals: occupancy since 2002
The problem the microphones exposed is old. Ecology named it, and solved the statistical half of it, long before passive acoustics existed. A 2002 paper by MacKenzie and colleagues opens with the sentence the field has repeated ever since: “Nondetection of a species at a site does not imply that the species is absent unless the probability of detection is 1.” The paper gave a likelihood-based estimator for occupancy when detection probability falls below one, and ran it on 32 Maryland wetland sites surveyed in 2000 under the Frogwatch USA programme. Estimated American toad occupancy came out at 0.49 — 44% higher than the naive proportion of sites where the species was actually observed. For spring peeper the estimate was 0.85 against an observed 0.83. One survey, two species, two very different sizes of correction. The gap between what a survey records and what is there is itself a quantity that varies. It is not a constant you can assume away.
Better detectors do not dissolve it either. BirdNET is a 157-layer ResNet-derived model with more than 27 million parameters, covering 984 North American and European species, and it was reported in 2021. Its authors tested it on three datasets: 22,960 single-species recordings, 286 h of fully annotated soundscape data from an array of autonomous recording units, and 33,670 h from a single omnidirectional microphone near four eBird hotspots. Their own summary reads: “In summary, BirdNET achieved a mean average precision of 0.791 for single-species recordings, a F0.5 score of 0.414 for annotated soundscapes, and an average correlation of 0.251 with hotspot observation across 121 species and 4 years of audio data.” The Cornell Lab of Ornithology now states the released model recognises over 6,000 species.
Read those three numbers in order, because the order is the argument. 0.791 on clean single-species clips. 0.414 once the audio is real annotated soundscape. 0.251 when the output is compared with what human observers actually logged at the hotspots. That is a serious instrument, with published figures for how often it is right. It is still an instrument. What it returns is detections, and a count of detections that ignores detectability is a count of your microphones.
Seven things stand between the animal and the file
Between an animal and a line in a detection log sit seven separate conditions. The microphone test changed one of them without going near a bird. Bioacoustic systems detect, classify, localize, or summarize animal and environmental sounds. Whether any particular call reaches a file depends on whether the animal is there, whether it calls, how the sound travels, the weather, whether the sensor is working, where somebody put it, and how sensitive the detector is. Only the first of those is the thing an ecologist wants to know about. Move the last one and, in the microphone study's words, “the sampled vocalisation activity increased by a factor of 1.6 for birds, and 9.7 for bats.”
That is the operational boundary, and it is worth stating flatly. Detection counts are not direct abundance estimates. A silent species may be present. One individual may call repeatedly. Recording conditions can change detectability across space and time. The defensible version of the work starts by defining the ecological estimand — the quantity in the world you are trying to estimate — then models observation bias against it, and says out loud who owns each end.
Counting detections and calling it abundance turns a quiet week, a dead sensor or a better microphone into an ecological finding.
Key idea
The recorder signature travelled with the species
Swapping hardware leaves the confound in plain sight. Training a detector on your own recordings moves it inside the model, where it is far harder to see. Background noise and recorder signatures identify a site the way a fingerprint identifies a hand. A detector that has learned the recorder rather than the species can still score well on every site and season it is tested against, because the held-out recordings carry the same hardware.
The first Bird Audio Detection challenge measured exactly that, and reported in 2018. The public development data was 7,690 freefield1010 items plus 8,000 Warblr items. The private test data was 10,000 Chernobyl items plus 2,000 Warblr items. Thirty teams entered, and the best score was 88.7% AUC, from team 'bulbul' (Thomas Grill). That is the headline the paper reports: “Multiple methods were able to attain performance of around 88% AUC (area under the ROC curve), much higher performance than previous general-purpose methods.” The instructive numbers sit underneath it. Of the paper's two baseline classifiers the stronger reached over 85% AUC in matched conditions but dropped to 79% AUC under mismatched conditions, and the simpler GMM baseline degraded dramatically further. Re-analysed per site instead of pooled, the second-placed 'cakir' system would have ranked tenth on average per-site AUC.
1) Site leakage through background or recorder signatures. 2) Sensitive species locations exposed through metadata. 3) Using one region’s calls as universal species templates. 4) Reporting detector counts as population change without an observation model.
The fourth is the least technical of the four and the easiest to commit. A trend in counts becomes a trend in population only through an observation model that says how the counts were produced. Without one, “detections rose” and “the population grew” are the same claim written twice. The second version is the one that reaches a decision-maker.
A baseline over 85% AUC in matched conditions fell to 79% when the conditions changed — and pooled scoring still ranked second a system that per-site scoring put tenth.
Comparison
A pattern found, or an ecological claim
BirdNET's “a F0.5 score of 0.414 for annotated soundscapes” and a rising detection count answer two different kinds of question, and only one kind is ecological. Call detection finds occurrences of an acoustic pattern and claims nothing beyond that. An F0.5 on annotated soundscapes measures precisely that, and measures it honestly.
Occupancy inference and a soundscape index are ecological claims, and they are not equally well supported. Occupancy has an estimator and a detection probability attached to it. A soundscape index has a correlation. A 2022 meta-analysis pooled 364 effect sizes from 34 studies covering the 11 most used acoustic indices. Against diversity metrics it got r = 0.33, CI [0.23, 0.43]. The evidence an ecological claim needs is ecological. That is why no detection score, however good, can be handed over in its place.
Call detection
Finds occurrences of an acoustic pattern.
- Decision focus: Define the ecological estimand
- Useful evidence: Event precision and recall under independent sites and seasons
- Watch for: Site leakage through background or recorder signatures
- Best used when its assumptions are documented for bioacoustics and environmental monitoring
Occupancy inference
Estimates whether a species is present while modeling imperfect detection.
- Decision focus: Design the sensor network
- Useful evidence: Occupancy or activity estimates with detection uncertainty
- Watch for: Sensitive species locations exposed through metadata
- Best used when its assumptions are documented for bioacoustics and environmental monitoring
Soundscape index
Summarizes acoustic properties and may correlate with, but not equal, biodiversity.
- Decision focus: Build detection and review
- Useful evidence: Sensor uptime, calibration, weather, and coverage
- Watch for: Using one region’s calls as universal species templates
- Best used when its assumptions are documented for bioacoustics and environmental monitoring
Example
Ecological terms with hard definitions
Those ecological claims get made in four words, and each of the four has a hard definition that a detector count does not meet. Occupancy, detectability, soundscape, and acoustic index each carry their own evidence, their own unit, and their own limit on what may be concluded. Borrowing one of them to describe a detector’s output is how a count starts sounding like a finding.
- Occupancy is the probability or state of presence within a defined site and period. That is why the site and the period have to be written down before the word is used at all. At 32 Maryland wetland sites surveyed in 2000, the estimated American toad occupancy of 0.49 stood 44% above the naive proportion of sites where the species was observed.
- Detectability is the probability of observing a target given that it is present. It is exactly the quantity a microphone alters. From the lowest to the highest signal-to-noise ratio model in the 12-microphone field test, the sound detection space area grew by a factor of 1.7 for audible sound and 10 for ultrasound.
- A soundscape is the whole collection of biological, geophysical, and human-made sounds in an environment, not only the species anyone is listening for.
- An acoustic index is a summary statistic derived from recordings, and it is not a direct biodiversity measure. The 2022 meta-analysis says so: “Overall, acoustic indices had a moderate positive relationship with the diversity metrics (r = 0.33, CI [0.23, 0.43]), and showed an inconsistent performance, with highly variable effect sizes both within and among studies.” Around 40% of the studies showed pseudoreplication. The best performer was the acoustic entropy index H, at r = 0.50, CI [0.36, 0.62]. The authors conclude the indices are far from being direct proxies for biodiversity.
Visual
Weather and uptime belong in the model
Hardware and placement are not the only conditions that move. Weather and sensor uptime shift a detection rate as effectively as a new microphone does, and they shift it week to week rather than once. They belong inside the observation model as terms, not in the discussion afterwards as caveats. A caveat cannot be subtracted from a count.
Spotted owls and barred owls put a number on what that costs. A 2019 study reported: “Based on ~6-night passive-acoustic surveys, site occupancy and detection probabilities were 0.43 and 0.50, respectively, for the common but declining species (the spotted owl), and 0.09 and 0.67, respectively, for the rare but increasing competitor (the barred owl).” A detection probability of 0.50 says that roughly half the six-night surveys at occupied spotted owl sites came back with nothing in them. The two species also differ: 0.50 against 0.67. The same recording effort sees them at different rates, so a comparison between their counts is a comparison between two detection probabilities unless the model says otherwise.
The design consequence is where the numbers bite hardest. Simulations showed what it takes to detect a 2% annual decline in spotted owl occupancy within 10 years with more than 80% power: 1,000 sites surveyed three times per season, or 1,500 sites surveyed twice. That is the real price of the trend claim, and it is set before any audio is recorded. The path below keeps measurement, modeling, decision, and verification visible as separate steps rather than collapsing them into one number. That is what lets you point at the step that failed when the count and the ecology disagree.
1. Define the ecological estimand
Choose presence, occupancy, activity, call rate, species composition, or another target.
2. Design the sensor network
Control placement, schedule, calibration, weather, failure, and spatial coverage.
3. Build detection and review
Use species models, unknown handling, human verification, and active learning.
4. Model observation bias
Separate ecological change from detectability, effort, hardware, and environment.
Observation bias is modeled against the ecological estimand you named earlier, so a badly chosen estimand is corrected for rather than reconsidered.
Example
Independent sites, independent seasons
Event precision and recall under independent sites and seasons is the right figure to ask for first. What it cannot do on its own is tell you whether the detector is keying on background noise or on the recorder itself. A detector that has learned the hardware takes that advantage with it into every held-out site. The Bird Audio Detection challenge showed the cheap countermeasure: score per site instead of pooling. Pooled, one system finished second out of thirty teams. Scored on average per-site AUC, the same 'cakir' system would have ranked tenth.
So report what the number counts and which sites and seasons it covers. Report how uncertain it is, and what the weather and the recorders were doing while it was measured. Report it alongside occupancy or activity estimates that carry detection uncertainty. “Site occupancy and detection probabilities were 0.43 and 0.50, respectively, for the common but declining species (the spotted owl)”, as the owl study put it, is what a reportable pair looks like: the ecological quantity and the observation quantity, side by side, neither standing in for the other.
- For core task evidence, event precision and recall under independent sites and seasons — scored site by site, since pooling moved one challenge system from second place to tenth.
- For system behavior, occupancy or activity estimates that carry their detection uncertainty instead of burying it in a point value, as the paired 0.43/0.50 and 0.09/0.67 figures do.
- For a robustness slice, sensor uptime, calibration, weather, and coverage — the conditions most likely to have moved between the seasons being compared, alongside the recorder models themselves, worth a factor of 1.6 in bird activity on their own.
- For lifecycle evidence, human review yield, the unknown-event rate, and whether ecological validation happened at all — BirdNET's 0.251 average correlation with hotspot observation is what that validation looks like when someone runs it.
Report event precision and recall under independent sites and seasons together with human review yield, unknown-event rate, and ecological validation.
Steps
Design a defensible monitoring study
A monitoring study is defensible when another team can take it apart and still not find a shortcut in the background noise or the recorders. The artifact that makes that possible is three lines long. Write down what defining the ecological estimand assumes. Write down one counterexample. Write down what modeling observation bias then requires.
For a project comparing counts across seasons, those three lines are short. The assumption is that detections are comparable from one season to the next. The counterexample is 12 microphone models tested in the field, across which sampled bird vocalisation activity varied by a factor of 1.6 and bat activity by 9.7. The requirement is that hardware sensitivity and sensor placement enter the observation model as terms, so that a change in either is corrected for rather than read as birds.
The last decision is not statistical at all. It is about who else reads the coordinates. “Poaching has been documented in species within months of their taxonomic description in journals”, Lindenmayer and Scheele wrote in Science in 2017, and “over 20 newly described reptile species have been targeted in this way, potentially leading to extinction in the wild.” Their answer is a risk-tiered rule rather than a blanket one: withhold location data outright for high-value species, buffer spatial data to broad locations at moderate risk, and publish unrestricted only where the risk of perverse outcomes is low. A detection log for a rare species is a coordinate list. Which of those three tiers it falls into is a decision the study makes on purpose or by accident.
1. Choose the ecological question
State the population, site, season, and inference unit.
2. Randomize or stratify sensors
Avoid confounding habitat and hardware.
3. Hold out geography and time
Evaluate transfer to new sites, seasons, and devices.
4. Protect locations
Limit access to sensitive coordinates and derived species evidence.
Decide per species which tier the coordinates fall in — withhold, buffer, or publish unrestricted — before the first detection log leaves the team.
Key takeaways
- Whatever a bioacoustic system is doing — detecting, classifying, localizing, or summarizing animal and environmental sounds — what comes out the far end is a detection, and a detection is not yet an observation of an animal. BirdNET's own figure against human hotspot observation was an average correlation of 0.251 across 121 species and 4 years.
- Never let a detection count stand in for an abundance estimate. Across 12 field-tested microphone models, sampled vocalisation activity moved by a factor of 1.6 for birds and 9.7 for bats with no change whatever in the animals.
- Name the ecological estimand before modelling anything, then model observation bias against it, because nondetection does not imply absence unless the probability of detection is 1. At 32 Maryland wetland sites the estimated American toad occupancy of 0.49 sat 44% above the naive observed proportion.
- Ask which question you are actually answering. Call detection, occupancy inference, and a soundscape index are related and not interchangeable, and the third is the weakest — r = 0.33, CI [0.23, 0.43] across 364 effect sizes from 34 studies.
- Assume until tested otherwise that a detector may be scoring hardware rather than animals. A baseline above 85% AUC in matched conditions fell to 79% under mismatched conditions, and pooled scoring hid a system that per-site scoring ranked tenth.
- Ship no ecological claim on event precision and recall alone. Pair it with human review yield, unknown-event rate, and ecological validation, and price the study honestly — a 2% annual decline in spotted owl occupancy needed 1,000 sites surveyed three times a season to detect within 10 years at more than 80% power.