Research
LabelBuddy: An Open Source Music and Audio Language Annotation Tagging Tool Using AI Assistance
Overview Research area: Music Information Retrieval (MIR), audio annotation tooling, and Human-in-the-Loop (HITL) machine learning. Technical level: Intermediate. The paper assumes familiarity with an

- arXiv
- 2603.04293
- Published
- 2026-03-04
- Authors
- Ioannis Prokopiou, Ioannis Sina, Agisilaos Kounelis, Pantelis Vikatos, Themos Stafylakis
AI summary
Overview
Research area: Music Information Retrieval (MIR), audio annotation tooling, and Human-in-the-Loop (HITL) machine learning.
Technical level: Intermediate. The paper assumes familiarity with annotation workflows, containerization, and audio-language models, but the system description itself is accessible.
Scope: This paper is a system description of LabelBuddy, an open-source audio annotation tool that separates its user interface from pluggable, containerized model backends so that AI models can pre-annotate audio for human verification.
What This Paper Is About
Creating datasets that capture the subjective, linguistic, and perceptual nuances of audio is slow, laborious, and currently spread across disconnected tools — waveform editors for segmentation, separate platforms for text, and separate software for subjective listening tests. LabelBuddy is proposed as a single open-source, collaborative annotation platform whose interface is decoupled from model inference, allowing managers to plug in custom Dockerized models for AI-assisted pre-annotation. The goal is to shift annotator effort from creating labels from scratch to verifying and correcting machine output, while supporting multi-user consensus and, eventually, subjective preference evaluation for reinforcement learning from human feedback (RLHF).
Key Contributions
-
Decoupled AI-assistance architecture. The interface (Django) is separated from compute-intensive inference (Docker), with models declared through declarative YAML files that specify the Docker image, input/output schema, and resource constraints. Predictions are returned through a RESTful Flask API, giving sandboxing and scalability (inference can run on remote cloud nodes such as AWS/Azure).
-
Collaborative consensus and role-based access control. Native multi-user roles — manager, annotator, reviewer — with Role-Based Access Control to prevent data leakage. Annotators and reviewers are restricted to their assigned task queues, and a review interface supports playback of specific regions with approve/reject decisions and feedback.
-
Hybrid workflow support. An architecture designed to support both region-based tagging and subjective preference aggregation, with pre-trained models listed for pre-annotation: YOHO, musicnn, PANNs, and LALMs such as Music Flamingo.
-
A reference case study for music captioning. A demonstrated workflow for building a Music Captioning Dataset, plus a stated evaluation plan for validating utility on DCASE 2024 data.
Main Findings
-
LabelBuddy is presented as uniquely combining four properties. Table 1 compares LabelBuddy against Audino, BAT, Aubio, Gecko, Prodigy, and Label Studio (CE). LabelBuddy is the only tool marked in all four columns: audio-specific, decoupled AI-assist, open source, and collaboratory consensus.
-
Most existing audio tools lack decoupled AI assistance. Audino, BAT, Aubio, and Gecko are each marked as audio-specific and open source, but not as decoupled AI-assist and not as collaboratory consensus.
-
General-purpose platforms restrict collaboration. Prodigy is marked as decoupled AI-assist but not audio-specific, not open source, and not collaboratory consensus. Label Studio (CE) is marked as decoupled AI-assist and open source, but not audio-specific and not collaboratory consensus. The paper states that general-purpose HITL platforms often restrict reviewer roles and consensus metrics to paid enterprise tiers, and lack native support for musical structures such as bars and beats.
-
Annotations are stored as JSON. The stored objects contain temporal boundaries and ontology tags, and managers can export consensus data to CSV/JSON for direct integration with ML training pipelines.
-
The captioning case study is illustrative, not benchmarked. A Music Flamingo checkpoint is wrapped in a Docker container via a YAML configuration mapping text output to a caption region. The example candidate caption returned is "A lo-fi hip-hop track with a slow tempo and vinyl crackle," which the annotator corrects (for example, changing "vinyl crackle" to "rain sounds") and adjusts timestamp boundaries. Finalized data is exported as JSONL or CSV containing aligned (audio_path, text_caption) pairs.
-
No empirical results are reported in this paper. The only quantitative evaluation described is a proposed plan, not measured outcomes: a pilot study on DCASE 2024 data measuring (a) time reduction versus de novo labeling, (b) inter-annotator agreement using Fleiss' Kappa, and (c) downstream PSDS gains for baseline SED models trained on LabelBuddy-curated data.
Methodology in Plain English
The authors approached the problem as a software engineering challenge rather than an experimental one. They identify a "coupling problem": annotation interfaces are usually hard-coded to a specific model backend, which makes them obsolete as models evolve. Their solution is a three-part split.
First, a Django web server with a relational database manages three primary entities — Projects, Users, and Tasks — and enforces role-based access control so that annotators and reviewers only see their assigned work.
Second, model inference runs in isolated Docker containers. A manager uploads a YAML file naming the Docker image, the input and output schemas, and resource needs such as GPU availability. When an annotator requests a prediction, the backend sends the audio to the container over a RESTful Flask API.
Third, the frontend renders returned predictions as editable waveform regions using wavesurfer.js, so the human's job becomes verification rather than creation. Completed tasks move to a review interface where reviewers play back regions and approve or reject annotations with feedback. The paper then walks through a captioning example to show how the pieces fit together end to end.
Why This Matters
Impact on research. The paper argues that MIR is shifting from static tag classification toward generative and reasoning-based approaches, driven by Large Audio-Language Models (LALMs) such as Music Flamingo, Qwen-Audio, and Audio Flamingo 3. It also cites a "crisis of metrics," in which objective scores such as Fréchet Audio Distance (FAD) fail to correlate with human perception, pushing the field toward RLHF and subjective evaluation. LabelBuddy is positioned as infrastructure for producing the human-aligned datasets that these approaches require.
Real-world applications:
- Building music captioning datasets that pair audio with natural language descriptions, for fine-tuning audio-to-text generation models.
- Curating sound event detection (SED) datasets, particularly temporal region labeling for events in audio.
- Running subjective listening studies — MUSHRA, GoListen, or pairwise preference testing — within the same tool used for annotation, rather than in separate standalone software such as WebMUSHRA.
- Producing consensus-based ground truth for model fine-tuning, using the review loop to resolve disagreements between annotators.
Industry relevance. The tool targets teams that need reviewer roles and consensus metrics, which the paper says general-purpose commercial platforms often restrict to paid enterprise tiers. Because inference is containerized and can run on remote cloud nodes, the same platform can serve small research groups and production data pipelines. The work was supported by Google Summer of Code (GSoC), the Open Technologies Alliance (GFOSS), and the European Union's Horizon Europe AIXPERT project (Grant Agreement No. 101214389).
Future Directions
-
Agentic and conversational reasoning. The authors plan to extend the backend API to support more reasoning-capable models such as Qwen-Audio, letting annotators query the model and receive chain-of-thought justifications, transforming the workflow from tag verification into collaborative reasoning.
-
Integrated subjective evaluation for RLHF. They aim to implement a native pairwise preference interface (Dataset A versus Dataset B) inside the review loop, and to integrate Bayesian Bradley–Terry (BBQ) models to handle noisy human raters and aggregate preferences robustly.
-
Enhancing perceptual validity. To counter models relying on text priors rather than audio content — a flaw highlighted by the RUListening benchmark — they plan timestamp-required QA templates that force both models and humans to ground every semantic claim in specific spectral regions.
-
Validating utility empirically. The proposed pilot study on DCASE 2024 data would test whether AI-assisted pre-annotation reduces labeling time, how it affects inter-annotator agreement (Fleiss' Kappa), and whether it improves downstream PSDS for SED models.
Target Audience
This paper is most useful to MIR and audio machine learning researchers who need to build labeled datasets; data curation and annotation teams looking for an open-source alternative to commercial labeling platforms; engineers who want to attach custom audio models to an annotation frontend; and researchers working on HITL pipelines, RLHF alignment, or subjective evaluation of generative music systems. Readers looking for empirical benchmark results will not find them here, since the paper presents a system description and a proposed evaluation plan.
Authors’ abstract
The advancement of Machine learning (ML), Large Audio Language Models (LALMs), and autonomous AI agents in Music Information Retrieval (MIR) necessitates a shift from static tagging to rich, human-aligned representation learning. However, the scarcity of open-source infrastructure capable of capturing the subjective nuances of audio annotation remains a critical bottleneck. This paper introduces \textbf{LabelBuddy}, an open-source collaborative auto-tagging audio annotation tool designed to bridge the gap between human intent and machine understanding. Unlike static tools, it decouples the interface from inference via containerized backends, allowing users to plug in custom models for AI-assisted pre-annotation. We describe the system architecture, which supports multi-user consensus, containerized model isolation, and a roadmap for extending agents and LALMs. Code available at https://github.com/GiannisProkopiou/gsoc2022-Label-buddy.