Research
ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering
Overview Research area: Natural Language Processing / Music Information Retrieval — specifically retrieval-augmented generation (RAG) and domain-specific question answering for music. Technical level:

- arXiv
- 2512.05430
- Published
- 2025-12-05
- Authors
- Daeyong Kwon, SeungHeon Doh, Juhan Nam
AI summary
Overview
Research area: Natural Language Processing / Music Information Retrieval — specifically retrieval-augmented generation (RAG) and domain-specific question answering for music.
Technical level: Intermediate. The paper is readable without deep music-theory background, but familiarity with RAG pipelines, BM25 retrieval, rerankers, and LoRA fine-tuning helps.
One-sentence scope: The paper releases a music-specific vector database (MusWikiDB) and an artist-centric multiple-choice benchmark (ArtistMus), then uses them to measure how retrieval, reranking, and RAG-style fine-tuning change LLM accuracy on music questions.
What This Paper Is About
General-purpose LLMs have thin music knowledge in their pretraining data, so they hallucinate or give generic answers when asked about artists, genres, and music history. At the same time, no existing music benchmark focuses on artist metadata — discographies, career evolution, collaborations — which is what everyday listeners actually ask about. The authors build two resources (a music-specialized retrieval database and a globally diverse artist benchmark) and systematically test whether retrieval grounding can close the gap between open-source and proprietary models on music QA.
Key Contributions
- MusWikiDB — a music-specific vector database of 3.2M passages drawn from 144K music-related Wikipedia pages, spanning seven categories: artists, genres, instruments, history, technology, theory, and forms. Passages are segmented to 256 tokens with 10% overlap and indexed with BM25.
- ArtistMus — a benchmark of 1,000 validated multiple-choice questions covering 500 globally diverse artists from 163 countries, annotated with metadata including genre, debut year, country, topic category, and the supporting text passage.
- Systematic RAG evaluation — the first application of RAG-based evaluation to the music domain across closed-source, open-source, and music-specific models under three settings (zero-shot, RAG with BM25, and RAG plus BGE reranker).
- Comparison against general Wikipedia retrieval — showing that a domain-specialized corpus yields higher accuracy and faster retrieval than the general Wikipedia corpus.
Main Findings
- Retrieval dramatically lifts factual accuracy. On the ArtistMus factual questions, Llama 3 8B rose from 39.0% zero-shot to 85.4% with RAG (+46.4 pp), and Qwen3 8B rose from 35.0% to 84.8% (+49.8 pp). The abstract reports open-source gains of up to +56.8 pp (Qwen3 8B: 35.0 → 91.8).
- Reranking adds more on top. Adding the BGE reranker yielded gains of up to +7.6 pp compared to RAG alone. Best open-source results reached 92.2% factual (Llama 3 8B) and 93.0% contextual (Llama 3 3B).
- Factual questions are harder than contextual ones. In the zero-shot setting, GPT-4o outperformed Llama 3.1 8B by 28.4 pp on factual questions, but only 8.2 pp on contextual questions.
- Retrieval narrows the open-source / proprietary gap. With retrieval support, smaller open models approached much larger API-based models. GPT-4o still led overall at 95.4% factual / 94.0% contextual with reranking.
- RAG helps facts more than reasoning. Average factual gains from RAG were +40.4 pp, versus +4.8 pp for contextual questions — suggesting retrieval supplies facts while model parameters handle reasoning.
- The music-specific model underperformed. ChatMusician (7B) scored around 26% factual accuracy in zero-shot, close to random guessing on a four-option task, which the authors attribute largely to instruction-following failures. It reached only 36.4% factual / 47.4% contextual with reranking.
- Standard QA fine-tuning can hurt. Fine-tuning Llama 3.1 8B Instruct on 8K multiple-choice QA pairs raised factual accuracy only slightly (39.0% → 40.8%, +1.8 pp) but dropped contextual accuracy from 84.8% to 75.6% (a 9.2 pp decline), suggesting overfitting.
- RAG-style fine-tuning is the strongest strategy. Training on 8K (context, question, answer) triples lifted contextual accuracy to 93.0% and, combined with the reranker, produced the best overall ArtistMus score of 93.1% (92.2% factual / 94.0% contextual).
- Gains transfer out of domain. On TrustMus (built from a different source, The Grove Dictionary Online), Llama 3 1B improved from 25.0% to 36.0% with RAG and 38.0% with reranking; every model tested showed improvement (+6% to +13%).
- Music-specific retrieval beats general Wikipedia retrieval. MusWikiDB delivered a +6.0 pp absolute accuracy improvement and 40% faster retrieval than the general Wikipedia corpus.
- BM25 plus reranking at 256 tokens was the best retriever configuration. Dense Contriever degraded badly on factual questions as passage size grew (80.8% at 128 tokens down to 54.6% at 512 tokens), while BM25 stayed stable; the best single result was BM25 + reranker at 256 tokens (92.2% factual / 91.2% contextual).
- Scale matters less once retrieval is in place. With RAG, Qwen3 0.6B and Qwen3 8B averaged 77.4% vs 87.4%, a much smaller gap than without retrieval.
Methodology in Plain English
Building the database. The authors started from seven representative Wikipedia pages (one per category: artists, genres, instruments, history, technology, theory, forms) and followed hyperlinks outward three layers deep (depth 1, 2, and 3). They split each page into Wikipedia sections, discarded sections shorter than 60 tokens, and cut the rest into passages of up to 256 tokens with 10% overlap. They indexed everything with BM25 for fast keyword-based retrieval, and used a BGE reranker as a second stage.
Building the benchmark. From the crawled data they kept artists whose Wikipedia infoboxes contained both a genre and a year. They ranked artist page sections by frequency and selected the five most common topics: biography, career, discography, artistry, and collaborations, keeping passages between 500 and 2,000 tokens. Genres were normalized (lowercased, spaces/hyphens/slashes removed), starting from 48 root genres and mapping the 300 most frequent to 20 final labels. Country was extracted by prompting Llama 3.1 8B Instruct with each page's first paragraph, validated against the pycountry library. They chose 500 artists using an inverse-frequency strategy that prioritized underrepresented countries, producing coverage of 163 countries with roughly 43% of artists from outside the U.S. and Europe.
Writing and validating questions. GPT-4o generated one factual question and one contextual question per artist from the corresponding section text. Questions passed a two-stage validation checking Music Relevance and Faithfulness (whether the answer could be derived from the given text), followed by human validation, yielding 1,000 multiple-choice questions. Correct-answer positions were balanced so each of the four options is correct 250 times.
Experiments. All evaluations used multiple-choice accuracy with temperature 0. Models tested in the main experiments: GPT-4o, Gemini 2.5 Flash, ChatMusician, Llama 3 (1B/3B/8B), and Qwen3 (0.6B/1.7B/4B/8B, all with thinking mode disabled). Retrieval used top-4 passages of 256 tokens for RAG, and the BGE reranker for the re-ranked setting. For the retriever configuration study, the token budget was fixed at 1024 by varying passage counts (top-8 at 128 tokens, top-4 at 256, top-2 at 512). Fine-tuning used LoRA with 8-bit quantization for one epoch, batch size 2, gradient accumulation 4, learning rate 3e-5, weight decay 0.005, warmup ratio 0.1, cosine schedule, AdamW, LoRA rank 16, alpha 16, and dropout 0.1.
Why This Matters
Impact on research. The paper provides the first retrieval resource and artist-centric evaluation suite purpose-built for music QA. It isolates why RAG helps (factual grounding) from why it helps less (reasoning), and shows that RAG-style training on context-question-answer triples outperforms conventional QA fine-tuning. The finding that a 0.14M-page domain corpus beats a 3.2M-page general corpus on both accuracy and speed is a strong argument for domain-specialized retrieval.
Real-world applications:
- Music streaming and discovery assistants — answering listener questions about artist background, career evolution, and collaborations without hallucinating.
- Editorial and music journalism tools — grounding draft articles or fact-checks in a verifiable music knowledge base.
- Music education and cultural heritage platforms — supporting questions about genres, instruments, and regional traditions across 163 countries, including underrepresented regions.
- Recommendation and catalog enrichment — attaching accurate, source-linked metadata to artist profiles as new releases appear, without retraining the underlying model.
Industry relevance. The results suggest that smaller, cheaper open models combined with a curated retrieval index can reach accuracy comparable to large proprietary APIs for music QA. Because RAG requires no retraining to reflect new releases or career milestones, the approach fits the constantly changing nature of music data, and the 40% retrieval speed advantage is relevant for real-time interactive applications.
Future Directions
- Expand the benchmark. The 1,000-question scale is described as sufficient for establishing trends but not for fine-grained analysis across subgenres and eras; open-ended, comparative, and multi-hop questions remain unaddressed because of the multiple-choice format.
- Broaden the knowledge source. MusWikiDB relies exclusively on Wikipedia, and the hyperlink-depth crawling may leave coverage gaps for niche genres, regional traditions, and recently emerging artists. Integrating structured music databases is suggested as future work.
- Evaluate on more benchmarks. The authors state that results on MusicTheoryBench and ZIQI-Eval will be included in the camera-ready version, so those comparisons are not yet reported here.
- Address bias and reliability. The ethics statement notes that Wikipedia's crowdsourced biases around non-Western music persist even in a 163-country benchmark, and that RAG can still produce incorrect information, especially for less-documented artists — raising the question of how to audit and mitigate this.
Target Audience
Researchers and engineers working on retrieval-augmented generation, domain-specific QA, and music information retrieval; practitioners building music or cultural-knowledge assistants who need a ready-made retrieval index and evaluation set; and computational musicology researchers interested in how LLMs handle artist-centric, globally diverse knowledge. Readers primarily interested in music audio or symbolic music generation will find this less relevant, since the benchmark is text-only and metadata-driven.
Authors’ abstract
Recent advances in large language models (LLMs) have transformed open-domain question answering, yet their effectiveness in music-related reasoning remains limited due to sparse music knowledge in pretraining data. While music information retrieval and computational musicology have explored structured and multimodal understanding, few resources support factual and contextual music question answering (MQA) grounded in artist metadata or historical context. We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic. These resources enable systematic evaluation of retrieval-augmented generation (RAG) for MQA. Experiments show that RAG markedly improves factual accuracy; open-source models gain up to +56.8 percentage points (for example, Qwen3 8B improves from 35.0 to 91.8), approaching proprietary model performance. RAG-style fine-tuning further boosts both factual recall and contextual reasoning, improving results on both in-domain and out-of-domain benchmarks. MusWikiDB also yields approximately 6 percentage points higher accuracy and 40% faster retrieval than a general-purpose Wikipedia corpus. We release MusWikiDB and ArtistMus to advance research in music information retrieval and domain-specific question answering, establishing a foundation for retrieval-augmented reasoning in culturally rich domains such as music.