Research
When AI Becomes Routine: A Decade of Public AI Mediation in Korean Go Commentary
Overview Research area: AI and society / AI ethics (the paper positions itself for the AIES community), computational social science, human–AI interaction, and media studies of expertise — specificall
- arXiv
- 2607.28332
- Published
- 2026-07-30
- Authors
- Haewoon Kwak
AI summary
Overview
- Research area: AI and society / AI ethics (the paper positions itself for the AIES community), computational social science, human–AI interaction, and media studies of expertise — specifically the public mediation of machine judgment in Korean Go commentary on YouTube.
- Technical level: Intermediate. The methods are keyword-anchored text measurement and standard statistical contrasts (Welch and paired t-tests, GEE logistic regression), not new model development, so the concepts are accessible; the clustered statistical caveats require some care.
- Scope: A decade-long (2016–2025) study of roughly 1,900 hours of Korean Go commentary across institutional broadcasters and creator-led channels, tracking how AI becomes verbally routine, or recedes into the unmarked background, after AlphaGo and KataGo made superhuman analysis a standard tool.
What This Paper Is About
The paper asks what happens to public talk about AI once a superhuman system stops being a novelty and becomes ordinary expert infrastructure. Using Korean Go commentary as the case — where AI winrate graphs sit on screen throughout broadcasts — the author measures how often commentators overtly name or reference AI versus rendering its outputs through familiar interfaces and metrics. The goal is to show that how machine judgment is made publicly intelligible and attributable is a governance-relevant layer distinct from what the model outputs.
Key Contributions
- A conservative, keyword-anchored "AI-salient" measurement. The paper introduces a precision-first subset that treats explicit AI naming as a marked surface form and checks its sensitivity with alternative per-word and per-minute denominators. A stratified manual audit of 150 strict-backbone hits found 145/150 true positives (96.7%, 95% Wilson CI [92.4, 98.6]).
- A conceptual account of public AI mediation. Commentators are framed as public interpreters of AI output — intermediaries who name, translate, soften, or resist machine judgment for audiences — building on gatekeeping and explainability-in-practice research.
- A mediation-form typology. The paper distinguishes source-foregrounding from source-receding mediation, grounded in qualitative inspection of recurring micro-patterns, and argues each preserves different "hooks of contestability" — discursive anchors through which audiences can recognize and question the machine source.
- A compositional finding about verbal mediation. The strongest evidence is a shift from explicit naming toward interface rendering, with creator-led commentary leaning further toward interface rendering than institutional commentary.
Main Findings
- The AI-salient subset grows across phases. In BadukTV, AI-salient sentences rise from an effectively zero pre-AlphaGo baseline to 0.38% in Phase 2, 1.35% in Phase 3, and 2.63% in Phase 4. A complete audit of all 13 strict hits in Phase 1 found them all to be false positives from generic percentage talk.
- The visual/verbal asymmetry is wide. AI winrate graphs are visible for an average of 97.99% of Phase 4 institutional broadcast time (median 99.83%; 51 of 54 videos ≥95% visible), yet AI-salient talk accounts for only 2.63% of BadukTV Phase 4 sentences. Under a Phase 3-specific detector applied to a separate BadukTV sample (N = 6), AI-graphic visibility was a mean of 86.9% (median 88.0%; range 81.7–92.8%).
- What recedes is the source label, not the metric. Winrate and point-gap talk persists while "AI" itself goes unsaid. The paper argues this rules out a simple suppression mechanism (commentators citing AI less because the metric is always on screen, like a scoreboard).
- Composition shifts from naming to interface rendering. In Phase 2, the BadukTV subset splits into explicit-only 76.4% and interface-only 23.6% with no overlap (N = 72). By Phase 3, interface-only (55.5%) outnumbers explicit-only (40.9%), with 3.6% jointly anchored (N = 916). In Phase 4 BadukTV, the families become roughly comparable: 50.1% explicit-only / 46.3% interface-only / 3.5% jointly anchored (N = 2,098). Pooled with K-Baduk, the Phase 4 institutional sample is interface-dominated at 35.9% / 61.5% / 2.6% (N = 4,125).
- The growth is robust to alternative denominators. The strict subset rises from 0.50 to 1.43 to 2.43 per 1,000 words and from 0.041 to 0.090 to 0.188 per minute across Phases 2–4.
- Field-level volume depends on the denominator. Per sentence, Phase 4 institutional videos devote 2.38% of sentences to AI-salient talk versus 2.48% in the creator field (t = −0.69, p = 0.493). Per 1,000 words the creator field is higher (2.97 vs. 2.18; t = 3.38, p < 0.001), and per minute (0.222 vs. 0.176; t = 2.39, p = 0.018). The paper reads this as faster solo narration rather than a different rate of overt naming.
- Composition, not volume, is the firmer contrast. In Phase 4, institutional subsets are 35.9% explicit-only / 61.5% interface-only / 2.6% jointly anchored, whereas the broader creator field is 25.2% explicit-only / 71.3% interface-only / 3.5% jointly anchored.
- The composition contrast is modest and clustering-sensitive. A GEE logistic model of explicit-only vs. interface-only at the sentence level (N = 6,797) with cluster = video (172 clusters) gives OR = 0.66, 95% CI [0.48, 0.91], p = 0.010. With cluster = channel (7 clusters), the estimate has the same direction (OR = 0.66, [0.26, 1.66]) but is not significant (p = 0.37). An aggregated chi-square test ignoring clustering produces p < 10⁻¹⁶ and, the paper states, overstates certainty. Per-video Welch tests agree (explicit-only 33.2% institutional vs. 24.8% creator, t = −2.40, p = 0.016; interface-only 64.6% vs. 73.0%, t = 2.32, p = 0.020).
- Sub-type decomposition sharpens the field contrast. Within Phase 4 interface-only AI-salient sentences (N = 8,786), creator-led commentary cites winrate markers at 58.2% versus 30.5% institutional, and bluespot (AI-recommended-move) markers at 18.4% versus 11.2%. Institutional commentary references the graph itself at 59.8% versus 21.1% in the creator field. Winrate (OR = 4.66, 95% CI [2.32, 9.35], p < 10⁻⁴) and graph (OR = 0.11, [0.04, 0.35], p = 0.00013) both survive Bonferroni correction at α = 0.05/4. Bluespot is significant at the video-clustered level (p = 0.003) and nominally at the channel level (OR = 1.65, [1.04, 2.61], p = 0.033) but does not survive correction. Point-gap shows no field effect.
- Within-speaker volume is stable, composition moves slightly. For Lee Hyunwook across 48 temporal pairs with non-zero AI-salient discourse on both sides, overall AI-salient share is 2.07% personal versus 1.98% television (paired t = 0.45, p = 0.653; p = 0.854 per 1,000 words, p = 0.309 per minute). Interface-only is 70.3% personal versus 64.2% television (paired t = 2.21, p = 0.032), which the paper calls suggestive rather than confirmatory.
- Audience uptake is higher around AI-salient events. In an exploratory event-linked chat analysis (48 event–baseline pairs, 619 hand-coded messages), 20.4% of messages around AI-salient commentary events engaged AI versus 5.8% in matched baseline windows (sign test p < 0.001).
- Three recurrent mediation micro-patterns. Reportive relay (Korean reportive forms such as "~라고 하네요" and "~라고 합니다" that relay machine-backed recommendations without repeatedly naming the source), translational mediation (re-encoding machine outputs into viewer-legible game language), and human–machine contrast.
Methodology in Plain English
The author assembled three datasets from four Korean-language YouTube channels — two institutional broadcasters (BadukTV, K-Baduk) and two creator-led professional-player channels (ProYeonwoo, LeeHyunWookTV) — plus a follow-up creator set adding ChoHyeyeon, RyuSihunWorld, and DongneBaduk. Across the three datasets the study uses 614 dataset entries corresponding to 609 unique source videos.
The longitudinal backbone, D_Long, is a stratified sample of 400 high-visibility videos (~1,394 hours) drawn from more than 28,000 candidates across four phases; the selected videos have substantially higher average view counts than unselected candidates (80,147 vs. 10,596; p < 0.001). D_Case pairs 57 BadukTV television appearances by Lee Hyunwook with 57 temporally proximate streams from his personal channel. D_Creator adds 100 videos across five creator channels, balanced across 2023 and 2024 (10 videos per year within channel).
Transcripts were generated with Whisper Base and split into sentences with the Kiwi Korean morphological analyzer. A targeted 120-row STT keyword audit found acceptable transcription in 98.3% of items and 100% preservation of the AI-anchor terms that feed the keyword pipeline. The AI-salient denominators are 400 listed / 398 transcribed / 673,557 sentences after non-game filtering for D_Long, and 100 / 100 / 116,202 sentences for D_Creator; an exploratory topic analysis used SBERT and a BERTopic-style pipeline over 876,353 sentences across D_Long and D_Case.
Rather than classifying whole sentences by meaning, the author flags a sentence as AI-salient if it contains an explicit AI-reference marker (AI, 인공지능, KataGo, Jueyi) or interface/metric language (graph, bluespot, winrate terms, bounded percentages, decimal point-gap estimates such as 1.6집), excluding record/performance and advertisement contexts and requiring local anchoring for broader point-gap talk. Coding reliability was checked on a stratified 100-sentence sample (95 effective for intra-rater, 94 for inter-rater): intra-rater Cohen's κ = 0.98 for the binary classification (95% CI [0.94, 1.00]) with marker-level κ between 0.87 and 1.00; a second coder gave κ = 0.98 binary with marker-level κ between 0.42 and 0.89, the low values concentrated on one Go-context-dependent marker reported descriptively only. Bootstrap checks on the precision audit give 95% CI [0.93, 0.99] and on the false-negative audit a miss-rate 95% CI of [0.000, 0.027].
A separate OCR-based check on 54 Phase 4 institutional broadcasts measured how long winrate graphs were on screen, using per-video manually annotated regions of interest. The author states the primary coding design was single-coder and bounds what each reliability check can support, and that no field-level claim about mediation norms is supported at the channel-cluster level.
Why This Matters
Impact on research. The paper argues that accountability and contestability depend not only on what a model outputs but on how that output is made publicly intelligible and attributable after deployment. It offers a way to study AI's social integration after adoption, rather than through anticipation, attitudes, or short-run reactions, and it treats the receding source label as the communicative signature of domestication.
Real-world applications:
- Public communication of AI-assisted decisions. The typology of source-foregrounding versus source-receding mediation gives a vocabulary for evaluating whether audiences retain an anchor to question a machine source.
- Domain transfer to higher-stakes settings. The paper argues that in domains where AI is less reliable than in Go and errors carry real-world harm, source naturalization raises governance concerns that are functional-and-benign in low-stakes Go viewing.
- Media and broadcast practice. The contrast between multi-commentator institutional formats with editorial oversight and solo creator narration suggests format design shapes whether the machine source stays marked.
- Audience-facing design of AI explanations. Because viewers in the coded windows sometimes name the AI around interface-only events — re-marking a source the commentator left unmarked — the analysis speaks to how explanation can be accomplished in the communicative layer rather than by the model alone.
Industry relevance. The findings speak to any setting where AI-assisted judgment is displayed to an audience — evaluation overlays, dashboards, live analytics — and where the practical question is whether the source label survives routine use. The paper's caution is that the retreat of explicit naming is not directly evidence of a deeper change in how experts treat AI judgment: whether it reflects that, or merely a broadcast convention against narrating an on-screen graphic, is underdetermined by verbal data alone.
Future Directions
- Disentangle convention from changed expert orientation. The paper states that whether the compositional shift reflects a deeper change in how experts treat AI judgment or a broadcast convention against narrating an on-screen graphic is underdetermined by verbal data alone. Non-verbal or interview-based evidence could address this.
- Test whether mediation form affects contestability in practice. The paper reports that challenge cases in chat are sparse, so its claim that overt AI activation leaves viewers a discursive anchor for both deference and challenge is offered as illustration, not as evidence of how often the anchor is used.
- Scale the field-level and audience analyses. Field-level claims are supported at the video-clustered level but not at the channel-clustered level with only seven channels; the audience layer rests on 48 event–baseline pairs and 619 messages with single-coder categorization, both flagged as extensions rather than backbones.
- Transfer the framework to domains where AI is less reliable than in Go. The paper explicitly says the stakes of the distinction rise where AI lacks Go's near-oracle status and errors carry real-world harm.
Target Audience
Researchers in AI ethics and AI and society (AIES), human–AI interaction, explainability and contestable-AI research, and computational social science studying expertise and media; media scholars interested in mediatization, gatekeeping, and domestication of technology; and practitioners who design or narrate AI-assisted analysis for a public audience — including broadcasters, streamers, and product teams building evaluation overlays. The paper is written for readers comfortable with quantitative text measurement and regression, but the conceptual framing, typology, and qualitative micro-patterns are accessible to a general social-science audience.
Authors’ abstract
When AI systems surpass elite human performance and settle into everyday expert practice, the question that follows is how machine judgment is made publicly intelligible and attributable. We study Korean Go commentary on YouTube, where AI systems such as KataGo became standard analytic tools after AlphaGo. Our corpus spans a decade (2016--2025) and approximately $1{,}900$ hours of footage across institutional broadcasters and creator-led channels, in four phases of AI availability. We document a widening asymmetry between visual and verbal AI presence: AI winrate graphs are visible for about $98\%$ of late-period institutional broadcast time, yet AI-salient talk accounts for only $2.63\%$ of sentences. What recedes is the source label, not the metric: winrate and point-gap talk persists while ``AI'' itself goes unsaid. We read this recession as the communicative signature of domestication. Our strongest evidence is a compositional shift in verbal mediation: explicit naming gives way to interface rendering, and creator-led commentary leans further toward it than institutional commentary. We develop a typology distinguishing source-foregrounding from source-receding mediation, and argue that the two preserve different hooks of contestability: discursive anchors through which audiences can recognize and question the machine source. The stakes of that difference rise in domains where AI is less reliable than in Go.