Skip to content
AI.info

Research

Expectations and Practices around AI Disclosure in CS Research

Overview Research area: AI safety and ethics — specifically research-integrity policy, AI disclosure norms, and the sociology of generative AI use in computer science publishing. Technical level: Inte

arXiv
2608.23271
Published
2026-08-24
Authors
Arati Mohapatra, Danish Pruthi

AI summary

Overview

Research area: AI safety and ethics — specifically research-integrity policy, AI disclosure norms, and the sociology of generative AI use in computer science publishing.

Technical level: Intermediate. The paper is a mixed-methods study (policy document analysis, a human survey, and an LLM-assisted audit of disclosure statements) with a linear mixed effects model at its statistical core, but no specialized ML background is required to follow it.

Scope: An analysis of AI disclosure policies at 65 top computer science conferences, a survey of 109 computer science researchers on when disclosure is warranted, and an audit of 13,867 real disclosure statements from ICLR 2026 and EMNLP 2025.

What This Paper Is About

Publishing venues in computer science have rapidly adopted policies requiring authors to disclose generative AI use, but it is unclear whether those policies actually say enough to guide authors, and whether authors' real disclosures match what readers and reviewers want to see. The authors ask three questions: how prevalent and how well-specified these policies are, for which research tasks and levels of human involvement disclosure is considered necessary, and how well current disclosure practice aligns with those expectations. They find that policies are widespread but vague, that researchers' expectations vary enormously by task, and that real disclosures concentrate on exactly the tasks researchers consider least in need of disclosure.

Key Contributions

  1. A policy audit of 65 CS conferences and four major societies. The authors document which venues have AI disclosure policies, trace how society-level policies (AAAI, ACL, ACM, IEEE) have evolved since early 2023 using the Internet Archive, and map conference policies back to the societies they borrow from.

  2. A survey of 109 CS researchers measuring disclosure necessity. Participants rated 21 research tasks across 5 chronological research phases at 3 levels of human involvement on a 5-point Likert scale, producing task-wise and phase-wise necessity scores plus stated preferences for what a disclosure should contain.

  3. An empirical audit of 13,867 real AI disclosure statements. Using Gemini 2.5 Flash as an annotator, the authors extracted and labeled disclosures from 12,577 ICLR 2026 submissions and 1,290 EMNLP 2025 Main and Findings papers, then compared observed practice against surveyed expectations.

  4. Concrete policy recommendations, including a boilerplate template. The authors propose task-based disclosure tiers (mandatory, recommended, optional) derived from necessity scores and a standardized disclosure template capturing the details most respondents said they expect.

Main Findings

  • Policies are common but under-specified. All major publishing societies (AAAI, ACL, ACM, IEEE) have AI disclosure guidelines, and 35 out of 65 top CS conferences do. Of those 35, 29 borrow closely from society-level policies, and 21 directly link to, mention, or verbatim quote the ACM authorship policy. Society policies have undergone only 1 to 4 changes since first appearing in early 2023, with revisions being minor (restructuring content, adding reviewer guidelines or sanctions for prompt injections).

  • No society policy says what details to include. ACL and ACM give conditions for when disclosure is needed (based on novelty of generated text, or the distinction between AI assistance for research versus writing), while AAAI and IEEE remain largely open-ended, only stating that any generative AI use should be disclosed. None of the society-level policies specify what details a disclosure statement must contain.

  • Disclosure is seen as moderately necessary overall. The mean necessity rating across all tasks and conditions was 2.95 (95% CI [2.78, 3.13]).

  • Research design ranks highest; writing and reporting rank lowest. Disclosure is considered most necessary for research design tasks and least necessary for writing and reporting. Idea generation tasks were rated relatively low, at a mean of 2.86, contrary to the authors' expectation given the emphasis on novel ideas in research.

  • Task-level variation within a phase is enormous. In idea generation, "Propose new hypotheses" scored a mean of 3.54 while "Identify relevant literature" scored 2.46. In data collection, "Generate synthetic data sets" scored 4.02 while "Transcribe recordings of research material" scored 2.93.

  • Human involvement moves expectations in both directions. Low human involvement raises disclosure expectations (β = 0.49, 95% CI [0.42, 0.55], p < 0.001), while high human involvement lowers them (β = −0.44, 95% CI [−0.51, −0.38], p < 0.001). The asymmetry differs by phase: in data analysis the gap between low and assumed involvement is 1 point while the gap between assumed and high is under 0.5 points; in idea generation and research design the pattern reverses, with a 1-point drop from assumed to high involvement but only a 0.5-point rise from assumed to low.

  • Disclosure rates are moderately high. 64% of research papers submitted to ICLR 2026 and 40% of EMNLP 2025 Main and Findings papers contained disclosure statements.

  • Practice favors the tasks researchers consider least important. "Edit a research paper" was disclosed in 96.5% of ICLR and 81.4% of EMNLP disclosures, and "Create or edit software code" in 16.1% of ICLR and 23.3% of EMNLP disclosures. Higher-rated tasks appeared rarely: "Generate synthetic data sets" at 2.1% (ICLR) and 2.5% (EMNLP), and "Translate a research paper" at 1.6% (ICLR) and 2.0% (EMNLP).

  • The most expected detail is the most often missing. 71% of survey participants expect disclosures to include a statement of author responsibility, yet only 2% of EMNLP and 23% of ICLR disclosures include one — meaning 98% and 77% respectively omit it.

  • Other mismatches between expectation and practice. Nearly 50% of ICLR disclosures state tasks for which generative AI was not used, a detail participants rated least necessary, while 62% of participants expect information about human oversight. Fewer than 50% of participants wanted the specific AI model named. Expected length was "a few sentences or one short paragraph," but most disclosed statements ran 1 to 2 sentences; 91% of EMNLP disclosures were short, versus 41% of ICLR disclosures being a few sentences long.

  • Duplicate and plausibly performative disclosures exist. One identical disclosure statement appeared verbatim in 95 unique ICLR submissions, and the authors confirmed the submissions did not share authors. The AI text detector Pangram flags the statement as AI-generated.

Methodology in Plain English

The authors used three complementary methods.

First, they took the list of 65 top computer science conferences from CSRankings and checked each one's website for an AI disclosure policy, looked at the policies of the four major societies the conferences affiliate with, and used the Internet Archive's Wayback Machine to see how society policies changed over time. This work was first done January 5–19, 2026 and reviewed and revised May 18–20, 2026.

Second, they ran an online survey on LimeSurvey with 109 respondents recruited through professional networks, university mailing lists, and departmental channels using purposive and snowball sampling. Participants rated how necessary disclosure is for 21 research tasks spanning 5 phases (idea generation, research design, data collection, data analysis, and writing and reporting) at 3 levels of human involvement — "assumed" (no explicit mention), "high" (author leads, AI output closely supervised), and "low" (AI leads, loosely supervised) — on a 5-point scale. The task taxonomy was adapted from prior work. To ground the multiple-choice questions about disclosure details, the authors first examined 110 random ICLR 2026 disclosure statements, which is how the "purpose of disclosure" and "non-use of AI" options entered the survey. They modeled the ratings with a linear mixed effects model, using task and human involvement as fixed effects and participant as a random intercept, to separate genuine task effects from between-person differences. Participants were paid $10/€10/₹500 depending on location, and the study was IRB approved.

Third, they pulled accepted, withdrawn, rejected, and desk-rejected ICLR 2026 papers via the OpenReview API and EMNLP 2025 checklists and papers from the ACL Anthology, then used Gemini 2.5 Flash to detect, extract, and label disclosure statements. They validated the LLM judge against human annotation: 100% F1 on detecting disclosures across 100 papers, and average micro F1 of 96.5% for details and 90.6% for tasks across 100 disclosures annotated by three researchers.

Why This Matters

AI disclosure policies are now a fixture of CS publishing, but this paper shows they were written broadly and checked rarely, and that the disclosures authors actually write cluster around the least consequential uses of AI while omitting the details (especially responsibility statements) that readers most expect. The result is a transparency mechanism that can look like compliance without conveying much information — illustrated starkly by one disclosure text appearing verbatim in 95 unrelated ICLR submissions.

Real-world applications:

  • Conference and journal policy design. Program chairs and publication committees can use the task-based tiers (mandatory at rating ≥ 3.5, recommended at 2.5 to < 3.5, optional below 2.5) to write more specific requirements.
  • Submission systems and templates. The proposed boilerplate disclosure, embedded directly in conference LaTeX files and paired with a dedicated checklist field for mandatory tasks, gives authors a concrete default rather than a blank section.
  • Peer review and editorial triage. A standardized format makes it faster for reviewers and area chairs to see what AI did and who verified it.
  • Author self-protection. Because disclosures are made in advance, they can serve as evidence of the actual extent of AI use if a paper is later flagged by an AI text detector.

Industry relevance: AI tool vendors and enterprise research organizations that publish or sponsor research face the same disclosure-consistency problem; standardized, task-based disclosure norms give them a template to adopt internally, and the finding that authors often disclose model names only when it matters gives vendors a clearer sense of what information is actually expected from users.

Future Directions

  • Replicating the survey at other venues and in other fields. The authors explicitly encourage policymakers to reproduce their survey to generate task-wise necessity ratings for their own communities, and note that extending the characterization beyond computer science is a natural next step since disclosure expectations likely differ by discipline.

  • Reducing self-selection bias in the sample. The authors acknowledge that convenience-based recruitment likely attracted respondents who already had opinions about AI disclosure; they suggest stratified sampling across CS subdomains or randomized recruitment via venue mailing lists could yield weaker or more diffuse expectations.

  • Handling the privacy consequences of disclosure. Legitimate uses such as translation may reveal that authors are non-native English speakers, which risks reinforcing existing biases during review; the authors propose masking language-support disclosures during the review period as a partial fix.

  • Closing the expectation-practice gap. Whether task-based tiers and a boilerplate template actually change author behavior, and whether disclosures can be made non-performative, remain open questions the paper poses but does not resolve.

Target Audience

This paper is most useful to conference program chairs, publication ethics committees, journal editors, and society policy writers who set or revise AI disclosure rules. It also serves researchers studying metascience, research integrity, and the sociology of AI adoption; authors who need to write disclosure statements and want to know what readers expect; and reviewers and area chairs who interpret disclosures during peer review.

Authors’ abstract

As generative AI tools find increasing use in research workflows, ongoing debates on their impact, appropriateness and responsible use have led policymakers to enact policies to disclose AI use at multiple publishing venues. However, are current AI disclosure policies and practices reflective of their purpose? In this work, we first investigate disclosure policies of top computer science venues and find that despite their prevalence, they remain highly under-specified. Secondly, through a survey of computer science researchers (N=$109$), we characterize the necessity of disclosures across different research tasks and levels of human involvement. We learn that researchers find disclosures most necessary for tasks involving research design, and for tasks when the human involvement is low. We also compile expectations that researchers have about the information to be conveyed in AI disclosure statements. Lastly, through an analysis of $13867$ disclosure statements from EMNLP $2025$ and ICLR $2026$, we reveal a large disconnect between these expectations and AI disclosures in practice---a prime example being writing assistance which is deemed less necessary but is frequently disclosed. We conclude with recommendations to align AI disclosure policies and practices with expectations, suggesting a categorization of research tasks by perceived necessity and a boilerplate template capturing expected details.

Read the original paper