Skip to content
AI.info

Research

More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers

Overview Research area: Computational-resource equity and scholarly impact in natural language processing (NLP), combining bibliometrics, scientometrics, and hardware-resource auditing of published pa

arXiv
2608.21806
Published
2026-08-22
Authors
Shuai Chen, Tong Bao, Jitong Peng, Chengzhi Zhang

AI summary

Overview

  • Research area: Computational-resource equity and scholarly impact in natural language processing (NLP), combining bibliometrics, scientometrics, and hardware-resource auditing of published papers.
  • Technical level: Intermediate. The paper uses regression models, field-normalized citation percentiles, and concentration statistics, but the framing and conclusions are accessible to readers without a statistics background.
  • Scope: An analysis of reported GPU resources in 13,921 ACL, EMNLP, and NAACL main-conference papers published between 2020 and 2025, testing whether greater reported GPU capability aligns with citations and paper awards.

What This Paper Is About

As language models have scaled up, access to GPUs has become a central ingredient in NLP research, and concerns have grown that researchers without large hardware budgets may be structurally disadvantaged. This paper asks a direct empirical question: are papers that report more computational resources actually more likely to be cited or to win awards? To answer it, the authors extract reported GPU models and counts from the full texts of over thirteen thousand leading NLP conference papers, convert them into a comparable hardware-capability measure, and compare how computational resources and scholarly impact are distributed.

Key Contributions

  1. A validated paper-level dataset of reported GPU configurations. The authors construct and validate a dataset covering reported GPU setups from 13,921 ACL, EMNLP, and NAACL main-conference papers, with human annotation of 400 papers and an LLM-based extraction pipeline evaluated against those annotations.
  2. A characterization of how reported GPU capability is distributed. They document temporal, topical, and institutional patterns, and show that the concentration of reported GPU capability substantially exceeds the concentration of citations and awards.
  3. Field-normalized, covariate-adjusted estimates of the compute–impact relationship. These estimates reveal limited incremental explanatory power, with a tenfold increase in aggregate reported GPU capability adding only 0.0042 to the model R² for the primary citation-percentile outcome.
  4. A decomposition of aggregate capability into GPU count and hardware generation. This separation shows the two dimensions relate differently to citations and awards, with GPU count showing the more consistent associations.

Main Findings

  • GPU reporting rose but stayed incomplete. The share of papers reporting at least one standardized GPU model increased from approximately 30% in 2020 to 57% in 2025, while the share reporting both a model and a count rose from approximately 15% to 49%.

  • Reported capability grew through newer hardware and medium-scale multi-GPU setups. The share of papers reporting one or two GPUs declined from 49.5% in 2020 to 35.9% in 2025, and configurations with nine or more GPUs accounted for 11.7% of papers with reported GPU counts in 2025. V100-class hardware was most common in earlier years, A100-class GPUs became dominant from 2023 onward, and H100-class GPUs began to appear in 2024 and 2025 without becoming dominant.

  • Median reported capacity rose sharply. Median paper-level reported GPU capacity increased from approximately 91 TFLOP/s in 2020 to 1,248 TFLOP/s in 2025.

  • Capability concentration far exceeded citation and award concentration. Between 2020 and 2023, the annual top 20% of GPU-quantifiable papers accounted for 83.9%–89.9% of total reported GPU capability, but only 27%–32% of citations and 20%–33% of awards.

  • The overlap between high-capability and high-impact papers was positive but limited. Of high-capability papers (top 20% by reported GPU capability), 14.5% fell into the citation top 10%, compared with 9.1% among other papers, a 1.59× higher high-impact rate. Most high-capability papers (85.5%) were not highly cited, and most highly cited papers lay outside the high-capability group.

  • Threshold choice did not overturn the pattern. Across GPU-capability cutoffs at the top 10%, 20%, and 30% and citation cutoffs at the top 5%, 10%, and 20%, the high-citation rate among high-capability papers was 1.49–2.14 times that among lower-capability papers.

  • Adjusted associations were positive but explained little. A tenfold increase in aggregate reported GPU capability was associated with a 3.52-percentage-point increase in the primary NLP topic–year citation percentile (95% CI [1.27, 5.77], p = 0.002), but increased model R² by only 0.0042. The OpenAlex field-normalized estimate was 1.26 percentage points (95% CI [-0.46, 2.98], p = 0.151), with an incremental R² of 0.0009.

  • Count-based and high-citation outcomes were also positive. A tenfold increase in reported GPU capability was associated with 18.8% higher log(1 + citations), a 61.2% increase in expected citation count under PPML, and a 3.84-percentage-point higher probability of belonging to the citation top 10%. The aggregate association with awards was 0.86 percentage points and did not reach the conventional 0.05 threshold (p = 0.056).

  • GPU count behaved differently from hardware generation. In a joint specification, a tenfold increase in GPU count was associated with a 4.57-percentage-point difference in the primary percentile (95% CI [1.96, 7.18], p < 0.001), while Ampere-or-newer hardware was associated with a 4.16-percentage-point difference (95% CI [1.51, 6.81], p = 0.002). Across count-based and high-citation specifications, GPU count was the more consistent predictor; newer hardware generation was associated with citation intensity but not with top-10% citation status.

  • Award evidence was dimension-specific. A tenfold increase in GPU count was associated with a 1.33-percentage-point higher award probability in the linear probability model, while the hardware-generation coefficient was close to zero. In a Firth rare-event model, a tenfold increase in GPU count was associated with 1.71 times the odds of receiving an award (95% CI [1.19, 2.44], Holm-adjusted p = 0.0085), whereas hardware generation remained statistically unsupported.

  • Richer controls further weakened the estimate. Using a common sample of 2,077 papers, the baseline association with the NLP topic–year percentile was 3.13 percentage points (95% CI [0.81, 5.45], p = 0.008). Adding pre-publication author citation history, team publication experience, institutional citation visibility, and collaboration structure reduced it to 2.74 percentage points (95% CI [0.37, 5.12], p = 0.024), and a further specification including public-artifact availability yielded 2.65 percentage points (95% CI [0.28, 5.02], p = 0.028). The incremental R² attributable to reported GPU capability declined from 0.0032 to 0.0023 and 0.0022 across these specifications.

  • Every joint linear model remained low in explanatory power. Incremental R² stayed below 0.01 for every reported linear model under the joint specification.

  • Findings-track replication matched the main track. The citation analyses replicated on Findings papers showed similar results across tracks, with no significant slope differences and similarly modest incremental explanatory power.

Methodology in Plain English

The authors gathered every ACL, EMNLP, and NAACL main-conference paper from 2020 through 2025 with an accessible PDF, ending with 13,921 papers. They parsed each full text with a tool called MinerU, then linked each paper to bibliographic metadata (authors, affiliations, citations) from OpenAlex using its DOI, plus official conference award records and topic labels.

To identify GPU usage, they first manually annotated 400 papers to create a human-validated evaluation set, with two annotators independently labeling 120 overlapping papers. Agreement was high: Cohen's kappa of 0.94 for identifying valid GPU-resource evidence, with exact match rates of 90.83% for GPU model and 87.50% for GPU count. They then tested an LLM (DeepSeek-v3.2) against this set; it achieved a GPU-name F1 of 0.933 and an exact model-and-count F1 of 0.879, which supported applying the pipeline to the full corpus.

Because GPU mentions varied widely in form, the authors normalized raw names to a standard hardware catalog and linked each model to memory capacity, GPU family, generation, and theoretical peak Tensor FP16/BF16 throughput. Specifications came from the Epoch AI Machine Learning Hardware dataset, vendor sources, and manual verification. A paper's reported GPU capacity was defined as its largest observable configuration: the maximum over reported models of the number of GPUs of that model multiplied by that model's per-card peak throughput. Papers reporting multiple configurations were not summed, to avoid double counting.

Each paper was assigned one primary topic from a closed taxonomy of 29 categories adapted from ACL Rolling Review area keywords. Two analysis samples were defined: a "model-reported" sample of 6,900 papers (49.6%) reporting at least one standardized GPU model (with count set to one when unstated, as a conservative lower bound), and a "strict" sample of 5,360 papers (38.5%) reporting both a model and an explicit count.

For the impact analysis, the primary outcome was the citation percentile within NLP topic–year cells, with citation models restricted to 2020–2023 (N = 2,194) to reduce citation-window truncation, and award models using N = 5,357. All models included publication-year-by-venue fixed effects, primary-topic fixed effects, team-size controls, and organization-count controls. Complementary outcomes included the OpenAlex field-normalized percentile, log(1 + citations), raw citation counts, top-10% citation status, and awards. The authors explicitly describe the estimates as conditional associations, not causal effects.

Why This Matters

Impact on research. The paper reframes a widely assumed relationship. Rather than confirming that compute buys influence, it shows that the top 20% of papers by reported GPU capability held 83.9%–89.9% of reported GPU capability but only 27%–32% of citations and 20%–33% of awards. Reported GPU capability is described as neither necessary nor sufficient for high citation impact, and it adds little explanatory power beyond observable publication, topical, team, and institutional characteristics. The paper also argues that resource concentration and impact concentration are not equivalent, so scholarly contribution should not be inferred from hardware scale.

Real-world applications:

  • Research evaluation and hiring. Institutions and funders can use the finding that hardware scale explains little citation variation as a caution against treating large reported GPU budgets as a proxy for research quality or likely influence.
  • Infrastructure policy. Universities and national funders weighing investment in shared GPU clusters get evidence that broadening access supports participation, while the link between compute and eventual influence is weak and uneven.
  • Reproducibility and reporting standards. The paper documents incomplete GPU reporting and recommends that venues and authors report GPU models, counts, runtime, utilization, and externally provided compute, which would let future work separate available capability from actual consumption.
  • Allocation of reviewer and program-committee attention. Conference organizers assessing submissions on large-model topics can see that high reported capability clusters in specific topics such as LLM agents, code models, and language modeling, and in industry-involved research.

Industry relevance. The finding that industry and industry–academia collaborations reported higher median capacity and were more frequently represented in the annual high-capacity tail is directly relevant to discussions of the industry–academia compute divide. The paper also notes that growing use of LLM APIs shifts the constraint from researcher-owned GPUs to platform-mediated access, meaning API-only studies may depend on substantial upstream compute that is invisible to hardware-based measures. This makes platform transparency and access to models, interfaces, and budgets an emerging axis of resource inequality alongside GPU ownership.

Future Directions

  • Measure consumption rather than reported capability. The authors state that their measure excludes GPU hours, training FLOPs, cost, energy use, and utilization. A manual audit found that only 92 of 240 GPU-reporting papers (38.3%) contained a consumption-related signal, and those signals were too heterogeneous for a comparable measure. Developing a usable consumption metric is the most direct next step.

  • Make API-mediated compute visible. Because API-only studies may rely on substantial upstream compute while reporting no local hardware, future work needs ways to account for platform-mediated model access so that compute dependence in application-oriented research is not invisible.

  • Extend the corpus and metadata scope. The current findings cover only ACL, EMNLP, and NAACL main-conference papers from 2020 to 2025 and may not generalize to workshops, journals, arXiv preprints, industrial technical reports, or other NLP and machine-learning venues. Affiliation metadata and full-counting rules also cannot identify resource ownership, researcher mobility, or access to shared and cross-national infrastructure.

  • Move from association toward causal identification. The authors emphasize that unobserved author, institutional, and project characteristics may remain correlated with both reported resources and impact, so the estimates are conditional associations. Designs that can isolate causal effects, and that address the selection introduced by incomplete reporting, remain open questions.

Target Audience

This paper is most useful to scientometrics and metascience researchers studying resource inequality and cumulative advantage in AI; to NLP researchers and program committee members curious about how compute relates to visibility in their own field; to research administrators, funders, and university infrastructure planners deciding how to invest in shared compute; and to policymakers and journalists tracking the industry–academia compute divide. Readers interested in reproducibility, hardware reporting standards, and citation-based research evaluation will also find the dataset, the reported-resource measurement framework, and the open code and data (available at the linked GitHub repository) directly relevant.

Authors’ abstract

Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact remains unclear. We analyze 13,921 ACL, EMNLP, and NAACL main-conference papers published between 2020 and 2025, using GPU resources as our operational measure of computational resources. From full texts, we extract GPU models and counts, standardize each paper's largest reported configuration into a comparable hardware-capability measure, and link these data to citation, award, topic, and institutional metadata. GPU reporting became more common but remained incomplete, while reported capability increased mainly through newer hardware generations and medium-scale multi-GPU configurations. Resource concentration substantially exceeded impact concentration: the annual top 20% of GPU-quantifiable papers accounted for 83.9%-89.9% of reported GPU capability, but only 27%-32% of citations and 20%-33% of paper awards. In adjusted models, a tenfold increase in aggregate reported GPU capability was associated with a 3.52-percentage-point increase in within-NLP topic-year citation percentile, but increased model R^2 by only 0.0042. GPU count showed more consistent positive associations with citation and award outcomes than newer hardware generation. Overall, reported GPU resources are associated with scholarly impact but provide little standalone explanation of research influence.

Read the original paper