Skip to content
AI.info

Research

Fostering the Ecosystem of AI for Social Impact Requires Expanding and Strengthening Evaluation Standards

Overview Research area: AI/ML for social impact (AISI) — specifically the meta-science question of how such research is evaluated, reviewed, and incentivized at publication venues. Technical level: In

Fostering the Ecosystem of AI for Social Impact Requires Expanding and Strengthening Evaluation Standards
arXiv
2510.18238
Published
2025-10-21
Authors
Bryan Wilder, Angela Zhou

AI summary

Overview

Research area: AI/ML for social impact (AISI) — specifically the meta-science question of how such research is evaluated, reviewed, and incentivized at publication venues. Technical level: Intermediate. No new algorithms or experiments are presented, but the paper assumes familiarity with randomized controlled trials, event studies, and standard ML publication norms. Scope: A position paper arguing that the AISI research ecosystem is distorted by a narrow definition of "impact" — and proposing three concrete reforms to how contributions are recognized and how deployments are evaluated.

What This Paper Is About

There has been growing research interest in machine learning and AI for social impact, and conferences such as AAAI and IJCAI now run dedicated tracks with tailored review criteria. The authors argue that these criteria most concretely reward projects that simultaneously achieve real-world deployment and novel ML methodological innovation, which pressures researchers to force every collaboration into that same mold. The paper's goal is to widen what counts as a valid contribution and to raise the bar for how deployed systems are actually evaluated.

Key Contributions

  1. An argument against the "idealized project" framing. The paper claims the pervasive image of the ideal AISI project — engage practitioners, invent novel ML methods, deploy a system — generates harmful incentives for researchers, partner organizations, and the field, even though it remains a valuable aspiration.
  2. A taxonomy of non-method contributions. The authors enumerate scientific contributions that come from improving a partner's use of existing, standard ML tools, including how ML enters organizational decision processes, what value ML provides in a given domain, which constraints shape feasible use, and what level of problem complexity is actually warranted.
  3. A taxonomy of non-deployment methodological contributions. This covers work that changes applied researchers' minds about estimation and evaluation strategies, methodology designed for maintainability (e.g. nomograms, scorecards, categorical prioritization), and the case for exceedingly simple single-variable benchmarks.
  4. A set of rigor recommendations for three deployment types. The paper specifies best practices for pilot tests, randomized controlled trials, and non-randomized deployments, framed as both author obligations and reviewer/venue requirements.

Main Findings

  • Deployment plus methodological novelty is the single clearest path to high scores, not the only accepted work. The IJCAI Multi-Year Track on AI and Social Good evaluates papers on "contribution to state-of-the-art AI" and "collaboration with stakeholders/partners," with "potential deployment/implementation opportunities in the future" as a key desiderata. The AAAI Special Track on AI for Social Impact's review criteria prioritize papers that are deployed or close to deployment ("Scope and promise for social impact") and methodologically novel ("Novelty of approach").
  • Exclusive focus on the ideal project forecloses valuable work. The authors identify two neglected categories: projects that deploy a non-novel ML method, and projects that develop practice-motivated methods without reaching deployment.
  • Deployment is treated as a finish line, but rigorous evaluation of deployments is not consistently demanded. The authors argue this pushes researchers to fit all projects into one framework regardless of partner needs.
  • Preregistration is not required at ML/AI conferences. It is not required at machine learning or AI conferences, nor covered in checklists like the NeurIPS Paper Checklist, even though it is a publication requirement for field experiments at many journals.
  • Naive before-after comparisons can reverse the sign of the estimated effect. In the paper's illustrative figure, the organization's outcomes were already improving before deployment; compared to an extrapolated counterfactual time series, average post-deployment outcomes indicate a net negative causal effect. The authors present interrupted time series, differences-in-differences, and synthetic control as increasingly information-hungry alternatives depending on whether there is no control, one control unit, or many comparable control units.
  • Simple baselines can perform comparably to complex ML. The paper cites recent work showing that single-variable predictions have comparable performance to more complex ML approaches, and argues that strong performance of simple baselines should not be surprising or diminish the value of ML innovations.
  • Underpowered trials waste partner and participant effort. The authors state that trials which cannot detect plausible effect sizes yield noisy results and potentially waste effort from partner organizations and participants.
  • Heterogeneity analysis is prone to false positives. Detecting heterogeneous effects is more data-intensive than measuring averages and particularly susceptible to false positives due to the number of comparisons made, so strategies for detecting it should be preregistered.
  • Algorithms often function as resource allocation mechanisms rather than direct interventions. This creates a randomization dilemma: in principle the trial should randomize groups, but this may be difficult when the implementing partner consists only of a single "group." Recent strategies can, under additional assumptions, use data from a single population to estimate treatment effects on those the algorithm selects.
  • The paper reports no empirical study, dataset, or quantitative benchmark result of its own. It is a position paper; all evidence is argument, example, and citation. Effect sizes, sample sizes, and error rates are discussed as design requirements, not as reported findings.

Methodology in Plain English

This is a position paper, so there is no experiment, dataset, or model evaluation. The authors build their case through structured argument: they identify the prevailing "idealized project" template, trace where it appears (conference tracks, review criteria, educational materials), and then reason through the downstream incentives it creates for researchers and partner organizations. They ground the argument in previously published case examples, including a collaboration with the government of Greece to deploy a reinforcement learning system for targeting COVID-19 testing of incoming travelers, a collaboration with a food rescue nonprofit on a machine learning plus optimization system for volunteer dispatching, pretrial risk assessment models in criminal justice, and early sepsis detection systems in a hospital. They also explicitly address five alternative views or counterarguments and respond to each in turn. For the deployment section, they import established best practices from fields where experiments are core to the discipline, such as economics and medicine, and adapt them to algorithmic interventions. One figure illustrates how naive before-and-after comparisons can be confounded by pre-existing trends.

Why This Matters

Impact on research. The authors frame the situation as a coordination problem: many kinds of AISI contributions are underappreciated, which distorts how research is presented and evaluated. Individual researchers cannot fix this alone, but collective action by authors, reviewers, and publication venues can. They also warn that if researchers must include technical innovation in every project, they may steer collaborations toward problems likely to yield methodological novelty even when those problems are not the partner's highest-value priority — a disservice to partners investing limited time and resources. They additionally note that in some sensitive domains, deployment depends on the idiosyncratic capacities and priorities of a few specific entities and executive leaders, so luck plays a role alongside skill.

Real-world applications discussed in the paper:

  • Targeting COVID-19 testing for incoming travelers with a reinforcement learning system, developed with the government of Greece.
  • Volunteer dispatching for a food rescue nonprofit, using a machine learning plus optimization system.
  • Pretrial risk assessment models in criminal justice.
  • Early sepsis detection systems deployed in a hospital.
  • Behavioral science-informed communication changes with public sector partners, where projects deployed beyond an initial pilot were those leveraging existing projects and infrastructure rather than requiring additional resources.

Industry relevance. The paper's arguments about maintainability apply directly to how tools get adopted outside academia: partner organizations in public health and social services rarely have teams of ML PhD graduates who can maintain complex pipelines, and the simpler the tool is, the more likely it is to be adopted. The emphasis on "total cost of ownership" and on restricting classification tools to forms practitioners already use, such as nomograms, scorecards, and categorical prioritization, speaks to anyone deploying models where the operating organization is not an ML company.

Future Directions

  • Developing strategies to evaluate algorithm-mediated interventions. The authors state that improving empirical strategies to test and evaluate interventions mediated by algorithms is itself a promising area for future research, particularly where group-level randomization is infeasible or where the implementing partner is a single group.
  • Extending validity scrutiny to outcome selection in field experiments. Recent work has drawn attention to the social-scientific concept of validity in decisions like what outcome a model predicts; the authors argue this scrutiny should be extended to the choice of outcomes for field experiments.
  • Ecosystem-level changes at venues. Venues might display examples of the AISI-specific contributions the paper delineates, in the way review guidelines have expanded at major AI/ML conferences to include examples of contributions for more typical AI/ML papers. The authors also propose that venues require authors to declare whether field experiments and analyses were preregistered.
  • Normalizing simple baselines. Comparing against domain-driven single-variable baselines is described as currently disincentivized; shifting this requires changed expectations and practices on the part of reviewers, authors, and publication venues, because AISI application areas have high Bayes error, meaning low signal-to-noise regimes due to individual variations.

Target Audience

Authors, reviewers, and publication venue organizers in machine learning and AI — especially those involved in AISI tracks and societal impact review processes. It is also relevant to researchers building cross-disciplinary partnerships with nonprofits or government agencies, to partner organizations deciding whether academic collaboration is worth their limited time and resources, and to department and hiring committees that evaluate researchers whose work spans applied and methodological contributions.

Authors’ abstract

There has been increasing research interest in AI/ML for social impact, and correspondingly more publication venues have refined review criteria for practice-driven AI/ML research. However, these review guidelines tend to most concretely recognize projects that simultaneously achieve deployment and novel ML methodological innovation. We argue that this introduces incentives for researchers that undermine the sustainability of a broader research ecosystem of social impact, which benefits from projects that make contributions on single front (applied or methodological) that may better meet project partner needs. Our position is that researchers and reviewers in machine learning for social impact must simultaneously adopt: 1) a more expansive conception of social impacts beyond deployment and 2) more rigorous evaluations of the impact of deployed systems.

Read the original paper