AI literacy basics
Evaluating AI Products and Claims
Learn how to inspect AI claims using baselines, representative evaluation, pilots, documentation, costs, failure evidence, and production outcomes.
By the end you can
- Distinguish demos, benchmarks, pilots, and production evidence
- Evaluate whether a baseline and test population match the claim
- Identify missing information about cost, failure, human work, and system boundaries
- Conduct a structured due-diligence review of an AI product
Example
Four claims, four very different evidence burdens
The stronger and broader the claim, the more demanding the evidence should become.
The measurement claim is the one that most often hides its own weakness. Hugh Zhang and fourteen colleagues at Scale AI published “A Careful Examination of Large Language Model Performance on Grade School Arithmetic”. That was in 2024. They commissioned GSM1k, a fresh set of grade-school arithmetic problems. It was built to mirror the style and complexity of the widely quoted GSM8k benchmark. They re-ran leading open- and closed-source models on it. They report “accuracy drops of up to 8%, with several families of models showing evidence of systematic overfitting across almost all model sizes”. They also report a positive relationship, with Spearman’s r^2 = 0.36. The more likely a model was to reproduce a GSM8k example, the further its score fell on the new problems. Their conclusion is deliberately two-sided: “many models, especially those on the frontier, show minimal signs of overfitting”. A benchmark score is a claim about one dataset. The dataset may already be inside the model.
- “The model can summarize these ten sample reports.” This is a narrow demonstration claim.
- “The model scores 92 percent on our benchmark.” This is a measurement claim tied to one dataset and metric.
- “Agents resolve tickets faster in a controlled pilot.” This is a workflow claim requiring a credible comparison.
- “The system improves support quality at scale.” This is a production claim involving adoption, cost, failures, and downstream outcomes.
- “The assistant understands customers like an expert.” This is an anthropomorphic claim that must be translated into observable tasks.
The evidence ladder: each rung answers a different question
A demo shows that something can happen. A benchmark estimates performance on a defined dataset. A controlled pilot tests a workflow with selected users, while production evidence measures behavior under real scale, incentives, drift, incidents, and costs.
Higher rungs do not make lower rungs useless. They answer broader questions and expose failure modes that laboratory evidence cannot reveal.
Medical imaging shows how thin the upper rungs can be. The claims still sit at the top. Myura Nagendran and colleagues published a systematic review in the BMJ on 25 March 2020. It compared diagnostic deep-learning algorithms with expert clinicians. “Only 10 records were found for deep learning randomised clinical trials, two of which have been published”. Of the 81 non-randomised studies identified, “only nine were prospective”. Just six “were tested in a real world clinical setting.” The comparison group was thin too. “The median number of experts in the comparator group was only four (interquartile range 2-9).” Yet “61 of 81 studies stated in their abstract that performance of artificial intelligence was at least comparable to (or better than) that of clinicians”. And “the overall risk of bias was high in 58 of 81 studies”.
Figure
Match the scope of the conclusion to the scope of the evidence.
Comparison
Do not ask one kind of evidence to do another kind of work
Product evaluation becomes clearer when each evidence type is treated as a bounded instrument.
Demo
Shows selected capability in a curated interaction.
- Fast and vivid
- Easy to cherry-pick
- Weak on frequency and edge cases
- Question: can this happen?
Benchmark
Measures defined tasks on a dataset with a metric and protocol.
- Supports comparison
- Depends on representativeness
- Can saturate or be contaminated
- Question: how does it perform here?
Pilot
Tests the system in a bounded real workflow against a baseline.
- Reveals usability and human work
- Needs clear success and stop criteria
- May still exclude hard conditions
- Question: does it improve this process?
Production evidence
Observes sustained operation across users, shifts, incidents, costs, and outcomes.
- Most relevant to service claims
- Harder to attribute causally
- Requires monitoring and governance
- Question: does value persist at scale?
Key idea
A result without a baseline may be a mirage
“Eighty percent accurate” sounds informative until you learn that a trivial rule reaches eighty-five. “Cuts drafting time by half” means something else once you learn that users spend the saved time correcting factual errors. Both numbers were true.
The baseline should be the actual alternative — the current human workflow, the existing software, a simple model, or no action. Anything weaker inflates the novelty. It also hides whether the system earns its complexity.
The mirage has been measured. Maurizio Ferrari Dacrema, Paolo Cremonesi and Dietmar Jannach took 18 neural top-n recommendation algorithms and tried to rebuild them. Those algorithms “were presented at top-level research conferences in the last years”. Their paper, “Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches”, was the Best Long Paper at the ACM Conference on Recommender Systems in 2019. “Only 7 of them could be reproduced with reasonable effort. For these methods, it however turned out that 6 of them can often be outperformed with comparably simple heuristic methods, e.g., based on nearest-neighbor or graph-based techniques.” The seventh beat the heuristics. But it “did not consistently outperform a well-tuned non-neural linear ranking method”. What the authors put in the dock is not fabricated numbers. It is “the choice of the baselines when proposing new models”.
A product claim is incomplete until you know what would have happened without the product.
Position
A number without a baseline is marketing, and should be read as such
A performance figure with no stated alternative is not evidence, and the polite convention of discussing it as though it were is how weak systems get bought. The question is never whether eighty percent is good. It is: eighty percent against what, measured by whom, on which population — and what does the workflow you already run score on the same test?
This is not scepticism as a pose. The recommender analysis in this lesson found tuned simple baselines matching or beating published neural methods. The medical-imaging review found that of eighty-one non-randomised studies, six had been tested in a real clinical setting and the median comparison group was four experts — while most of the abstracts claimed performance at least comparable to a clinician. In both cases the headline number was real and the evidence under it was thinner than the claim, which is the more common failure by a wide margin. Until the seller names the alternative, you have a claim about a system. You do not yet have a reason to buy it.
“Eighty percent accurate” is not a result. It is half of one.
Visual
A claim audit from sentence to decision
The audit treats every claim as a proposition with a scope, evidence base, comparison, and cost of being wrong.
- 1
Parse the claim
Identify the task, population, condition, metric, time horizon, and promised outcome.
- 2
Inspect the evidence
Check data provenance, evaluation protocol, sample size, slices, and whether the system saw similar cases during development.
- 3
Check the baseline
Compare with current practice, simple alternatives, and the cost of doing nothing.
- 4
Find missing work
Look for review, correction, escalation, integration, data preparation, and incident response omitted from the headline.
- 5
Find the failure boundary
Ask where performance degrades, which users are excluded, and what happens after a mistake.
- 6
Make a bounded decision
Accept, reject, pilot, narrow, request evidence, or postpone with explicit reasons.
Case
Where a regulator found an entire workforce
Step four, “find missing work”, is where a regulator once found an entire workforce. In an order dated 14 January 2025 (Securities Act Release No. 11352), the US Securities and Exchange Commission settled charges against Presto Automation over its drive-thru voice-ordering product, Presto Voice. Investor presentations had reported “automated order completion” rates “of 95% to 99%” and “non-intervention” rates “greater than 95%”. The order finds that those rates measured orders completed without restaurant staff involvement, not without human involvement: Presto “hired, trained, and supervised human order takers located abroad (primarily in the Philippines and India), who processed the vast majority of drive-thru orders placed through Presto Voice.” Only in a prospectus supplement filed on 17 November 2023 did the company disclose that “over 70% of orders taken by our Presto Voice solution require human agent intervention”. The headline metric was true to its own definition and useless as a description of the workflow.
Steps
Questions a serious buyer or product team should ask
The questions are useful whether the system is built internally, purchased from a vendor, or accessed through an API.
- 1
Intended and excluded uses
What tasks, users, languages, environments, and decision rights were evaluated or explicitly ruled out?
- 2
Data and model lineage
What sources, versions, adaptation, documentation, and dependencies shape the system?
- 3
Evaluation quality
Which metrics, slices, human studies, stress tests, baselines, and uncertainty estimates support the claim?
- 4
Operational behavior
What are latency, cost, availability, limits, update policy, monitoring, and rollback capabilities?
- 5
Risk and control
How are privacy, security, harmful bias, unsafe outputs, permissions, incidents, and appeals handled?
- 6
Contractual reality
Which claims are guaranteed, measurable, auditable, and supported when the system changes?
Analogy
Why software keeps changing after you buy it
A polished exterior and a smooth test drive matter when you buy a used car, and so do maintenance records, brakes, ownership history, repair cost, and behavior under conditions unlike the showroom.
AI due diligence similarly combines demonstration with provenance, stress testing, and operating evidence. A purchased car is the car you inspected. An AI service can change remotely — through model, data, policy, or vendor updates — after you sign.
That difference has been measured on a hosted service. Lingjiao Chen, Matei Zaharia and James Zou compared the March 2023 and June 2023 versions of GPT-3.5 and GPT-4. They ran seven task sets, noting that “when and how these models are updated over time is opaque”. Their study appeared in Harvard Data Science Review on 12 March 2024. On identifying prime versus composite numbers, GPT-4 (March 2023) reached 84% accuracy. GPT-4 (June 2023) reached 51%. The authors attribute the drop partly to the June model becoming less willing to follow chain-of-thought prompting. GPT-3.5 moved the other way on the same task. Their summary is the line to keep on file: “the behavior of the ‘same’ LLM service can change substantially in a relatively short amount of time, highlighting the need for continuous monitoring of LLMs.”
Inspection must include the update process, not only the version demonstrated today.
Documentation tells you what to verify
Dataset sheets, model cards, system cards, evaluation reports, and incident records can reveal intended use, limitations, subgroup behavior, provenance, and known risks. Their absence is informative, but their presence is not proof of quality.
Good documentation makes claims specific enough to challenge. Independent testing and contract terms still matter, especially when the product affects rights, safety, confidential information, or critical operations.
The model card is not a vague genre. It has authors, a date and a definition to hold vendors to. Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji and Timnit Gebru proposed the format. Their paper, “Model Cards for Model Reporting”, was presented in Atlanta on 29–31 January 2019. In their definition, model cards “provide benchmarked evaluation in a variety of conditions”. Those conditions run “across different cultural, demographic, or phenotypic groups (e.g., race, geographic location, sex, Fitzpatrick skin type) and intersectional groups (e.g., age and race, or sex and Fitzpatrick skin type) that are relevant to the intended application domains”. Model cards “also disclose the context in which models are intended to be used, details of the performance evaluation procedures, and other relevant information”. A document that reports one aggregate number and names no conditions is not a model card in that sense. That holds whatever it is titled.
Transparency improves the audit surface; it does not eliminate the need to audit.
Key takeaways
- Demos, benchmarks, pilots, and production evidence support claims of different scope.
- A result is meaningful only relative to an appropriate baseline and representative conditions.
- AI claims often omit review, correction, integration, escalation, incident response, and other human or operational work.
- Due diligence should inspect intended use, data and model lineage, evaluation, operations, risk controls, and update policy.
- Documentation creates an audit surface but does not certify that a system is trustworthy.
- A disciplined decision can accept, reject, narrow, pilot, postpone, or request stronger evidence without defaulting to hype or cynicism.