AI literacy basics
When AI Is the Wrong Tool
Learn when deterministic software, analytics, process redesign, or human service is more reliable and valuable than an AI model.
By the end you can
- Identify conditions that make AI unnecessary or irresponsible
- Compare AI with rules, analytics, process redesign, and human service
- Recognize when evidence, objectives, recovery, or value are too weak for deployment
- Apply a preflight decision process before investing in an AI solution
Key idea
Every AI feature creates an AI tax
AI can handle complexity that fixed software cannot, but it introduces uncertainty, evaluation work, monitoring, data governance, change management, and new failure modes. That ongoing cost is the AI tax.
The tax is worth paying when pattern-based capability creates enough value, and wasteful when a clear rule, a better form, a database query, or a process change solves the actual problem.
It is also paid whether or not anyone budgeted for it. In 2015 the Royal Free London NHS Foundation Trust gave Google DeepMind the records of about 1.6 million patients. The purpose was to build and test Streams, an app that alerts clinicians to acute kidney injury. On 3 July 2017 the UK Information Commissioner’s Office ruled against the trust. It had failed to comply with the Data Protection Act. Patients had not been adequately informed that their data would be used in the test. The trust signed an undertaking to establish a proper legal basis, complete a privacy impact assessment, and audit the trial. DeepMind’s own post that day put the bill in one sentence. “We underestimated the complexity of the NHS and of the rules around patient data, as well as the potential fears about a well-known tech company working in health.” The data-governance work was not avoided by skipping it. It was deferred, and then imposed.
The question is not “Can AI do this?” but “Does AI improve the whole solution after its additional costs are counted?”
Comparison
Four alternatives that often beat an AI-first plan
Choosing a simpler approach is not technological failure. It can be the more advanced engineering decision.
Google Flu Trends is the standing demonstration. It estimated influenza-like illness in the United States from search queries. It was widely presented as big data outperforming conventional surveillance. David Lazer, Ryan Kennedy, Gary King and Alessandro Vespignani wrote in Science on 14 March 2014. It had “missed high for 100 out of 108 weeks starting with August 2011”. The first version had been “part flu detector, part winter detector”. And “even 3-week-old CDC data do a better job of projecting current flu prevalence than GFT”. The cheaper approach was to take the health agency’s own lagged surveillance counts and project them forward. It was also the more accurate one.
Rules and deterministic software
Use exact logic when requirements are stable, auditable, and complete enough to specify.
- Predictable outputs
- Strong testability
- Low uncertainty
- Example: eligibility date calculation
Analytics and search
Give people clearer information when the main problem is visibility rather than prediction.
- Preserves human interpretation
- Often cheaper to maintain
- Makes evidence inspectable
- Example: searchable policy dashboard
Process redesign
Remove delays, duplication, unclear ownership, or bad incentives before automating them.
- Addresses root causes
- Can improve every case
- Reduces data confusion
- Example: one intake form instead of three
Human service
Use trained people when cases are rare, relational, morally contested, or difficult to reverse.
- Supports negotiation and empathy
- Handles novel context
- Can explain and adapt
- Example: complex crisis support
Five stop signs for an AI proposal
Stop when the objective cannot be defined well enough to recognize success. Stop when there is no credible evidence, no representative data, or no way to observe outcomes.
Stop when errors are severe and irreversible but cannot be detected or appealed. Stop when the proposed automation would obscure accountability. Stop when the expected value does not justify integration, monitoring, review, and maintenance.
A postponed AI project can be a successful risk decision.
Case
Zillow Offers: the stop sign that arrived after launch
The stop sign can also arrive after launch. On 2 November 2021 Zillow Group announced that it would wind down Zillow Offers, the arm in which the company itself bought and resold homes on the strength of its own price forecasts. Co-founder and chief executive Rich Barton gave the reason without decoration: “We’ve determined the unpredictability in forecasting home prices far exceeds what we anticipated and continuing to scale Zillow Offers would result in too much earnings and balance-sheet volatility.” The wind-down was expected to take several quarters and to include a reduction of Zillow’s workforce by approximately 25%. Reaching that judgment before the build is the same decision taken earlier and more cheaply.
Example
Situations where “no model” is the stronger design
These examples show different reasons to decline AI rather than one universal prohibition.
Vendors sometimes reach the same conclusion about their own products. On 21 June 2022 Microsoft aligned its Azure Face service with its newly published Responsible AI Standard. It said it was “retiring capabilities that infer emotional states and identity attributes such as gender, age, smile, facial hair, hair, and makeup”. It added that it would “not provide open-ended API access to technology that can scan people’s faces”. Such technology would “purport to infer their emotional states based on their facial expressions or movements”. The feature worked in the narrow sense that it returned a label. What it could not do was make the label mean what a buyer would assume it meant.
- Exact compliance calculation: the governing formula is explicit and must be applied identically, so deterministic code is easier to audit.
- Tiny rare dataset: a high-impact outcome has only a handful of historical cases, making confident learning claims unjustified.
- Broken intake process: teams disagree about the meaning of fields, so a model would automate inconsistent definitions.
- Irreversible personal decision: affected people cannot understand, challenge, or recover from an error.
- Low-value novelty: a generative feature adds impressive demos but increases review time and support burden.
- Sensitive relationship: the core value comes from trust, negotiation, or human presence rather than information processing alone.
Position
For most problems an organisation actually has, the answer is still no
Worth saying plainly, because no vendor will: for the large majority of problems inside a real organisation, the right amount of machine learning is none. Not because the technology is weak. Because the problem is a definition problem, a process problem or an incentive problem, and a model fitted to a broken process industrialises the breakage instead of repairing it.
The cases here are not exotic ones chosen to make the point. Google Flu Trends missed high for a hundred weeks out of a hundred and eight, while three-week-old data from the conventional system predicted current flu better. Zillow priced houses with a model and wound down the business built on it. The stop signs are not there to make you cautious in general; they exist so that “we should use AI for this” has to survive one question — what would we do if the model were not available, and why exactly is that worse? A proposal that cannot answer it is not ready to be built.
Choosing the simpler tool is the more advanced engineering decision, and the one that never makes it into a case study.
Visual
A preflight decision tree
The decision tree is not a formula. It forces the proposal to survive questions that technology enthusiasm often skips.
- 1
Is the user problem real and specific?
If no, investigate needs and workflow before selecting technology.
- 2
Can rules or better information solve it?
If yes, prefer the simpler approach unless AI adds clear measurable value.
- 3
Is there credible evidence to learn or evaluate?
If no, improve measurement, narrow the task, or avoid automation.
- 4
Can errors be detected, contained, and remedied?
If no, reduce authority or keep the decision human.
- 5
Does the full-system value exceed the AI tax?
Include integration, latency, review, monitoring, privacy, security, and maintenance.
- 6
Can ownership and stopping conditions be assigned?
If no one can pause or retire the system, it is not ready for deployment.
Analogy
Power tools and the wrong cut
Someone picks a power saw because it is faster. Then the task turns out to need one careful cut beside fragile wiring, where a hand tool is slower per movement and safer for the actual job.
AI resembles a power tool when it scales pattern-based work. A bad cut stays where it was made. AI errors can be opaque, copied instantly, and amplified through feedback.
Tool sophistication should match the task, consequence, and recovery path.
Case
Tay, live for a single day
Microsoft’s Tay is the compact case. The chatbot went live on Twitter on Wednesday 23 March 2016 and learned from the people who talked to it. Two days later Microsoft corporate vice-president Peter Lee published the post-mortem: “in the first 24 hours of coming online, a coordinated attack by a subset of people exploited a vulnerability in Tay”, and, as a general property, “AI systems feed off of both positive and negative interactions with people”. Microsoft took Tay offline and apologized: “We are deeply sorry for the unintended offensive and hurtful tweets from Tay.” A saw that has made the wrong cut does not keep making it faster because bystanders cheered.
Steps
Write a one-page no-hype preflight memo
Requiring a short written case makes assumptions visible before procurement or model development begins.
- 1
Problem and current baseline
Describe the workflow today, including costs, delays, errors, and who is affected.
- 2
Why AI is uniquely useful
State which complexity or uncertainty simpler software cannot handle adequately.
- 3
Evidence and evaluation
Name available data, truth signals, pilot design, and rejection criteria.
- 4
Failure and remedy
List common and severe errors, detection, fallback, appeal, and recovery.
- 5
Operating cost
Estimate integration, human review, latency, monitoring, security, and maintenance.
- 6
Decision
Choose proceed, narrow, measure first, use a hybrid, or do not use AI.
Simple first does not mean simple forever
A rule-based baseline or manual pilot can create the data and understanding needed for a later AI system. Starting simple makes the benefit of added complexity measurable.
Hybrid designs are common: deterministic policy for hard constraints, AI for ambiguous cases, and human judgment for high-impact exceptions. The goal is not ideological purity; it is dependable value.
Earn complexity by demonstrating which limitation it solves.
Key takeaways
- AI creates ongoing costs for evaluation, monitoring, governance, integration, review, and maintenance.
- Rules, analytics, process redesign, and human service often solve problems more reliably than AI.
- Undefined objectives, weak evidence, irreversible errors, missing remedy, and unclear ownership are stop signs.
- A full comparison must count the AI tax rather than comparing only model capability.
- Simple baselines and manual pilots create evidence for deciding whether added complexity is justified.
- Hybrid systems can reserve AI for ambiguity, deterministic controls for policy, and people for consequential exceptions.