Skip to content
AI.info

AI literacy basics

When AI Is the Wrong Tool

Learn when deterministic software, analytics, process redesign, or human service is more reliable and valuable than an AI model.

By the end you can

Key idea

Every AI feature creates an AI tax

AI can handle complexity that fixed software cannot, but it introduces uncertainty, evaluation work, monitoring, data governance, change management, and new failure modes. That ongoing cost is the AI tax.

The tax is worth paying when pattern-based capability creates enough value, and wasteful when a clear rule, a better form, a database query, or a process change solves the actual problem.

It is also paid whether or not anyone budgeted for it. In 2015 the Royal Free London NHS Foundation Trust gave Google DeepMind the records of about 1.6 million patients. The purpose was to build and test Streams, an app that alerts clinicians to acute kidney injury. On 3 July 2017 the UK Information Commissioner’s Office ruled against the trust. It had failed to comply with the Data Protection Act. Patients had not been adequately informed that their data would be used in the test. The trust signed an undertaking to establish a proper legal basis, complete a privacy impact assessment, and audit the trial. DeepMind’s own post that day put the bill in one sentence. “We underestimated the complexity of the NHS and of the rules around patient data, as well as the potential fears about a well-known tech company working in health.” The data-governance work was not avoided by skipping it. It was deferred, and then imposed.

The question is not “Can AI do this?” but “Does AI improve the whole solution after its additional costs are counted?”

Comparison

Four alternatives that often beat an AI-first plan

Choosing a simpler approach is not technological failure. It can be the more advanced engineering decision.

Google Flu Trends is the standing demonstration. It estimated influenza-like illness in the United States from search queries. It was widely presented as big data outperforming conventional surveillance. David Lazer, Ryan Kennedy, Gary King and Alessandro Vespignani wrote in Science on 14 March 2014. It had “missed high for 100 out of 108 weeks starting with August 2011”. The first version had been “part flu detector, part winter detector”. And “even 3-week-old CDC data do a better job of projecting current flu prevalence than GFT”. The cheaper approach was to take the health agency’s own lagged surveillance counts and project them forward. It was also the more accurate one.

FigureComparison · 4 columns

Rules and deterministic software

Use exact logic when requirements are stable, auditable, and complete enough to specify.

  • Predictable outputs
  • Strong testability
  • Low uncertainty
  • Example: eligibility date calculation

Analytics and search

Give people clearer information when the main problem is visibility rather than prediction.

  • Preserves human interpretation
  • Often cheaper to maintain
  • Makes evidence inspectable
  • Example: searchable policy dashboard

Process redesign

Remove delays, duplication, unclear ownership, or bad incentives before automating them.

  • Addresses root causes
  • Can improve every case
  • Reduces data confusion
  • Example: one intake form instead of three

Human service

Use trained people when cases are rare, relational, morally contested, or difficult to reverse.

  • Supports negotiation and empathy
  • Handles novel context
  • Can explain and adapt
  • Example: complex crisis support

Five stop signs for an AI proposal

Stop when the objective cannot be defined well enough to recognize success. Stop when there is no credible evidence, no representative data, or no way to observe outcomes.

Stop when errors are severe and irreversible but cannot be detected or appealed. Stop when the proposed automation would obscure accountability. Stop when the expected value does not justify integration, monitoring, review, and maintenance.

A postponed AI project can be a successful risk decision.

Case

Zillow Offers: the stop sign that arrived after launch

The stop sign can also arrive after launch. On 2 November 2021 Zillow Group announced that it would wind down Zillow Offers, the arm in which the company itself bought and resold homes on the strength of its own price forecasts. Co-founder and chief executive Rich Barton gave the reason without decoration: “We’ve determined the unpredictability in forecasting home prices far exceeds what we anticipated and continuing to scale Zillow Offers would result in too much earnings and balance-sheet volatility.” The wind-down was expected to take several quarters and to include a reduction of Zillow’s workforce by approximately 25%. Reaching that judgment before the build is the same decision taken earlier and more cheaply.

Example

Situations where “no model” is the stronger design

These examples show different reasons to decline AI rather than one universal prohibition.

Vendors sometimes reach the same conclusion about their own products. On 21 June 2022 Microsoft aligned its Azure Face service with its newly published Responsible AI Standard. It said it was “retiring capabilities that infer emotional states and identity attributes such as gender, age, smile, facial hair, hair, and makeup”. It added that it would “not provide open-ended API access to technology that can scan people’s faces”. Such technology would “purport to infer their emotional states based on their facial expressions or movements”. The feature worked in the narrow sense that it returned a label. What it could not do was make the label mean what a buyer would assume it meant.

  • Exact compliance calculation: the governing formula is explicit and must be applied identically, so deterministic code is easier to audit.
  • Tiny rare dataset: a high-impact outcome has only a handful of historical cases, making confident learning claims unjustified.
  • Broken intake process: teams disagree about the meaning of fields, so a model would automate inconsistent definitions.
  • Irreversible personal decision: affected people cannot understand, challenge, or recover from an error.
  • Low-value novelty: a generative feature adds impressive demos but increases review time and support burden.
  • Sensitive relationship: the core value comes from trust, negotiation, or human presence rather than information processing alone.

Position

For most problems an organisation actually has, the answer is still no

Worth saying plainly, because no vendor will: for the large majority of problems inside a real organisation, the right amount of machine learning is none. Not because the technology is weak. Because the problem is a definition problem, a process problem or an incentive problem, and a model fitted to a broken process industrialises the breakage instead of repairing it.

The cases here are not exotic ones chosen to make the point. Google Flu Trends missed high for a hundred weeks out of a hundred and eight, while three-week-old data from the conventional system predicted current flu better. Zillow priced houses with a model and wound down the business built on it. The stop signs are not there to make you cautious in general; they exist so that “we should use AI for this” has to survive one question — what would we do if the model were not available, and why exactly is that worse? A proposal that cannot answer it is not ready to be built.

Choosing the simpler tool is the more advanced engineering decision, and the one that never makes it into a case study.

Visual

A preflight decision tree

The decision tree is not a formula. It forces the proposal to survive questions that technology enthusiasm often skips.

FigureProcess · 6 steps
  1. 1

    Is the user problem real and specific?

    If no, investigate needs and workflow before selecting technology.

  2. 2

    Can rules or better information solve it?

    If yes, prefer the simpler approach unless AI adds clear measurable value.

  3. 3

    Is there credible evidence to learn or evaluate?

    If no, improve measurement, narrow the task, or avoid automation.

  4. 4

    Can errors be detected, contained, and remedied?

    If no, reduce authority or keep the decision human.

  5. 5

    Does the full-system value exceed the AI tax?

    Include integration, latency, review, monitoring, privacy, security, and maintenance.

  6. 6

    Can ownership and stopping conditions be assigned?

    If no one can pause or retire the system, it is not ready for deployment.

Analogy

Power tools and the wrong cut

Someone picks a power saw because it is faster. Then the task turns out to need one careful cut beside fragile wiring, where a hand tool is slower per movement and safer for the actual job.

AI resembles a power tool when it scales pattern-based work. A bad cut stays where it was made. AI errors can be opaque, copied instantly, and amplified through feedback.

Tool sophistication should match the task, consequence, and recovery path.

Case

Tay, live for a single day

Microsoft’s Tay is the compact case. The chatbot went live on Twitter on Wednesday 23 March 2016 and learned from the people who talked to it. Two days later Microsoft corporate vice-president Peter Lee published the post-mortem: “in the first 24 hours of coming online, a coordinated attack by a subset of people exploited a vulnerability in Tay”, and, as a general property, “AI systems feed off of both positive and negative interactions with people”. Microsoft took Tay offline and apologized: “We are deeply sorry for the unintended offensive and hurtful tweets from Tay.” A saw that has made the wrong cut does not keep making it faster because bystanders cheered.

Steps

Write a one-page no-hype preflight memo

Requiring a short written case makes assumptions visible before procurement or model development begins.

FigureProcess · 6 steps
  1. 1

    Problem and current baseline

    Describe the workflow today, including costs, delays, errors, and who is affected.

  2. 2

    Why AI is uniquely useful

    State which complexity or uncertainty simpler software cannot handle adequately.

  3. 3

    Evidence and evaluation

    Name available data, truth signals, pilot design, and rejection criteria.

  4. 4

    Failure and remedy

    List common and severe errors, detection, fallback, appeal, and recovery.

  5. 5

    Operating cost

    Estimate integration, human review, latency, monitoring, security, and maintenance.

  6. 6

    Decision

    Choose proceed, narrow, measure first, use a hybrid, or do not use AI.

Simple first does not mean simple forever

A rule-based baseline or manual pilot can create the data and understanding needed for a later AI system. Starting simple makes the benefit of added complexity measurable.

Hybrid designs are common: deterministic policy for hard constraints, AI for ambiguous cases, and human judgment for high-impact exceptions. The goal is not ideological purity; it is dependable value.

Earn complexity by demonstrating which limitation it solves.

Key takeaways