Skip to content
AI.info

Responsible AI

Intended Purpose, Context of Use, and Foreseeable Misuse

Specify intended purpose, excluded uses, operating context, foreseeable misuse, and change controls before deployment.

By the end you can

Risk lives in the use, not in the model

Risk attaches to a system’s intended purpose and its actual context, not merely to the underlying model. The same component can be low impact in one workflow and unacceptable in another. Populations, stakes, incentives and available recourse differ. A useful purpose statement names the user, the affected population, the decision and the timing. It then names the operating environment, the required evidence, the excluded uses and the expected human role. Foreseeable misuse includes the incentive to repurpose the system, not only malicious attack.

What happens when a model meets a population its evidence was never gathered on has been measured. Epic’s proprietary Sepsis Model — the ESM — was validated externally at Michigan Medicine by Wong and colleagues. The cohort ran to 27,697 patients and 38,455 hospitalizations, between 6 December 2018 and 20 October 2019. Sepsis occurred in 2,552 hospitalizations: 6.6%, rounded to 7% in the published abstract. The model’s discrimination was an area under the receiver operating characteristic curve of 0.63 (95% CI, 0.62–0.64). At a score of 6 or higher it alerted on 6,971 of the 38,455 hospitalizations (18%). It did not identify 1,709 patients with sepsis (67%).

The authors call that “substantially worse than that reported by Epic Systems (AUC, 0.76-0.83)”. Their conclusion, published in JAMA Internal Medicine in 2021, is flat: “This external validation cohort study suggests that the ESM has poor discrimination and calibration in predicting the onset of sepsis.” Nothing about the model changed between the developer’s numbers and Michigan Medicine’s. The population and the setting did.

The sentiment model below passed every technical review. What it never had was a written line saying that performance review was an excluded use.

Risk assessment done at the component level answers a narrower question than the one that matters: what this particular decision does to this particular population.

Comparison

Component reuse, Purpose expansion, or Context drift?

Component reuse, purpose expansion and context drift can look identical in a change log. What they oblige the company running the system to do is very different.

Component reuse keeps a technical module inside the approved purpose — an updated sentiment encoder still doing support triage. Population and consequence remain bounded. Evidence can sometimes transfer, and compatibility testing may still be required. Context drift is the mirror image: the declared purpose stays constant while reality changes under it. New languages, incentives or populations appear. Monitoring has to detect the boundary violation, and the remedy is restriction or retraining rather than a compatibility test. Crisis-period usage patterns are the standard example.

Purpose expansion — the output beginning to influence a new decision — is the one whose consequences are written into statute. Under Article 25(1)(c) of Regulation (EU) 2024/1689, a distributor, importer, deployer or other third party stands in the provider’s shoes where “they modify the intended purpose of an AI system, including a general-purpose AI system, which has not been classified as high-risk and has already been placed on the market or put into service in such a way that the AI system concerned becomes a high-risk AI system in accordance with Article 6.”

The organisation that did the repurposing is then considered the provider, and assumes the provider obligations of Article 16. Under Article 25(2) the provider that originally placed the system on the market ceases to be the provider of it. So “changes the impact and legal context” is not a caution about extra paperwork. It is a transfer of legal identity. The vendor drops out. The new purpose is inherited whole by whoever chose it. Old metrics may become irrelevant and fresh stakeholder and risk analysis is required — but the analysis is now the new provider’s to produce.

FigureComparison · 3 columns

Component reuse

A technical module is reused within the approved purpose.

  • May still require compatibility testing
  • Evidence can sometimes transfer
  • Population and consequence remain bounded
  • Example: updated sentiment encoder in support triage

Purpose expansion

The output begins to influence a new decision.

  • Changes the impact and legal context
  • Requires new stakeholder and risk analysis
  • Old metrics may become irrelevant
  • Example: support score used for employment

Context drift

The declared purpose stays constant while reality changes.

  • New languages, incentives, or populations appear
  • Monitoring must detect boundary violations
  • May require restriction or retraining
  • Example: crisis-period usage patterns

Example

The sentiment model that moved to performance reviews

A customer-support sentiment model is later reused to score employee attitude during performance reviews. The model, the data and the interface remain similar. The consequence and the power relationship change completely. And in at least one jurisdiction the reuse becomes conditionally unlawful the moment it happens.

That jurisdiction is New York City. Local Law 144 of 2021 “prohibits employers and employment agencies from using an automated employment decision tool unless the tool has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates.” It was enacted on 11 December 2021 and took effect on 1 January 2023. The Department of Consumer and Worker Protection began enforcing it on 5 July 2023.

An AEDT, in that law, is any computational process derived from machine learning, statistical modeling, data analytics or artificial intelligence. It has to issue simplified output — a score, a classification or a recommendation. And that output has to substantially assist or replace discretionary decision making for employment decisions. A support sentiment score that starts substantially assisting an appraisal meets all three.

  • Original purpose: Prioritize support conversations for faster service recovery. Customers were the population, queue order was the decision, and triage was the only thing the evidence was ever gathered about.
  • New use: Influence compensation and employment decisions. Annex III point 4(b) of Regulation (EU) 2024/1689 names that use in its own words: “AI systems intended to be used to make decisions affecting terms of work-related relationships, the promotion or termination of work-related contractual relationships, to allocate tasks based on individual behaviour or personal traits or characteristics or to monitor and evaluate the performance and behaviour of persons in such relationships.” Point 4(a) separately covers recruitment and candidate evaluation.
  • Population shift: Customers and employees produce language under different incentives and constraints. The evaluation evidence collected on support conversations describes a population that is no longer the one being scored.
  • Power shift: Employees cannot freely opt out of workplace monitoring. That is why the obligations attaching to the new use are public ones — a published bias audit, a notice to the people scored — rather than internal quality checks.
  • Control failure: The inventory tracked the model version, not the approved purpose or the prohibited reuse. Under Article 25(1)(c) the team that made the move would itself be considered the provider of a high-risk system. The inventory held no field in which that could even be recorded.

Visual

What an approval actually approves

An approval approves a chain, not a model. The intended task: the specific function and decision the system may support. The approved context: population, location, language, timing and operational conditions. The excluded uses: decisions or domains for which evidence is absent or risk is unacceptable. The foreseeable misuse: likely repurposing, overreliance, gaming and secondary use. And the change triggers: the events that require renewed assessment or approval.

That chain is not a local invention. The NIST AI Risk Management Framework 1.0 puts context before measurement, and it was published on 26 January 2023. The first category of its MAP function is ‘Context is established and understood’. Its first subcategory reads: “MAP 1.1: Intended purposes, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented.”

Purpose, law, norm and deployment setting are named in one numbered requirement, to be documented together. The same framework lists “anticipating risks of the use of AI systems beyond intended use” among the outcomes broad engagement is meant to produce. That is the misuse link in the chain, given a home in the same document.

FigureProcess · 5 steps
  1. 1

    Intended task

    The specific function and decision the system may support.

  2. 2

    Approved context

    Population, location, language, timing, and operational conditions.

  3. 3

    Excluded use

    Decisions or domains for which evidence is absent or risk is unacceptable.

  4. 4

    Foreseeable misuse

    Likely repurposing, overreliance, gaming, and secondary use.

  5. 5

    Change trigger

    Events that require renewed assessment or approval.

Key idea

Where a purpose and context governance control can still fail

A broad label such as “analytics” or “decision support” is not an intended-purpose statement. Vague purpose makes every later use appear authorized. It also stops anyone from evaluating it meaningfully.

The alternative is not abstract, and it has been specified and dated. A model card carries a standing field for what a model is not for. The framework is from 2019, by Mitchell, Gebru and seven co-authors, and it opens on the problem this section is about: “In order to clarify the intended use cases of machine learning models and minimize their usage in contexts for which they are not well suited, we recommend that released models be accompanied by documentation detailing their performance characteristics.”

In that framework “Out-of-scope uses” sits in the Intended Use section, beside primary intended uses and primary intended users. It is modelled explicitly on warning labels. The kind of entry it asks for is as narrow as “not for use on text examples shorter than 100 tokens”. An exclusion at that resolution is testable against an observed use. “Decision support” is not.

Even a precisely written purpose cannot predict every emergent use. Controls need telemetry, access restrictions, contractual limits, and people empowered to challenge repurposing. Nothing technical changed when the sentiment model moved to performance reviews. The sentence naming that as an excluded use had simply never been written.

Vague purpose language authorizes the next use in advance; what catches the drift afterwards is telemetry, access and contractual limits, and someone empowered to object.

Example

Four things a purpose statement should survive

These drills attack the vaguest part of any approval, which is usually the purpose sentence itself. Each one has a published model to imitate. MAP 1.1 supplies the numbered documentation requirement. The model card supplies the “Out-of-scope uses” field. Article 25(1)(c) supplies the value-chain rule. And the Michigan Medicine validation of the ESM supplies the boundary conditions, exposed by measuring the same model on a different population.

  • Purpose sentence: Write a one-sentence purpose specific enough to test. Name the population, the setting and the decision — MAP 1.1 asks for purposes, context-specific laws, norms and expectations and prospective deployment settings, documented together.
  • Misuse story: Describe a realistic internal repurposing, not only an external attacker. The support-to-appraisal move is the shape to aim for: it lands a model in Annex III point 4(b). Note who becomes the provider under Article 25(1)(c) when it happens.
  • Access control: Identify who can connect the model to a new workflow. Then write the exclusion at warning-label resolution, in the manner of a model card that would refuse a model “not for use on text examples shorter than 100 tokens”.
  • Change signal: Define telemetry that would reveal use outside the approved boundary. A population arriving that the evaluation never covered is the condition under which an AUC of 0.63 (95% CI, 0.62–0.64) can sit behind a developer’s reported 0.76-0.83, with nothing in the software having changed.

Steps

From the exact decision to the change gate

Naming the decision comes first. Gating changes comes last. Most unauthorized reuse is caught in the three steps between them. Name the decision: describe the exact action or judgment the output influences. Define operating conditions: record populations, languages, environments and timing. Write exclusions: prohibited domains, decisions and forms of automation. Model misuse incentives: ask who benefits from reuse, surveillance or scope expansion. Only then gate changes, requiring renewed review when purpose, population, consequence or autonomy changes.

A regulator has already built that last step into a submission. On 4 December 2024 the US Food and Drug Administration announced its final guidance on predetermined change control plans for AI-enabled device software functions. It was issued jointly by CDRH, CBER, CDER and the Office of Combination Products. Its instruction has three named components: “This guidance recommends that a PCCP describe the planned AI-DSF modifications, the associated methodology to develop, validate, and implement those modifications, and an assessment of the impact of those modifications.”

A manufacturer that wants to change an AI-enabled device after clearance submits that plan in advance. FDA reviews the plan as part of the marketing submission. Each modification authorised within it then does not require a new submission. The change gate is itself the object of prior review, rather than a promise to reassess afterwards.

FigureProcess · 5 steps
  1. 1. Name the decision

    Describe the exact action or judgment influenced by the output.

  2. 2. Define operating conditions

    Record populations, languages, environments, and timing.

  3. 3. Write exclusions

    List prohibited domains, decisions, and forms of automation.

  4. 4. Model misuse incentives

    Ask who benefits from reuse, surveillance, or scope expansion.

  5. 5. Gate changes

    Require renewed review when purpose, population, consequence, or autonomy changes.

A purpose you cannot test is only a label

A purpose that cannot be tested against an observed use is not a boundary. It is a label.

Write down in advance which changes force redesign, which force restriction, which trigger remedy, and which end the deployment.

Both halves of that boundary have a definition in law. Regulation (EU) 2024/1689 was signed on 13 June 2024 and published in the Official Journal on 12 July 2024. Article 3(12) defines intended purpose as the use for which an AI system is intended by the provider, including the specific context and conditions of use. It then says where that has to be written down: the instructions for use, promotional or sales materials and statements, and the technical documentation. That middle clause — “including the specific context and conditions of use” — is the one usually dropped in paraphrase. It is also the one that makes context part of the declaration rather than a commentary on it.

Article 3(13) then does the governance work: “‘reasonably foreseeable misuse’ means the use of an AI system in a way that is not in accordance with its intended purpose, but which may result from reasonably foreseeable human behaviour or interaction with other systems, including other AI systems.” It is written by reference to the intended purpose, and it reaches past it — to the uses that human behaviour and system interaction will produce anyway. Together the two definitions turn a marketing sentence about scope into something an assessment can be run against.

Key takeaways