Skip to content
AI.info

Responsible AI

Labor, Environmental, and Supply-Chain Impacts

Evaluate worker conditions, data labor, energy, water, hardware, sourcing, and environmental externalities across the AI lifecycle.

By the end you can

Key idea

A lower-energy model that runs far more often

A lower-energy model is not automatically the responsible choice. If the product drives far more usage, surveillance, or low-quality automated decisions, the saving is spent and then some. A team has to weigh efficiency against necessity and total demand. Environmental and labor data are often incomplete, vendor-controlled, or hard to allocate among products. Disclose that uncertainty. Do not convert it into a precise sustainability claim.

The uncertainty starts before anyone disputes the data. One median text prompt to Gemini Apps takes 0.24 Wh. The same prompt takes 0.10 Wh. Both numbers come from Google's own production measurement. The first counts the full stack; the second counts active accelerator power on highly utilised machines. Same prompt, 2.4× apart, decided entirely by where the counting stops. And neither figure, at either width, records the shift worked by the data labelers Sama employed on behalf of OpenAI, at take-home wages of $1.32–$2 an hour.

Publish the 0.10 Wh without the boundary that produced it and the buyer inherits a sustainability claim the measurement was never scoped to support.

Visual

The footprint reaches past the data center

The footprint of a system reaches well past the data center it runs in. Five stages carry it: annotation labor, workplace deployment, compute operations, hardware supply chain, rebound and scale. Most of them have already been measured by somebody.

Annotation labor: around three dozen Kenyan workers labeling descriptions of sexual abuse, hate speech and violence, at quotas of 150–250 passages per nine-hour shift. Compute operations: 176 TWh of US data-centre electricity in 2023, 4.4% of US electricity. Hardware supply chain: the distance between BLOOM's 24.7 tonnes and its 50.5 is what counting equipment manufacturing adds. Workplace deployment: a jurisdiction fight at Nairobi over who employs 185 former Facebook content moderators.

Rebound and scale is the stage nobody has a meter for. It is also the stage that drops off the list.

FigureProcess · 5 steps
  1. 1

    Data and annotation labor

    Collection, labeling, translation, moderation, and quality control.

  2. 2

    Workplace deployment

    Job redesign, monitoring, deskilling, workload, and bargaining power.

  3. 3

    Compute operations

    Energy, cooling, water, utilization, and geographic grid conditions.

  4. 4

    Hardware supply chain

    Mining, manufacturing, logistics, lifespan, and e-waste.

  5. 5

    Rebound and scale

    Lower unit cost can increase total use and absolute footprint.

Comparison

Efficiency metric, Absolute footprint, or Lifecycle impact?

Efficiency per request can improve while the absolute footprint grows. Only a lifecycle view puts both numbers on the same page.

Running 1,000 inferences costs 0.002 kWh if the task is text classification. It costs 2.907 kWh if the task is image generation. That is a spread of over 1,450× inside the single column labelled joules per completed task. It is settled by which task you chose, not by how well you optimised it. Luccioni and colleagues measured it across 88 models in 2024.

The same study lays the efficiency metric alongside the absolute footprint. BLOOMz-7B would need 592,570,000 inferences before its deployment energy equalled the 59,257 kWh it took to train (51,686 kWh) and fine-tune (7,571 kWh) it. That sounds unreachable until you count users. The authors write: “Even assuming a single query per user, which is rarely the case, the energy costs of deploying it would surpass its training costs after a few weeks or months of deployment.” Training is a number you pay once. Serving is a number that keeps arriving.

FigureComparison · 3 columns

Efficiency metric

Measures resources per task, token, or request.

  • Useful for engineering optimization
  • Can improve while total impact grows
  • Depends on workload definition
  • Example: joules per completed task

Absolute footprint

Measures total resource or emissions burden.

  • Reflects scale and demand growth
  • Needs boundary and allocation choices
  • Can include location and time
  • Example: annual energy and water use

Lifecycle impact

Includes upstream and downstream consequences.

  • Captures hardware and disposal
  • Includes labor and supply-chain conditions
  • More difficult and uncertain
  • Supports procurement and product decisions

192 lbs for the model, 626,155 for the search that found it

Responsible AI includes labor and environmental conditions across data collection, annotation, content review. It includes them across hardware production, training, serving, maintenance, and disposal. Automation often redistributes work and cost rather than eliminating them. A lifecycle assessment should consider job quality, surveillance, exposure to harmful material, fair compensation, worker voice. It should then consider energy source, water use, utilization, embodied hardware impact, electronic waste, and rebound effects. Absolute impact and marginal product impact are different questions.

Training one Transformer (big) model emitted 192 lbs CO2e. Reaching that same architecture through neural architecture search emitted 626,155 lbs CO2e, at a cloud cost of $942,973–$3,201,722. That is roughly 3,260 times the published cost of the run itself. For scale, an average car emits 126,000 lbs including fuel over its whole lifetime.

Strubell and colleagues published both figures in 2019. They also stated what all that searching bought: “So et al. (2019) report that NAS achieves a new state-of-the-art BLEU score of 29.7 for English to German machine translation, an increase of just 0.1 BLEU at the cost of at least $150k in on-demand compute time and non-trivial carbon emissions.” Thousands of training runs' worth of compute. A tenth of a BLEU point.

Rebound effects close the list and are the item most often left off it. Cheaper moderation tends to buy more moderation, not less.

Scope the assessment to the training run and the answer is 192 lbs; scope it to the search that found the model and it is 626,155.

Example

Three dozen labelers, $200,000, and an appeal dismissed with costs

Around three dozen Kenyan workers read and labeled descriptions of sexual abuse, hate speech and violence, at quotas of 150–250 passages per nine-hour shift. They worked under three contracts worth about $200,000, signed by OpenAI with the outsourcing firm Sama in late 2021. TIME documented the arrangement on 18 January 2023: “The data labelers employed by Sama on behalf of OpenAI were paid a take-home wage of between around $1.32 and $2 per hour depending on seniority and performance.” Sama ended the work in February 2022, eight months early.

Who answers for such conditions is not a rhetorical question. It has been litigated. Meta argued that Kenya's Employment and Labour Relations Court had no jurisdiction over it. On 20 September 2024 the Court of Appeal at Nairobi dismissed that argument with costs. “The upshot of our above findings is that the appellants’ appeals numbers E232 of 2023 and E445 of 2023 are devoid of merit and both appeals are hereby dismissed with costs to the respondents,” the court held.

Daniel Motaung and the 185 former Facebook content moderators engaged through Sama may now proceed in that court. Motaung's petition covers working conditions, unfair labour practices and constitutional violations. The moderators have a separate constitutional petition over their mass redundancy.

  • Invisible labor: Around three dozen Kenyan workers read and labeled descriptions of sexual abuse, hate speech and violence, so that a downstream product could be described as automated.
  • Incentive pressure: Quotas of 150–250 passages per nine-hour shift, at take-home wages of $1.32–$2 per hour. The contracts were worth about $200,000, and Sama ended them in February 2022, eight months early.
  • Resource demand: The compute side of the same lifecycle carries its own boundary problem. BLOOM's final training reads 24.7 tonnes of CO2 equivalent counting dynamic power alone, and 50.5 tonnes counting equipment manufacturing and the rest of the operational chain.
  • Geographic burden: The labeling was contracted into Kenya, and the jurisdiction fight was heard at Nairobi. The electricity lands somewhere else again — 176 TWh of US data-centre demand in 2023, projected at 325–580 TWh by 2028.
  • Disclosure gap: Meta argued it was not the moderators' employer. The Court of Appeal refused to settle that at the interlocutory stage, leaving both petitions to proceed in the Employment and Labour Relations Court. Closing this gap has so far required a court.

Example

Counting the labor, the energy, and the demand

Three of these drills count things. The fourth asks the product question nobody schedules: whether this feature should generate anything at all.

The vendor-evidence drill is no longer a request for goodwill. The EU AI Act requires a general-purpose AI model provider's technical documentation to record the “known or estimated energy consumption of the model. With regard to point (e), where the energy consumption of the model is unknown, the energy consumption may be based on information about computational resources used.” The law has been in force since 1 August 2024, and it supplies its own fallback method. A provider who says the figure is unavailable is answering a question the law has already anticipated.

  • Labor map: List every human task required to produce, evaluate, and operate the system. Around three dozen labelers working under a $200,000 contract are a line in that list, not a footnote to a training-set description.
  • Resource baseline: Estimate requests, model size, utilization, energy source, and cooling context, and mark where serving overtakes training. BLOOMz-7B needs 592,570,000 inferences before its deployment energy equals the 59,257 kWh of training and fine-tuning behind it.
  • Demand question: Ask whether the product should reduce, batch, cache, or avoid generation. The distance between 0.002 kWh and 2.907 kWh per 1,000 inferences is a decision about what you generate, not about how efficiently you generate it.
  • Vendor evidence: Request subcontractor, worker-protection, resource, and hardware-lifecycle information. The AI Act already requires the model provider to hold the energy figure, though it exempts providers of models released under a free and open-source licence unless the model is classified as posing systemic risk.

Steps

Where published impact figures stop

The lifecycle traced here includes contractors and disposal. That is exactly where most published impact figures stop.

Step one is where the omission usually happens. The paper announcing a model reports the run, not the neural architecture search that found it. That is how 192 lbs CO2e and 626,155 lbs CO2e come to describe the same architecture.

Step four is what a court ends up doing when nobody wrote the protections in advance. Meta's position was that the Employment and Labour Relations Court had no jurisdiction over it. On 20 September 2024 that argument was dismissed with costs.

Step five now has a statutory calendar beside it. The AI Act obliges the Commission to report by 2 August 2028, and every four years thereafter, on progress on standardisation deliverables for the energy-efficient development of general-purpose AI models. The same regulation presumes the high-impact capabilities that trigger systemic-risk classification wherever training compute exceeds 10^25 floating point operations.

FigureProcess · 5 steps
  1. 1. Map the lifecycle

    Include contractors, reviewers, infrastructure, hardware, and disposal.

  2. 2. Identify burdens

    Assess working conditions, exposure, energy, water, materials, and local context.

  3. 3. Measure alternatives

    Compare non-AI, smaller-model, cached, batch, and demand-reduction options.

  4. 4. Contract protections

    Set labor, transparency, resource, subcontractor, and remediation requirements.

  5. 5. Track absolute impact

    Monitor scale, rebound, location, incidents, and retirement.

Report the absolute number, not the ratio

Efficiency gains that lower unit cost tend to raise total use. That is why the absolute number is the one worth reporting.

The United States has that claim as a measured national series rather than an assertion. Data-centre electricity use held near 60 TWh from 2014 to 2016, while efficiency work absorbed the growth. It reached 76 TWh in 2018 — 1.9% of US electricity. In 2023 it reached 176 TWh, or 4.4%. The projection for 2028 is 325–580 TWh, between 6.7% and 12.0%.

The numbers are Lawrence Berkeley National Laboratory's, published in 2024. Its executive summary names the break: “While many of these efficiency strategies continue to provide significant energy efficiency improvements in data center design and operation, the expansion of data center services into areas that require new types of hardware has ended the era of generally flat data center energy use.” Efficiency did not stop working. It stopped being enough to hold the total down.

Define when labor and environmental impact governance requires the team to redesign, restrict, remedy, or retire the system.

Case

BLOOM at 24.7 tonnes, or at 50.5

Two published estimates give this argument its absolute numbers. BLOOM's final training emitted about 24.7 tonnes of CO2 equivalent, counting dynamic power alone. Count equipment manufacturing and the rest of the operational chain, and the same run emitted 50.5 tonnes. Luccioni and colleagues published both figures for that one training run. The gap between them is the boundary argument, stated in tonnes. It is roughly a doubling.

The same argument holds at serving, and there the vendor made it itself. Google's production measurement puts the median Gemini Apps text prompt at 0.24 Wh, 0.03 gCO2e and 0.26 mL of water under a full-stack boundary. Under the narrower existing approach, which counts active accelerator power on highly utilised machines, the same prompt is 0.10 Wh. Idle machines and data-centre overhead contribute 0.02 Wh each. Google's own authors draw the conclusion: “The comprehensive approach reveals a total energy consumption that is 2.4 times greater than the estimate from the existing approach, demonstrating the need to standardize measurement approaches and boundaries to be more inclusive of all energy consumptive activities.”

NIST made the same point structurally. Environmental impacts are one of the twelve risks named in its Generative AI Profile of July 2024.

Figure

Slightly more than half of BLOOM’s published footprint sits outside the electricity the training run itself drew — the same run, counted at two boundaries.

Key takeaways