Skip to content
AI.info

Kinds of learning

Regression: Predicting Quantities Without Pretending to Know the Future

Understand point, interval, and quantile predictions, common regression targets, and the risks of averages, extrapolation, and misleading precision.

By the end you can

Analogy

A weather forecast with more than one number

Planning an outdoor event starts from a forecast. An expected temperature is useful, but a likely range and the chance of extreme weather may matter more than the average.

Regression has the same need to represent uncertainty around quantities. Many business decisions penalize late estimates far more than early ones, while weather intuition treats an error as equally bad in either direction.

The US National Hurricane Center publishes exactly this shape of answer, and it states the coverage it is claiming. Its track forecast cone is built by drawing circles around the forecast centre position, at 12, 24, 36, 48, 60, 72, 96 and 120 hours. In the agency’s own words, “the size of each circle is set so that two-thirds of historical official forecast errors over a 5-year sample fall within the circle.” Two habits there are worth copying. The interval is calibrated against measured past error rather than asserted. And the promise attached to it is two-thirds rather than certainty.

A single number can be mathematically neat and operationally incomplete.

What makes a task regression

Regression predicts a quantity: time, demand, cost, temperature, probability-adjusted loss, remaining life. The numeric scale carries information about distance. An error of ten units is usually larger than an error of one.

Not every numbered label is a regression target. Postal codes, product identifiers and ordinal ratings may use digits without representing continuous magnitude.

Example

Numeric predictions serving different decisions

Regression targets should be defined with the downstream action in mind.

  • Logistics: estimate arrival time so a dock can allocate staff and space.
  • Energy: forecast hourly demand to schedule generation and reserves.
  • Maintenance: predict remaining useful life while preserving high-risk uncertainty.
  • Insurance: estimate claim severity rather than merely predicting whether a claim occurs.
  • Agriculture: forecast yield under stated weather and management assumptions.
  • Customer operations: estimate handling time to plan queue capacity.

Visual

Three ways to communicate a numeric forecast

Different regression outputs answer different operational questions. A point estimate, a set of conditional quantiles and a prediction interval are three different promises. Only the last two say anything about how often the model expects to be wrong.

A stated coverage level is a testable claim, and one supervisory regime has been auditing such claims by counting misses since 1996. The Basel Committee on Banking Supervision’s backtesting framework, published in January 1996, compares a bank’s 99th-percentile one-day value-at-risk against 250 days of trading outcomes. Section II states the principle: “The backtests to be applied compare whether the observed percentage of outcomes covered by the risk measure is consistent with a 99% level of confidence.” Then it counts. Zero to four exceptions in 250 days is the green zone; “The range from five to nine exceptions constitutes the yellow zone.”; ten or more is the red zone. Those boundaries are not conventions. They are the binomial arithmetic of the claim itself: “For 250 observations, it can be seen that five or fewer exceptions will be obtained 95.88% of the time when the true level of coverage is 99%.”

EU law wrote the same counts down. Article 366 of Regulation (EU) No 575/2013 defines an overshooting as “a one-day change in the portfolio’s value that exceeds the related one-day value-at-risk number generated by the institution’s model”. It then tabulates the capital multiplier addend by the number of overshootings in the most recent 250 business days: 0.00 for fewer than 5, then 0.40 at 5, 0.50 at 6, 0.65 at 7, 0.75 at 8, 0.85 at 9, and 1.00 at 10 or more. An interval that misses more often than it promised is not merely embarrassing. It is priced.

FigureHierarchy · 3 levels
  • Point estimate

    One central prediction, often optimized for mean or median error.

    • Quantile estimates

      Several conditional percentiles describe asymmetric uncertainty.

      • Prediction interval

        A range aims to contain a stated proportion of future outcomes under defined conditions.

Comparison

What different losses ask the model to prioritize

The loss does two things. It changes which summary of the target distribution is favored, and it sets how strongly large errors matter. Squared error pushes toward the conditional mean and lets a few extreme residuals dominate. Absolute error pushes toward the conditional median and survives heavy tails better. Quantile loss weights over- and underprediction asymmetrically, and returns a chosen percentile instead of a centre.

Quantile loss is not a classroom refinement, and its evaluation is not left to taste. The M5 “Uncertainty” competition ran on Kaggle from 2 March to 30 June 2020, on Walmart data, and demanded exactly this shape of output. Entrants had to submit “nine different quantiles (0.005, 0.025, 0.165, 0.250, 0.500, 0.750, 0.835, 0.975, and 0.995), that can sufficiently describe the complete distributions of future sales”. That line is from the abstract of the results paper in the International Journal of Forecasting. Nine numbers, not one, for each of 42,840 hierarchically aggregated series, scored by the weighted scaled pinball loss, across 909 submissions. That is what a service-level objective looks like when someone has to keep score on it.

FigureComparison · 3 columns

Squared error

Large residuals receive rapidly increasing penalty.

  • Often targets the conditional mean
  • Sensitive to extreme observations
  • Smooth and widely used
  • Suitable when large misses are especially costly

Absolute error

Residuals contribute in direct proportion to magnitude.

  • Often targets the conditional median
  • More robust to large outliers
  • Less dominated by rare extremes
  • Useful with heavy-tailed noise

Quantile loss

Over- and underprediction receive asymmetric weights.

  • Targets a chosen conditional quantile
  • Supports service-level planning
  • Captures asymmetric risk
  • Requires evaluation of quantile coverage

Key idea

Regression can produce numbers far beyond its evidence

A flexible model may interpolate well inside the observed range and still behave unpredictably outside it. Forecasting unprecedented temperatures, prices, loads or ages requires assumptions that held-out random rows may never test.

The clearest objection to extrapolation ever recorded was made out loud, the night before it mattered. Challenger launched on 28 January 1986 in the cold. The Rogers Commission report states that “The ambient temperature at time of launch was 36 degrees Fahrenheit, or 15 degrees lower than the next coldest previous launch.” Morton Thiokol engineers had argued against flying. The same report records engineer Roger Boisjoly’s summary of the engineering position: “The conclusion was we should not fly outside of our data base, which was 53 degrees.” The launch proceeded at 36°F. Dalal and colleagues later fitted the 23 pre-accident launches, in a 1989 paper in the Journal of the American Statistical Association. They published the number nobody had produced that night: “a probabilistic risk assessment at 31°F, the temperature at which Challenger was launched, yields at least a 13% probability of catastrophic field-joint O-ring failure. Postponement to 60°F would have reduced the probability to at least 2%.” The data base ended at 53 degrees. The physics did not.

Extrapolation also has an invoice. On 2 November 2021 Zillow announced it would wind down Zillow Offers, the business that bought homes against its own price forecasts. Rich Barton, the company’s co-founder and CEO, gave the reason in the third-quarter 2021 results release: “We’ve determined the unpredictability in forecasting home prices far exceeds what we anticipated and continuing to scale Zillow Offers would result in too much earnings and balance-sheet volatility.” The same release prices the miss: “Included in the company’s third-quarter financial results is a write-down of inventory of approximately $304 million within the Homes segment as a result of purchasing homes in Q3 at higher prices than the company’s current estimates of future selling prices.” It warned of a further $240–265 million of losses in the fourth quarter. The wind-down included cutting about 25% of Zillow’s workforce.

Evaluate by time, region, and target range. When extrapolation is unavoidable, combine domain constraints, scenarios, and explicit uncertainty rather than trusting smooth output alone.

A plausible-looking number outside the training support is still an unsupported claim.

Case

April 2020: a forecast whose 95% intervals missed half the time or more

COVID-19 produced a public test of what happens when a curve is projected past its evidence. A University of Sydney report dated 8 April 2020 ran that test. Marchant and colleagues took the daily death counts predicted by the Institute for Health Metrics and Evaluation model and compared them with what US states actually recorded from 30 March to 2 April. The model was treated as a black box. The only question asked of it was whether reality landed inside its stated intervals. A correctly calibrated 95% interval should miss about 5% of the time. On 30 March, 27% of states fell inside it. Their summary: “we found that the observed percentage of death counts that lie outside the 95% PI to be in the range 49% - 73%, which is more than an order of magnitude above the expected percentage.” The point estimates were not the whole failure. The advertised uncertainty was.

Figure

The number attached to an interval is a claim about a rate, and it can be audited against what actually happened.

Transformations change interpretation

Logarithms and other target transformations can stabilize scale or reduce the influence of extreme values. Predictions transformed back to original units may require care. Averages do not simply invert through nonlinear functions.

Document the transformed objective, the back-transformation, and the metric domain. Otherwise a mathematically valid pipeline can produce a misleading business interpretation.

The skew that motivates all of this is measurable, and US healthcare spending is where the health-economics literature measured it. An AHRQ statistical brief published in March 2025 reports: “In 2022, the top 1 percent of people ranked by their healthcare expenditures accounted for 21.7 percent of total healthcare expenditures, while the bottom 50 percent accounted for less than 3 percent.” That top 1 percent averaged $147,071 each. The same brief adds that “People with the top 5 percent of expenses in 2022 accounted for about half (49.7 percent) of healthcare expenses.” A mean fitted on the log of a distribution shaped like that, exponentiated back, is not the mean of the dollars.

The problem has a name and a fix that predate machine learning by decades. Naihua Duan proposed the smearing estimate in 1983, in the Journal of the American Statistical Association: a nonparametric estimate of the expected response on the untransformed scale, where the regression itself is fitted on a transformed scale. It is needed because exponentiating a fitted mean of log y does not return the mean of y. Manning and Mullahy reopened the question for skewed cost data in 2001, in the Journal of Health Economics. Which retransformation is correct, they showed, depends on whether the log-scale errors are heteroscedastic. If a pipeline fits on a log scale and reports in currency, the step back is a modelling assumption. Someone has to own it. It is not arithmetic.

Steps

Review a numeric target before modeling

A regression design review should expose scale and decision consequences. Define the units. Inspect the shape of the target. Match the loss to the decision. Request the uncertainty the action actually needs. Test extrapolation on time, geography and range slices. Audit the precision at which the output is communicated.

FigureProcess · 6 steps
  1. 1. Define units

    Specify currency, time base, measurement device, and any normalization.

  2. 2. Inspect shape

    Study skew, zeros, censoring, bounds, and heavy tails.

  3. 3. Match the loss

    Connect the objective to mean, median, quantile, or asymmetric cost.

  4. 4. Request uncertainty

    Decide whether the action needs ranges, quantiles, or calibrated distributions.

  5. 5. Test extrapolation

    Create time, geography, and range slices that challenge support.

  6. 6. Audit precision

    Round and communicate outputs at a resolution justified by measurement quality.

The same error can have different business meaning

A ten-minute early estimate and a ten-minute late estimate have equal absolute size. Their operational consequences may differ. A missed upper tail can overflow capacity even when average error looks excellent.

Direction is the part an average hides. The most famous regression failure in public health is a failure of direction rather than of magnitude. Google Flu Trends was built to predict CDC influenza-like-illness rates, and its errors were almost all one way. Lazer and colleagues reported in Science on 14 March 2014: “GFT also missed by a very large margin in the 2011–2012 flu season and has missed high for 100 out of 108 weeks starting with August 2011”. By February 2013 it was predicting more than double the CDC’s proportion of doctor visits for influenza-like illness. An independent evaluation in PLOS Computational Biology found the updated model overshot 2012–13 ILI surveillance by 268% nationally, 208% regionally and 296% locally. A mean absolute error, averaged over 108 weeks and reported as one number, would have shown a large error. It would not have shown that 100 of those weeks were high.

Report error by target range, direction, time horizon, and decision slice. Numeric performance becomes actionable only after the loss is connected to the cost of being wrong.

Regression quality is conditional on where, when, and in which direction the estimate misses.

Key takeaways