Skip to content
AI.info

Research

A note on the area under the likelihood and the fake evidence for model selection

Overview Research area: Bayesian statistics, specifically model selection (Level-2 inference), the marginal likelihood / Bayesian evidence, and the use of improper and diffuse priors. Technical level:

A note on the area under the likelihood and the fake evidence for model selection
arXiv
2602.22965
Published
2026-02-26
Authors
L. Martino, F. Llorente

AI summary

Overview

Research area: Bayesian statistics, specifically model selection (Level-2 inference), the marginal likelihood / Bayesian evidence, and the use of improper and diffuse priors.

Technical level: Advanced. The paper is written for readers already comfortable with Bayesian notation, posterior distributions, marginal likelihoods, Bayes factors, and generalized linear models.

Scope: A short theoretical note (arXiv:2602.22965v1 [stat.ME], 26 Feb 2026, by L. Martino and F. Llorente, affiliated with Universita' di Catania, Catania, Italia, and Stony Brook University, New York, USA) arguing that improper priors can legitimately be used for one narrow kind of model selection, provided the resulting quantity is renamed a "fake evidence" or "area under the likelihood" rather than a Bayesian evidence.

What This Paper Is About

Bayesian model selection normally requires computing the Bayesian evidence (marginal likelihood) Z = p(y), an integral of the likelihood times the prior over the parameters. Improper priors, whose integral is infinite, are generally considered forbidden here because they leave Z specified only up to an arbitrary constant c, so Z = c Z_{\ell \times h}. The paper's goal is to clarify this rule, show that improper priors can still be used when comparing models that belong to the same parametric family (Type-2 model selection, i.e., tuning parameters or hyper-parameters of a parametric model), and warn that the quantity computed in that case is not a true evidence and cannot be recovered as a limiting case of a diffuse prior.

Key Contributions

  1. Terminology and clarification. The authors propose calling the quantity obtained with a non-uniform improper prior the "fake evidence" (Z_{\ell \times h}) and, in the special case of a uniform improper prior, "the area under the likelihood" S = ∫ ℓ(y|θ) dθ. They argue the common statement "improper priors are not allowed for model selection" is technically wrong; the correct statement is that improper priors are not allowed for computing the Bayesian evidence Z.

  2. A precise exception to the prohibition. The paper shows that in Type-2 model selection (several, possibly infinite, models belonging to the same parametric family), using the same improper prior for all models makes the arbitrary constant c cancel out in the Bayes factor, in the model-averaging weights, and in the pointwise estimator of the tuning parameter. So fake evidences are usable there.

  3. A negative result about diffuse priors. The authors show that the area under the likelihood S cannot be recovered by starting from a proper diffuse prior and increasing its scale parameter to infinity, unlike what typically happens at Level-1 of inference. As the prior spreads, Z → 0, hence Z ↛ S, since S > 0.

  4. Worked application to Bayesian regression. They develop the complete analysis for a generalized linear regression model with nonlinear bases under two priors, a uniform improper prior and a Gaussian prior, including full theoretical details and a numerical experiment (Section 6 of the paper) that confirms the statements. They also compare the approach with profile likelihood and with a full-Bayesian solution using a double (uniform) improper prior.

Main Findings

  • Improper priors leave Z undetermined: If g(θ) = c · h(θ) with ∫ h(θ) dθ = ∞, then Z = c Z_{\ell \times h}, where c > 0 is arbitrary. Hence Z is not completely specified and improper priors cannot be used to compute the Bayesian evidence.

  • Level-1 vs Level-2 roles of diffuse priors: At Level-1 (parameter estimation), diffuse priors are weakly informative and, as the support volume |B| → ∞, the posterior converges to the normalized likelihood ℓ(y|θ)/S, recovering a non-informative choice. At Level-2 (model selection), diffuse priors are strongly informative: if S is finite, Z = ∫_B ℓ(y|θ) dθ / |B| and Z → 0 as |B| → ∞, so a good model can receive a low Z simply because its prior was spread out, and a worse model can receive a higher Z with a concentrated prior.

  • Improper priors remain valid at Level-1 when Z_{\ell \times h} is finite: If Z_{\ell \times h} = ∫ ℓ(y|θ) h(θ) dθ < ∞, the posterior is still proper, since the arbitrary constant cancels in the normalized posterior.

  • The exception for Type-2 model selection: When the same type of improper prior is used for all compared models (g(θ|η_{p,i}) = c · h(θ|η_{p,i})), the Bayes factor equals the ratio of fake evidences, the model-averaging weights are well-defined, and the optimal tuning parameter can be written as η* = arg max ∫ ℓ(y|θ, η) dθ. The constant c must not be absorbed into the tuning vector η_p.

  • Nested models (Type-3) are more restrictive: In the nested-model scenario, an improper prior is allowed only for the first variable/parameter that is common and shared by all nested models.

  • The area under the likelihood S is more limited than the general fake evidence: Because h(θ) = 1 admits only a multiplicative constant, S can be used only for tuning parameters of the likelihood function and not for prior hyper-parameters.

  • S cannot be obtained as a limit of a proper prior: With S < ∞ and S > 0, increasing the scale of a diffuse prior drives Z → 0, so Z does not tend to S.

  • Interpretation of the tuning estimator: Maximizing S(y|η) = ∫ ℓ(y|θ, η) dθ with respect to η amounts to a frequentist treatment of η after integrating out θ, i.e., a Bayesian-frequentist combination related to empirical Bayes; the paper contrasts this with profile likelihood, which involves no integration and no explicit or implicit prior density.

  • Full-Bayesian alternative: If S_Z = ∫ S(y|η) dη < ∞, an improper uniform prior can be placed on η as well, giving p(η|y) = (1/S_Z) S(y|η); the procedure then draws η_1, …, η_R from p(η|y) and, for each r, θ_{r,1}, …, θ_{r,N} from p(θ|y, η_r) ∝ ℓ(y|θ, η_r), producing N·R samples. This can be read as a continuous Bayesian model averaging with ℓ̄(y|θ) = ∫ ℓ(y|θ, η) p(η|y) dη, where p(η|y) acts as an informative, data-dependent prior on η.

Methodology in Plain English

The authors proceed analytically rather than experimentally for the core argument. They set up standard Bayesian notation (likelihood ℓ(y|θ, M), prior g(θ|M), posterior π̄(θ|y, M), evidence Z), define a likelihood that is integrable over the whole parameter space so that its total area S is finite, and then compare what happens at Level-1 and Level-2 as a uniform prior on a bounded region B is allowed to grow until it covers the entire (unbounded) parameter space.

They then separate three model-selection scenarios: Type-1 (comparing genuinely different likelihoods), Type-2 (comparing models in the same parametric family, i.e., tuning parameters, which is the empirical Bayes setting), and Type-3 (nested models of increasing dimension, as in variable selection, order selection, or clustering with an unknown number of clusters). By tracking where the arbitrary prior constant c appears and cancels, they identify exactly which quantities remain well-defined under an improper prior.

To make the general claims concrete, they work out a Bayesian generalized linear regression model with N data points and M nonlinear bases (with M ≤ N, the case M = N being non-parametric; M > N is explicitly not considered). They derive the Gaussian likelihood for the coefficient vector θ with noise power σ_e², and then carry out the analysis twice: once with a uniform improper prior on θ and once with a Gaussian prior. This regression setup also supplies the theoretical support for a numerical experiment that checks the statements (the numerical results themselves fall in the truncated portion of the text provided here and are therefore not reported in this summary).

Why This Matters

Impact on research: The paper targets a persistent source of confusion in the applied Bayesian literature: which quantities are legitimate evidences and which are not. It clarifies that fake evidences and the area under the likelihood remain usable for tuning parameters while being invalid as Bayesian evidences, and it corrects the over-broad claim that improper priors are never allowed in model selection. It also warns that diffuse priors, often treated as neutral, actively penalize models at Level-2.

Real-world applications (fields cited in the paper as users of Bayesian methods):

  • Remote sensing.
  • Astronomy.
  • Cosmology.
  • Optical spectroscopy.

Within those and other fields, the specific model-selection problems the paper connects to include variable selection, order selection (for example in polynomial regression or ARMA models), and clustering when the number of clusters is unknown, along with dimension reduction.

Industry relevance: Any workflow that tunes hyper-parameters, basis functions, or model complexity by maximizing a marginal likelihood will be affected by whether that quantity is a true evidence or a fake evidence, and by how diffuse the prior is. Practitioners comparing models through Bayes factors, Bayesian model averaging weights, or empirical Bayes tuning need to know that the arbitrary prior constant cancels only in the Type-2 setting, and that increasing a prior's scale parameter does not converge to the area under the likelihood.

Future Directions

  • The case M > N. The paper's regression analysis assumes M ≤ N and explicitly excludes M > N, stating it requires specific observations and separate analysis. Extending the fake-evidence argument to that regime is an open step.
  • More than two levels of inference. The authors note that hierarchical Bayesian approaches involve additional levels beyond Level-1 (estimation and prediction) and Level-2 (model selection), but they focus mainly on Level-2; how the fake-evidence reasoning behaves in hierarchical settings is left open.
  • Tractable versus intractable marginal likelihoods. The authors state their theoretical and practical considerations hold both when the marginal likelihood can be computed analytically and when it is intractable; in the latter case only the extra computational problem of approximating the marginal likelihood remains, which points to further work on approximation in this setting.
  • Beyond uniform h(θ) and beyond the double-improper full-Bayesian scheme. The paper presents results for h(θ) = 1 and extends remarks to generic h(θ), and it sketches a full-Bayesian solution using improper priors twice; generalizing and stress-testing these constructions in applied models is a natural continuation.

Target Audience

Statisticians and methodologists working on Bayesian model selection and marginal likelihood computation; researchers in the applied fields listed above (remote sensing, astronomy, cosmology, optical spectroscopy) who choose priors and compare models; and graduate students or practitioners using empirical Bayes, Bayes factors, or Bayesian model averaging who need to know when an improper prior is defensible and when the resulting number is not a true Bayesian evidence.

Authors’ abstract

Improper priors are not allowed for the computation of the Bayesian evidence $Z=p({\bf y})$ (a.k.a., marginal likelihood), since in this case $Z$ is not completely specified due to an arbitrary constant involved in the computation. However, in this work, we remark that they can be employed in a specific type of model selection problem: when we have several (possibly infinite) models belonging to the same parametric family (i.e., for tuning parameters of a parametric model). However, the quantities involved in this type of selection cannot be considered as Bayesian evidences: we suggest to use the name ``fake evidences'' (or ``areas under the likelihood'' in the case of uniform improper priors). We also show that, in this model selection scenario, using a diffuse prior and increasing its scale parameter asymptotically to infinity, we cannot recover the value of the area under the likelihood, obtained with a uniform improper prior. We first discuss it from a general point of view. Then we provide, as an applicative example, all the details for Bayesian regression models with nonlinear bases, considering two cases: the use of a uniform improper prior and the use of a Gaussian prior, respectively. A numerical experiment is also provided confirming and checking all the previous statements.

Read the original paper