Mathematical foundations
Joint, Marginal, and Conditional Probability
Learn joint distributions, marginalization, conditioning, independence, conditional independence, and the chain rule of probability.
By the end you can
- Move between joint, marginal, and conditional distributions
- Distinguish independence from conditional independence and zero correlation
- Factor a joint distribution using the probability chain rule
- Detect selection and conditioning effects that can reverse apparent relationships
Example
Conditioning can create an association
A filter can manufacture a relationship out of nothing, and the size of the effect has been measured rather than assumed. Sackett measured it in 1979. He took household interviews of a random general-population sample — 2,784 people — and computed the relative odds between respiratory disease and disease of the bones and organs of movement. The answer was 1.06, which is to say nothing at all. He then restricted the same interviews to the 257 of those people who had been in hospital in the previous six months. Inside that subgroup the cross-product odds were 4.06.
Oxford's Catalogue of Bias records the result: “He analysed data from 257 hospitalized individuals and detected an association between locomotor disease and respiratory disease (odds ratio 4.06).”
Same people. Same two diseases. Same questionnaire. The only thing added was a filter, and the filter is what produced the fourfold association. This is Berkson's bias, and admission is the collider being conditioned on.
- Population: In the 2,784 household interviews, respiratory disease and disease of the bones and organs of movement carried relative odds of 1.06.
- Conditioning: Keep only the 257 of those people who had been in hospital in the previous six months.
- Observed pattern: Within that hospitalised subgroup the cross-product odds are 4.06 — a fourfold association where the population showed 1.06.
- Mechanism: Admission is a common effect of both conditions, sometimes called a collider; either illness can put a person in the sample.
- Lesson: Conditioning on a selection variable can produce a relationship that is absent in the full population, and 1.06 against 4.06 is how large the manufactured effect can be.
One variable rarely tells the whole story
A joint distribution describes several random variables together. Sum or integrate a variable out and you get a marginal distribution. Restrict attention to what is already known and you get a conditional one. These operations change the question. P(Y) describes overall frequency; P(Y|X=x) describes frequency within a specified context. Machine learning leans hard on conditional distributions, because prediction asks how uncertainty about a target changes after observing features.
The two answers are not obliged to agree. The rest of this lesson works through files where they did not: graduate admissions counts in which the marginal favours one group and four of six conditionals favour the other, a surgical series in which the marginal and both strata point opposite ways, and a sample of 257 hospital patients in which a filter alone moved an odds ratio from 1.06 to 4.06.
Marginalization removes information; conditioning adds information to the question.
Case
Berkeley, 1973: 12,763 applications, 4,526 in the table everyone uses, and a gap that changes sign
One admissions file made this concrete. The University of California, Berkeley received 12,763 applications for graduate study in 1973: 8,442 from men and 4,321 from women. Of the men, 3,738 were admitted, or 44 percent. Of the women, 1,494 were admitted, or 35 percent. That is a gap of nine percentage points, and it reads like one fact about one university. Three authors took the same file apart department by department and published the result in Science in 1975. The gap largely disappeared. Women had applied more often to departments that admitted fewer of everyone. A 2025 re-analysis finds the same 12,763 applicants spread across 101 departments, the men at approximately 44.2 percent and the women at approximately 34.6 percent.
The table almost every reader actually meets is a smaller one. R ships the six largest departments only, as the built-in UCBAdmissions dataset: 4,526 applicants, not 12,763. Its own documentation states the margin — “There were 2691 male applicants, of whom 1198 (44.5%) were admitted, compared with 1835 female applicants of whom 557 (30.4%) were admitted.” It reports “a sample odds ratio of 1.83”; the exact cross-product ratio of the four cells is 1.84.
Now condition on department. The aggregate gap does not merely shrink. In four of the six departments it changes sign. Department A admitted 62.06% of male applicants and 82.41% of female applicants. B, 63.04% against 68.00%. D, 33.09% against 34.93%. F, 5.90% against 7.04%. Only C, 36.92% against 34.06%, and E, 27.75% against 23.92%, run the other way. One array of counts, two summaries that contradict each other, and nothing in dispute except which distribution is being reported. 44.5% against 30.4% is a true marginal statement. 62.06% against 82.41% is a true conditional statement. Neither is a correction of the other.
Figure
Visual
Three views of the same joint model
A two-variable joint distribution supports several related calculations, and the six-department Berkeley table shows all of them on one object. The joint P(sex, department, admission) is the array of counts over 4,526 applicants. Marginalize department away and you get 1198 of 2691 men admitted, 44.5%, against 557 of 1835 women, 30.4%. Condition on department A instead, renormalize inside it, and you get 62.06% against 82.41%. Reverse the conditioning — ask what mix of departments a rejected applicant applied to — and Bayes' rule takes you back the other way through the same counts. No number was recomputed between those three statements. Only the denominator changed.
- 1
Joint P(X,Y)
Represents how X and Y vary together.
- 2
Marginal P(X)
Sum or integrate over Y to ignore it.
- 3
Conditional P(Y|X)
Renormalize the joint within a known value of X.
- 4
Reverse conditional P(X|Y)
Use Bayes' rule to update the other direction.
The denominator in conditioning turns a restricted slice of the joint distribution into a valid distribution.
Key idea
Conditioning on the wrong variable can answer the wrong question
In observational data, a feature measured after an intervention may be affected by both the intervention and the outcome process. Conditioning on it can block or create associations. Prediction models may legitimately use post-event variables when those variables are available at prediction time. Causal interpretations require a different analysis. So always state whether a conditional relationship is predictive, descriptive, or intended to support a causal claim.
How far the choice can move an answer was measured in 6,746 children across two Ghanaian hospitals. Krumkamp and colleagues published it in 2016. The sign of the malaria–invasive nontyphoidal Salmonella (iNTS) association depends entirely on which control group the analysis conditions on. Set the 6,301 children with no febrile bloodstream infection as controls and malaria looks protective: OR 0.4, 95% CI 0.3–0.7. Set the 285 children carrying non-iNTS bacteraemia as controls and malaria becomes a risk factor: OR 1.9, 95% CI 1.1–3.3. The 160 iNTS cases are the same children in both analyses.
Oxford's Catalogue of Bias summarises the first of the two: “In the first study, children with salmonella infection were classified as cases, and controls were uninfected. A protective association between malaria and salmonella infection was found: pooled OR = 0.4.”
Protective and harmful, from one dataset, decided by the conditioning set.
A valid conditional probability can still be irrelevant—or misleading—for the decision you meant to analyze.
Case
Pearson 1899, Yule 1903, Simpson 1951: the effect was older than the name, and the dates are disputed
The name arrived after the effect, and the authorities do not agree on when. The Stanford Encyclopedia of Philosophy puts it this way: “This phenomenon was first pointed out in papers by Karl G. Pearson (1899) and George U. Yule (1903), but it was Simpson’s short paper “The interpretation of interaction in contingency tables” (1951), discussing the interpretation of such association reversals, that led to the phenomenon being labeled as “Simpson’s Paradox”.” Sprenger and Weinberger wrote that entry.
A 2020 paper in the Journal de la Société Française de Statistique dates the early accounts to Yule 1900 and Pearson 1900 instead, and credits Colin Blyth, in 1972, with coining the phrase “Simpson's paradox”. A 2014 survey by P. Vellaisamy records that statisticians were aware of the issue at the beginning of the twentieth century, and that it was Simpson who popularized the paradox, earning his name. Some writers say Yule–Simpson for that reason.
The mathematics was never in doubt. The attribution is. Two respectable sources will hand a reader 1899 or 1900 for the same paper.
The probability chain rule factorizes any joint distribution
For variables X₁,…,Xₙ, the joint can be written as P(X₁)P(X₂|X₁)…P(Xₙ|X₁,…,Xₙ₋₁). This identity requires no independence assumption. Independence assumptions simplify the factors. A graphical model encodes which conditioning variables are retained and which are omitted, so a factorization is both a computational device and a claim about dependence structure.
The Berkeley table is a three-variable instance. P(sex, department, admission) factorizes as P(sex) P(department | sex) P(admission | sex, department), with no assumption at all. The middle factor is the one that carried the aggregate result: women applied more often to departments that admitted fewer of everyone. Drop it — assert that department is independent of sex — and the model predicts a single admission rate for each sex, which is the 44.5% against 30.4% margin. Keep it, and the model reproduces 62.06% against 82.41% in A and 5.90% against 7.04% in F. Choosing which conditioning variables to omit is choosing which of those two answers your model can express.
Analogy
Marginalizing and conditioning in a library catalog
A catalog contains every book. Marginalizing over genre counts books by language while ignoring genre. Conditioning on mystery keeps only mystery books and renormalizes the proportions. Independence means the language mix is the same in every genre. Conditional independence means a relationship disappears after filtering by another attribute. The Berkeley gap of 44.5% against 30.4% shrinks, and in four departments reverses, once department is fixed — which is what a strong dependence between sex and department looks like from inside the catalog.
Selection is a different problem. When observations affect which books enter the catalog at all, the joint distribution itself has changed. Sackett's 257 hospitalised interviewees are a catalog assembled by the very variable under study, and inside it 1.06 reads as 4.06.
Conditioning is a change of population, not merely a change of notation.
Comparison
Independence, conditional independence, and uncorrelatedness
These terms express different levels of separation, and a surgical series shows why the distinction is not academic. A 1986 report in the BMJ compared four treatments for kidney stones: “Success was achieved in 273 (78%) patients after open surgery, 289 (83%) after percutaneous nephrolithotomy, 301 (92%) after ESWL, and 15 (62%) after percutaneous nephrolithotomy and ESWL.” Marginally, then, percutaneous nephrolithotomy removed kidney stones more often than open surgery: 289 of 350 against 273 of 350, 83% against 78%.
Condition on stone diameter and the ordering inverts in every stratum. In the stratified table reproduced from Julious and Mullee, open surgery succeeded in 81 of 87 cases (93%) against 234 of 270 (87%) for stones under 2 cm, and in 192 of 263 (73%) against 55 of 80 (69%) for stones of 2 cm or more. Open surgery wins in both strata and loses overall.
Marginal independence, conditional independence and zero covariance are three separate statements about the same joint distribution. 78 against 83, reversing to 93 against 87 and 73 against 69, is what it costs to confuse them.
Independence
P(X,Y)=P(X)P(Y) for all relevant values.
- Knowing X does not change the distribution of Y
- Implies zero covariance when moments exist
- Strong distributional statement
- Often unrealistic without careful design
Conditional independence
X and Y become independent after conditioning on Z.
- Written X ⟂ Y | Z
- Supports graphical-model factorization
- Can be created or destroyed by conditioning
- Central to causal and probabilistic reasoning
Uncorrelated
Cov(X,Y)=0.
- Only rules out linear association
- Does not generally imply independence
- Equivalent to independence for some Gaussian settings
- Can hide strong nonlinear dependence
Steps
A dependence audit
Before simplifying a joint model, work through the sequence below: draw the variables, state the joint question, identify what is being conditioned on, test the proposed independences, and check the selection effects last.
The last step is the one with a published benchmark. UK Biobank is a selected sample by construction. Approximately 9.2 million people aged 40–69, living within 25 miles of one of 22 assessment centres in England, Wales and Scotland, were invited. Of those, 5.5% took part, giving a cohort of about 500,000. Set that against surveys with conventional response rates and the difference is stark — “Analytical sample of 499 701 people (response rate 5.5%) in analyses in UK Biobank; pooled data from the Health Surveys for England (HSE) and the Scottish Health Surveys (SHS), including 18 studies and 89 895 people (mean response rate 68%).” That comparison appeared in the BMJ in 2020.
A separate study in the American Journal of Epidemiology describes who the 5.5% are: more likely to live in less socioeconomically deprived areas, less likely to be obese, to smoke or to drink daily, and carrying fewer self-reported health conditions than the general population. A joint distribution estimated on that cohort is the joint distribution of the people who agreed to join it. That is the same structural fact as Sackett's 257 patients, at a scale of half a million.
1. Draw the variables
List targets, features, selection mechanisms, and timing.
2. State the joint question
Specify the population and variables under consideration.
3. Identify conditioning
Mark which information is known or used to filter data.
4. Test proposed independences
Compare conditional distributions and domain mechanisms.
5. Check selection effects
Ask whether conditioning opens or closes important paths.
Key takeaways
- Joint distributions represent variables together, while marginalization removes variables and conditioning changes the information set: 1198 of 2691 men and 557 of 1835 women, 44.5% against 30.4%, is a margin; 62.06% against 82.41% in Berkeley's department A is a conditional, and both are true of one table.
- Independence is stronger than zero correlation and must hold across the full joint distribution; zero covariance rules out only linear association.
- Conditional independence can differ sharply from marginal independence — the 1986 kidney-stone series' 78% against 83% reverses to 93% against 87% and 73% against 69% once stone diameter is fixed.
- The probability chain rule factorizes any joint distribution without assumptions; P(sex) P(department | sex) P(admission | sex, department) reproduces the Berkeley departments, and dropping the middle factor is what collapses them to a single 44.5%-against-30.4% claim.
- Conditioning on selection variables or common effects can create misleading associations: Sackett's 2,784 household interviews gave relative odds of 1.06, and the 257 of them admitted to hospital gave 4.06.
- Predictive conditional relationships should not be given causal meaning without a causal design and temporal audit — one Ghanaian hospital dataset gave OR 0.4 (95% CI 0.3–0.7) and OR 1.9 (95% CI 1.1–3.3) from the same 160 iNTS cases, purely by changing the control group.