Causality / Foundations
verifiedCausal Inference
Prediction asks what usually happens together. Causal inference asks what would happen if you changed something — a different question, needing more than the data. No amount of observation settles it on its own, because the same numbers are consistent with several stories about what produces what, and choosing between them takes an assumption you bring rather than one you measure.
This is the gap under most deployed models. A model trained on logged behaviour predicts the world it was logged in, and the moment it is used to decide something it changes that world — a recommender that learns what people click then determines what they see. The assumptions doing the work are always about what was left unmeasured, which is why causal claims are argued rather than computed.
Association is p(y|x); causation is p(y|do(x)). They differ whenever a common cause feeds both, and they coincide when x is randomised — which is what randomisation buys, and it buys nothing about the parts you did not randomise. Recovering the second from observational data needs an identification argument, a proof that the causal quantity can be written in terms of observable ones under the stated assumptions.
assumed holds 50% of the budget; rest holds the remaining 50%.
What the identification argument assumes against what the data supplies, in equal units. Drag the assumptions up to watch the answer come to rest on things nobody measured.
Reviewed by opendroid · 2026-08-18
- arXiv:2002.02770 — A Survey on Causal Inference
- arXiv:2102.11107 — Towards Causal Representation Learning
Origin · not linkable
- Pearl 1995 — Causal Diagrams for Empirical Research · Biometrika 82(4) · doi:10.1093/biomet/82.4.669
- Rubin 1974 — Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies · Journal of Educational Psychology 66(5) · doi:10.1037/h0037350