Critical Materials Atlas
Rigor · uncertainty propagation

How robust is that?

The hardest-cases scorecard draws hard lines — and a fair critic says a metal sitting just under a cutoff is hidden by pass/fail. So don’t draw the line once. Put an honest uncertainty on every input and run it 20,000 times: the binary cliff becomes a probability gradient, and the “only gallium and germanium” headline gets the soft edges it deserves.

The uncertainty model — the load-bearing assumption, stated

For each material we draw each axis input from a distribution and recompute the five failure flags, 20,000 times: companionality ~ N(mean, 12pp); top-producer share ~ N(mean, 8pp) (the cross-source spread measured at 6.4pp on this data); recycling ~ N(mean, 5pp); demand growth ~ mean×exp(N(0, 0.40)), because forward demand is deeply scenario-dependent; world tonnage ~ mean ±~10%. We report, per material, the probability each axis fails, the expected number of failures with a 90% interval, and P(hardest) = P(≥4 of 5 fail).

Honest about the honesty: those five spreads are themselves judgement calls — a wider demand sigma pulls more metals into contention, a tighter one sharpens the top. The point is not a precise probability but to show the ranking’s shape and which results survive perturbation. Deterministic (fixed seed). Inputs: the five layer JSONs → uncertainty.json.

P(hardest) — probability a material fails 4 or 5 of the 5 axes

Not a cliff, a gradient. The bracket marks the 90% interval on the number of axes failed. Two materials are almost certain; a cluster behind them is genuinely borderline — which the binary scorecard erased.

Per-axis failure probability

Shaded by probability, not on/off. Reading down the columns shows what actually drives vulnerability.

The objection this page could not answer — so we tested it

“A Monte-Carlo propagates input uncertainty. It cannot validate the model.” That is the sharpest thing anyone has said about this page, and it is correct. Everything above draws the five inputs — but the design was held sacred: the five axes are chosen rather than derived, the thresholds (66, 60, 10, 2.5, 50 kt) are hard constants that were never drawn, and “≥4 of 5” is a constant too. So “gallium 99%” only ever meant “99% of draws, under this design”. Reported as though it were the probability gallium is truly the hardest case, it launders a designer’s choices into false precision.

So we perturbed the design as well: drew the thresholds instead of fixing them, moved them ±15%, changed the rule to 3-of-5 and 5-of-5, and dropped each axis in turn — holding the bar at “fails all but one” so an ablation tests the axis rather than smuggling in a stricter rule. Eleven specifications. The results pull in opposite directions, and both belong on the page.

The axes are not independent

The scorecard counts five axes as if each were a separate way to be stuck. They are not.

What propagation changes