Critical Materials Atlas
Open challenge · adversarial review

This atlas has broken four of its own claims

Here is how — not as an apology, as the evidence. Zero retractions would not mean zero errors. It would mean nobody checked. Everything here is reproducible from committed code, so everything here is checkable, and four times that checking has come back and demolished a published headline of ours. Below: exactly what broke, the habit behind all four, and the claims still standing — each with the test that would kill it. Now come and break one.

This atlas cross-checks itself hard — robustness re-tests, an uncertainty Monte-Carlo, an independent second-source production check, a falsifier for every headline. But self-review has a ceiling: the author cannot referee his own work, and re-deriving a number a second way tests consistency, not the premise. So the honest next step is adversarial human review. If you know this domain — a trade economist, a minerals analyst, a supply-chain engineer — the fastest way to help is to try to break it. The record below is the argument that it is worth your time: the four claims this atlas has already demolished are its own.

The four claims this atlas has broken — and how

Judge a method by what it does when the answer is inconvenient, not by how many findings it can stack up. Findings are cheap. Every claim below was a published headline on this site; every one is now withdrawn or narrowed in place, original visible and struck, never quietly deleted. Each entry is a receipt: the claim, what actually broke it, and the page where the corrected version now lives.

1. “By-product metals are more volatile” — 37% vs 31%. withdrawn

A mean-versus-mean comparison with no control for market size — so it could not tell “supply is stuck” from “market is small”. By-product metals are ~174× smaller markets. Control for size and the effect dies (p=0.85) while size takes all the significance. Small primary metals — rare earths, tantalum, beryllium — are just as volatile, and companionality cannot explain those. The proxy was worse than the model: trade unit-value volatility tracks real price volatility at r=0.13. It was not a noisy measure of volatility; it was not a measure of it. → the retest

2. “The host decides the direction” — mean coupling 0.29, five metals beat the control. withdrawn

Three broken legs. The control was not a control (subtracting one correlation from another is not residualisation). It was never significant (at n≈21 you need |r|≈0.44; we had 0.29). And it cherry-picked — reporting the best-matching host out of two or three, then testing at 5%, which manufactures false positives at about the rate we “found” them. Done properly the coupling is 0.03. Antimony went +0.58 → −0.05: that was the commodity cycle, credited to the host. → the rewrite

3. Demand multiples were curated point estimates. corrected

Five of six sat above even the IEA’s Net-Zero scenario — magnets carried at 3.5× when the most aggressive published case is 1.9×. Recomputed from the IEA Critical Minerals Dataset, cobalt drops out of the structural-squeeze set entirely. The headline survived: gallium and germanium never leaned on the numbers that moved. → demand

4. “No clean open dataset of refining capacity by country exists.” false

It did. The IEA publishes one under CC BY. The claim was an assertion about the world made without checking it, and the dataset that disproved it also improved the page: lithium is mined in Australia (35%) and refined in China (70%). → refining

The failure mode, named

All four are the same mistake: we stopped at the first answer that agreed with us. Twice we ran a test, got the number we were hoping for, and did not ask what else could produce it — a confound we never controlled, a control that was not one. Twice we asserted something (a curated multiple, “no such dataset exists”) instead of checking whether open data already said otherwise. In every case the error survived precisely because the result was the one we wanted. That is not four unlucky findings. It is one habit, four times, and it is the most reliable way this project produces wrong answers.

It is also the most useful thing to attack. If a claim on this site flatters the thesis, that is exactly where to look — the record says our scrutiny drops when the answer is convenient.

Where else we looked — the contagion sweep

After two retractions from the same cause, the obvious question is how far it spreads. So we swept every generator. Trade unit values as a price proxy: three builders touch them — one is the retracted price test, one is the value-vs-volume page where unit values are the subject rather than a proxy, one does not use them as a price. Correlations presented as evidence: two others exist — the criticality and risk-methods rank correlations — and both are method-agreement checks (“does our ranking match the EU’s?”), not causal claims, so there is no confound to control for. Both already say they are “same family, not independent validation”.

Conclusion, offered for attack: the failure is concentrated in the arm that tried to turn correlations into mechanisms. The rest of the atlas is descriptive — who mines what, who refines it, where it flows — which fails differently and, we think, less easily. We would like to be wrong about that, and a good way to break this page is to show the sweep missed something.

What is still standing — and how to disprove each

These are the claims that have survived so far. That is a statement about effort, not about truth: the four above looked exactly this solid until someone checked properly. If any of these is wrong, a chunk of the atlas is wrong — so each one names the concrete test that would falsify it. The fastest way to help is to run one.

1. The “refiner illusion”: trade attributes supply to processors, not mines.

Break it: show a material where the reconciled trade-derived top exporter is also the genuine mine origin (so the origin gap is spurious), or show the mirror-reconciliation systematically mis-assigns direction. Test against a chain with a known mass balance.

2. By-product metals can’t scale supply to their own price (companionality).

Break it: find a metal we mark ~100% by-product (gallium, germanium, hafnium) that has responded to a price spike with new primary production on a <3-year horizon — or show the gallium mass balance (94% discarded, ~5.6% recovery) is off by enough to change the conclusion.

3. The producer geography is not a one-source artefact. narrowed

Someone ran this one. The invitation here used to read “show that agreement is circular (the two share a common source)” — and a reviewer did exactly that, arguing USGS and World Mining Data both just recompile the same national statistics. So we counted, because WMD tags every figure with its source. Across 1,903 tagged figures: national statistics 60.0%, company reports 28.5%, questionnaire 6.7%, USGS 0.8% (16 figures). The charge fails as put — WMD is not a repackaging of USGS, and 26/28 is not circular.

But it lands in a weaker form, and the claim is narrowed accordingly. Both rest on the same upstream returns, so they are independent compilations, not independent measurements: if a country misreports, both inherit it identically and agree perfectly. 26/28 now claims only compilation reliability — two teams reading the same primary returns made the same call. Break it further: show that the 6.4pp mean gap hides a material that actually flips, or that WMD's source tags misstate its own provenance.

4. Only gallium and germanium are “hardest cases” — and even that is a gradient. strengthened + narrowed

Someone ran both invitations here. This entry asked readers to argue the input spreads were mis-specified, or that the five axes double-count one underlying factor. A reviewer went further and named the real flaw: a Monte-Carlo propagates input uncertainty and cannot validate the model. Ours drew the five inputs but held the design sacred — chosen axes, hard thresholds, a fixed ≥4-of-5 rule — so “gallium 99%” only ever meant “99% of draws under this design”. Fair, and it stuck.

So we perturbed the design across 11 specifications — drawing the thresholds instead of fixing them, moving them ±15%, changing the rule, and dropping each axis in turn. The number died: gallium spans 49–100%, and we no longer report 99% as if it were a probability. The finding got stronger: gallium and germanium are top-2 in 11 of 11, including every single-axis ablation. And the double-count is real — companionality vs market size correlate at r=−0.70, so “can’t scale” and “thin market” are largely one factor (the volatility retest found the same thing independently). It inflates the contingent middle; it does not move the top two. Break it further: find a design where gallium is not top-2.

5. Concentration is not the binding constraint — elasticity and market thinness are.

Break it: show a material that is geographically concentrated and genuinely stuck despite being primary, recyclable, and large-market — i.e. a case where single-country dominance alone is the real problem, contra the scorecard.

Soft spots the atlas already flags

Not hidden — stated, and worth pressing on hardest.

Data & method

  • Trade value is a noisy proxy — nominal, re-export-contaminated, HS-misclassified, weak for by-products. The core engine is the most attackable link.
  • Companionality & demand are point estimates from literature rounds, dressed in interactive sliders that imply more precision than they have.
  • HS-code aliasing — gallium/germanium/hafnium share one code; other silent bundles may exist.

Epistemics

  • Internal cross-checks are consistency, not validation. Re-deriving a number a second way proves the arithmetic, not the premise. Treat every “cross-checked” note as such.
  • Scrutiny drops when the answer is convenient. All four retractions below share it: we stopped at the first result that agreed with us. If a claim here flatters the thesis, look there first.
  • Breadth over depth. Many layers share one spine; some re-test the same story. Press on whether any single result has independent bite.
  • No external human review yet. That is what this page is trying to fix.

How to challenge it

Open an issue on the public engine repository — it is anonymous to raise, and everything is versioned, so a fix is visible in the commit history.

Open an issue → Read the full limitations

A useful challenge names the page, the specific number or claim, and the test that would settle it. A template:

Page:        e.g. gallium.html
Claim:       the specific sentence / number you dispute
Why wrong:   the mechanism or counter-evidence
How to test: the check that would settle it (a source, a mass balance, a recompute)
Source:      any public data that supports your point

The commitment

Every substantive challenge gets one of three honest outcomes, in public: fixed (the code and pages change, logged in updates), bounded (acknowledged as a real limitation and added to limitations), or rebutted (with the reasoning shown). No silent deletions. The point of building this in the open was never to look finished — it was to be correctable.