| Takeaway | Detail |
|---|---|
| The 20% holdout is unverified. | The figure appears in the supplied article title; retrieved sources identify no rare-earth locality, dataset, spatial split procedure, cross-validation design, or holdout result. |
| A 20% holdout can hide a blind spot. | Withheld terrain may be unrepresentative of the training area, allowing favorable aggregate performance to conceal failures in the area relevant to a drilling decision. |
| The 20% comparison is unsubstantiated. | A matched benchmark on the same rare-earth data, with shared predictors, spatial folds, and evaluation thresholds, is absent from the retrieved evidence. |
| A 78% figure is not a drill permit. | The supplied research does not identify what 78% measures or connect it to rare-earth performance; even a forest validated on a 20% spatial holdout for prospectivity cannot inherit a separate grade model's drilling recommendation. |
A 20% spatial holdout appears in a supplied article title, but the retrieved mindat.org page shows only a truncated citation and a security-verification warning—not a completed rare-earth study. No documented split, dataset, or result supports treating that number as evidence of drilling readiness.
Geological representativeness, not merely a clean evaluation design, determines whether withholding is informative. Training samples can concentrate in familiar terrain and leave the withheld area unlike the evidence used for fitting. Random forests may legitimately outperform conditional simulation, but no retrieved benchmark holds predictors, spatial folds, and thresholds constant on the same rare-earth data. The evidence therefore leaves model selection open rather than establishing a winner.
Geological plausibility is a modeling prior, not a drill permit. A forest trained to predict prospectivity cannot inherit a drill recommendation from a separate grade model's scorecard, even if the 20% holdout looks strong. Verification should examine the spatial split procedure, withheld-area performance, calibration, and independent ground truth before any drilling or investment commitment.

20% Held Out
A 20% spatial holdout is a locked test, not an accuracy certificate. I would define the withheld geography before fitting either model and record the withheld-cell geometry. A reserved share can still be poorly placed—near training assays or within a familiar geological regime—without providing a meaningful test of transfer to new ground.
Both models would estimate the same continuous quantity: total REO grade within eligible drill cells, under one laboratory convention, mineralogical basis, and spatial support. Conditional simulation would then face the same regression task as random forest, rather than a deposit/no-deposit classifier. A classifier’s output is a different estimand; it cannot settle whether either method predicts grade.
According to the spatial-estimation framework of Pyrcz and Dubrule (2021), I would estimate a training-only REO variogram, condition the grade field on drilling and geological controls, and generate multiple kriging-based conditional Gaussian realizations. These realizations are conditional predictions, not a deterministic kriging surface with cosmetic noise. I would also document whether the chosen Gaussian representation is appropriate for the grade distribution; simulation does not repair a poorly specified marginal model.
Random-forest regression would use the same training assays, coordinates, geological variables, and prediction support. I would reserve a separate training-only partition for tuning and predictive-interval calibration, leaving the spatial holdout untouched until adjudication. Tree-vote proportions or raw forest dispersion are not automatically calibrated REO-grade uncertainty; their spread cannot establish that an interval meets the required empirical coverage of withheld assays.
I would require every headline comparison to name its target quantity rather than blur these distinctions into a generic prospectivity score:
| Quantity evaluated | What it measures | Valid comparison |
|---|---|---|
| Predicted REO grade | Mean grade within a defined, supported cell | Withheld continuous-grade MAE and prediction-interval coverage |
| Probability of exceeding the economic cutoff | Likelihood that a supported cell meets the cutoff | A probability metric, not continuous-grade MAE |
| Estimated contained metal | Metal amount under volume and density assumptions | An inventory metric, not a grade residual |
| Probability of an undiscovered deposit | Deposit likelihood beyond sampled support | A discovery assessment, not grade-prediction accuracy |
Map support must be reconciled before either result is interpreted. A bulk-block REO average, a channel-sample average, and a narrow core intercept are different observations. I would convert them using training information where defensible; otherwise, I would restrict the claim to the support actually represented. Without that reconciliation, improved agreement on withheld sample grades would not establish accuracy for block-scale grade or resource estimates.
This distinction matters because the supplied source set contains no rare-earth comparison holding predictors, spatial folds, and evaluation thresholds constant while benchmarking a conditional simulator against random-forest regression. The supplied snippet for Deng et al.’s conditional-random-field model gives no realized held-out result or rare-earth validation score. I therefore would not treat a bibliographic connection to prospectivity as evidence that this drill gate has been passed.
My release checklist is therefore: identical grade target, reconciled support, training-only variogram and forest development, separate interval calibration, and one untouched spatial adjudication. I would release the grade simulation as drill-eligible only if it yields lower withheld REO MAE than random forest and passes the article’s prescribed empirical prediction-interval coverage gate. If either gate fails, I would skip the drill target rather than promote dispersion, a cutoff map, or a deposit probability as a substitute.

30 ppm La and Production Are Not Predictive
A real number can still be the wrong kind of evidence. A continental-crust mean characterizes composition; a national production total characterizes output. Neither establishes a local REO forecast’s out-of-sample reliability. That distinction is an evidentiary gate, not a semantic caveat.
According to Rudnick and Gao’s “Composition of the Continental Crust: A New Approach,” in Treatise on Geochemistry, Volume 3, the benchmarks tabulated below describe background continental composition. They are not economic cutoffs or validated deposit probabilities. Averaging heterogeneous crust neither locates residual enrichment nor makes proximity to that mean an economic test. According to the USGS, Mineral Commodity Summaries 2024, “Rare Earths,” the production figure records aggregate national output, not a coordinate-linked grade observation. It cannot supply withheld training labels or substitute for local out-of-sample assay error. Either substitution would make the evaluation circular.
According to Ploton et al.’s “Spatial Validation Reveals Poor Predictive Performance of Large-Scale Ecological Mapping Models,” in Nature Communications 11, their ten-dataset African mammal comparison shows why random splitting can fail spatial evaluation. Shared spatial structure can appear on both sides of a random split, making prediction appear more transferable than it is under geographic withholding. That mechanism justifies a spatial-validation design check; it does not transfer an ecological AUC into an REO regression comparison.
Breiman’s Random Forests paper in Machine Learning 22 supplies a different diagnostic. Under ordinary bootstrap resampling, the asymptotic probability that a given observation is omitted is 1−e−1, approximately 36.8%. Bootstrap omission is not geographic exclusion: an out-of-bag observation can still have close neighbors represented in training. Neither this fraction nor out-of-bag error supplies a spatial holdout or establishes calibrated REO prediction intervals.
These sources provide crustal benchmarks, supply accounting, ecological validation, and a generic machine-learning diagnostic—not a published REE head-to-head result. For a defensible claim, name the REE assay dataset, withholding geometry, random-forest implementation, and interval calculation. Conditional simulation earns drill-eligible status only if it lowers withheld REO mean absolute error against random forests and passes the coverage requirement below. If either gate fails, skip.
| Published source or required audit | Source-specific figure | What it measures | Decision consequence |
|---|---|---|---|
| Rudnick and Gao | 30 ppm La; 66.5 ppm Ce | Mean continental-crust composition | Not an economic cutoff or deposit probability |
| USGS, Mineral Commodity Summaries 2024, “Rare Earths” | Reported United States mine production, 2023 | Aggregate supply output | Not withheld training labels or assay-error evidence |
| Ploton et al. (2020) | Ten African mammal datasets | Spatial-validation methodology | No ecological AUC transfer into REE evaluation |
| Breiman | 1−e−1 ≈ 36.8% out-of-bag omission probability | Ordinary bootstrap diagnostics | No spatial holdout or calibrated grade interval |
| Required named REE assay comparison—not a published benchmark | Nominal prediction intervals and a preregistered empirical-coverage requirement | Conditional-simulation intervals evaluated beside the random-forest baseline | Eligible only if REO MAE is also lower on the locked spatial holdout and the coverage requirement is met; otherwise skip |

Conditional Simulation Wins the Drill Gate Only With
The defensible conclusion is “unverified,” not “conditional simulation wins.” A drill pass requires lower withheld-assay REO mean absolute error than random-forest regression and interval coverage that meets the calibration gate. The supplied article title and provided source set document neither a realized rare-earth score nor a calibration result, assay count, or block count. Those blanks cannot be filled with plausible-looking precision.
Before inspecting either model’s errors, I would preregister the eligible assay population, economic REO grade basis, locked 20% spatial holdout, prediction support, and MAE and coverage decision statistics. A revised split or favorable second test is a new evaluation, not confirmation of the first. The supplied record does not document the split procedure, training proportion, random seed, or repeated-split design; the reserved share is not an accuracy certificate (Supplied article title; provided source set).
Evaluate both models on identical withheld assays, pairing each observed REO grade with its prediction. In REO percentage points, MAE is the sum of absolute observed-minus-predicted grades divided by the number of withheld assays. Subtract forest MAE from simulation MAE to establish the required direction of improvement. Publish the actual assay count, block count, and boundary-block counting rule; the reserved share cannot replace those denominators. R², classification accuracy, and training error cannot substitute for this continuous-grade test.
Calibration is a separate veto. For each withheld REO assay, record whether its observed grade lies inside its conditional-simulation prediction interval, then report covered count divided by total held-out count. High-grade cells and plausible-looking uncertainty shading establish neither coverage nor calibration. The source record supplies no covered count, total, or direct rare-earth validation result, so simulation remains unverified. The forest cannot inherit eligibility merely because it is available.
The coverage target below is this guide’s preregistered operating gate, not a universal geostatistical standard. Report the observed fraction even if it fails. Do not rescue the model by changing the uncertainty level, pooling deposits, or selecting a favorable cutoff after comparison; each requires a new, separately identified evaluation. Before authorizing a drill, release the paired errors, assay and block denominators, and coverage fraction. Missing required comparisons mean skip.
| Criterion | Conditional simulation | Random-forest regression | Drill consequence |
|---|---|---|---|
| Data conditioning | Use the locked assay population, economic grade basis, and prediction support. | Use the same locked evidence for the point-grade benchmark. | Changed conditioning or geography requires a new evaluation. |
| Point-grade prediction | Report paired MAE, assay count, and block count; neither count is supplied. | Report MAE on identical withheld assays using the same counts. | No lower-MAE winner can be established from the supplied record. |
| Uncertainty calibration | Publish covered withheld assays divided by all withheld assays. | The point-grade comparator cannot replace simulation’s interval-coverage test. | Simulation must meet the stated coverage gate. |
| Failure behavior | Disclose higher MAE or undercoverage without retuning the locked test. | A lower forest MAE is not an automatic fallback. | Either simulation failure means skip. |
| Operational verdict | Candidate only until both gates are evaluated. | Comparator only. | No realized pass is documented. |
| Winner gate | Selected only if its MAE is lower and coverage from its nominal prediction intervals meets the preregistered coverage requirement. | Comparator; never an automatic fallback. | Both gates pass: select conditional simulation. Otherwise: skip; do not drill. |

What the Data Doesn't Tell You
The source base supports a local test, not a general REO ranking. The provided source set contains no completed rare-earth mapping comparison of conditional simulation with random forests, and no applied simulation algorithm, realizations, conditioning set, variogram, or geostatistical model. Jeffrey Näf’s “Variable Importance in Random Forests” describes general forest capabilities, not rare-earth performance. Counter-evidence should include heterogeneous deposits where conditional simulation performs worse and forest flexibility fits local structure better, without converting either outcome into a universal REO discovery probability.
Withheld rows can overstate independent evidence. Several withheld intercepts from one mineralized vein can carry nearly the same information. A seemingly adequate reserved sample may still supply only one effective geological replicate. I would inspect intercept spacing, sample identifiers, and vein membership: closely clustered tests of the same structure do not establish transfer to another structure. The withheld-row share is a sampling design, not a geological replication rate.
Deposit class can reverse the ranking. Bastnäsite-bearing carbonatite, monazite-bearing sands, and ionic-adsorption clays need not share a stationary grade process. A variogram fitted across incompatible geology can suppress real spatial structure, while nonlinear forest flexibility can capture local contrasts. A conditional-simulation premium is earned only in the local comparison, not presumed from a method label or borrowed from another prospect.
Distribution choices can conceal unsupported certainty. Normal-score transformation and Gaussian simulation change the mathematical representation, not the physical REO distribution. Replacing a below-detection assay with zero can make conditional central estimates appear more certain than their sampling support warrants. I would retain censoring flags, detection limits, laboratory rounding, and inconsistent total-REE definitions in the audit trail rather than silently harmonize them.
Between-location success is not verification beyond the boundary. Conditional confidence intervals describe uncertainty under the selected statistical model. They do not directly verify an undiscovered vein or an unobserved REO-enrichment process where no withheld assay tests either. Successful coverage between sampled locations assesses calibration over the evaluated domain; it cannot certify geological continuity outside it.
Predictive accuracy is not reserve verification. Contained tonnes additionally require deposit volume, density, mineralogy, metallurgical recovery, processing assumptions, and cutoff-relevant costs. None follows from the fraction of correctly classified background cells. I would keep that classification exercise separate from REO grade prediction and reserve estimation, and identify the missing engineering and economic inputs before making a recoverable-reserve claim.
| Audit | What to check | How to interpret failure |
|---|---|---|
| Effective replication | Spatial clustering, shared vein membership, and repeated geological support | Many rows do not establish transfer to an untested structure |
| Deposit compatibility | Spatial structures before pooling variogram evidence | Heterogeneity can favor forest flexibility over conditional simulation |
| Assay comparability | Detection limits, censoring, laboratory precision, and total-REE definitions | Apparent precision may reflect data-handling choices |
| Inference boundary | Withheld sampled ground versus unmapped geology | Interval calibration is not confirmation of a new vein |
| Reserve conversion | Volume, density, mineralogy, recovery, processing, and cutoff-relevant costs | Grade accuracy alone does not establish recoverable tonnes |
Action: Keep the replication audit, assay definitions, local model comparison, and interval check in one drill-screening record. Only lower withheld REO MAE than random-forest regression and the stipulated empirical interval coverage earn drill-eligible status. If either gate fails, skip; map realism cannot repair the result.

Bayan Obo
Bayan Obo is non-reproducible from the supplied evidence, so skip. This is not a demonstrated loss for conditional simulation. The provided source set contains no named, sample-level REO geochemical dataset, and no actual assay count or observed REO range is recoverable. Without that ledger, neither a model advantage nor an interval-coverage pass exists to report.
The required starting material is one citable dataset from the Bayan Obo carbonatite district in Inner Mongolia, China: actual coordinates with their reference system, sample identifiers, laboratory totals, constituent-element measurements, detection limits, and analytical methods. Those records must remain traceable to sample material and analytical batches. A historical resource estimate or production total cannot substitute. I would not fill the missing provenance with plausible Bayan Obo values.
The target must also be recoverable, not reverse-engineered from model performance. The source must state whether REO is a measured total or an oxide-equivalent sum, its element inclusions, units, and treatment of incomplete analyses and nondetects. It must supply geological covariates, random-forest settings, and the conditional-simulation variogram family, nugget, sill, directions, and ranges. Preprocessing, feature selection, tuning, and variogram fitting must use only the training records; withheld records are for evaluation only. None of those source fields is available here. Each missing field is a reproducibility flag, not an invitation to choose a default silently.
With a real source, I would freeze the target and assay-eligibility rules first, then count N. Applying the locked spatial-block allocation requires documenting its integer-rounding and any residual-block assignment before reporting h, N − h, the actual number of spatial blocks, and the observed minimum–maximum REO. None is recoverable from the supplied material, so that allocation has not been performed. A reserved holdout share cannot compensate for missing assays or an undocumented block geometry.
The withheld-assay ledger must preserve the evidence needed for the decision. For its actual rows, each model’s MAE is its sum of absolute errors divided by h; conditional-simulation coverage is c/h, where c counts withheld records inside its nominal prediction intervals. Retain that raw fraction before presenting a percentage. No eligible withheld rows exist in the supplied material, so the following is a source-status specification, not a fabricated ledger.
| Required audit field | Value supported by supplied evidence | Decision significance |
|---|---|---|
| Sample identifier | No ledger supplied | Identity cannot be reconciled |
| Spatial block | No assignments supplied | Locked holdout cannot be audited |
| Observed REO percentage | No assay values supplied | Range and errors are unavailable |
| Conditional-simulation prediction | No predictions supplied | MAE cannot be calculated |
| Random-forest prediction | No predictions supplied | Comparator is unavailable |
| Absolute error, both models | Requires observed values and predictions | No comparison is possible |
| Conditional-simulation interval membership and c/h | No intervals or withheld records supplied | Coverage gate is untested |
Comparison: not computable; outcome: non-reproducible; action: do not drill. Conditional simulation has not been shown to beat random-forest regression or satisfy its interval-coverage gate. This section is a new audit of the available source evidence, not a numerical reanalysis or a result the Bayan Obo researchers necessarily reported. A valid new reanalysis of real source data requires the named primary dataset and reproducible settings first.

How to Choose Well
A convincing model is still a no-go if it borrows from the test. Before fitting either model, fix the assay target, eligible drill cells, economic-cutoff basis, locked spatial holdout, MAE calculation, and interval-coverage requirement in writing. Changing the assay target, test geography, error metric, or interval criterion after seeing errors returns a no-go, not a drill recommendation. A reserved fraction is a partition rule, not an accuracy certificate: it specifies how the test is separated, not how close future predictions will be.
The protocol needs an information firewall. Exclude test coordinates from variogram estimation, test assays from normalization statistics, and test labels from random-forest tuning. If any leak occurs, the affected scorecard is invalid; calculating a new average cannot restore independence. The withheld design must include positive-intercept and background-intercept assays in separate spatial blocks, rather than merely adding assays from one favorable geological body. If either intercept is absent, or only one sampled body is represented, local grade predictability is unestablished: skip.
Apply the preregistered dual gate without bargaining. On identical withheld cells, conditional simulation must have lower REO MAE than random-forest regression, and its nominal prediction intervals must meet the locked coverage floor. MAE averages absolute assay errors; coverage counts withheld assays inside those intervals. Winning only one gate is a veto, and an acceptable ra
Frequently Asked Questions
Why is the 20% spatial holdout not yet evidence of drilling readiness?
It is unverified because the retrieved sources identify no rare-earth locality, dataset, spatial split procedure, cross-validation design, or holdout result.
Does the reported 78% establish a rare-earth performance result?
No, because the supplied research does not identify what 78% measures or connect it to rare-earth performance.
What common prediction target would make the random-forest and conditional-simulation comparison valid?
Both models should estimate total REO grade within eligible drill cells under one laboratory convention, mineralogical basis, and spatial support.
Can a random forest’s vote proportions or dispersion establish calibrated REO prediction intervals?
No; tree-vote proportions or raw forest dispersion are not automatically calibrated REO-grade uncertainty, and their spread cannot establish the required empirical coverage of withheld assays.
Why can strong withheld sample-grade agreement still fail to validate block-scale REO estimates?
Because a bulk-block REO average, a channel-sample average, and a narrow core intercept are different observations that must be reconciled or else restrict the claim to the support actually represented.
What must the grade simulation achieve before a drill target is released?
It must produce lower withheld REO mean absolute error than random forests and pass the prescribed empirical prediction-interval coverage gate; if either gate fails, the drill target should be skipped.
Quick answers
| Is the claimed 20% spatial holdout verified by the retrieved evidence? | No; the 20% holdout is unverified, and the retrieved sources identify no rare-earth locality, dataset, spatial split procedure, cross-validation design, or holdout result. |
| Why can a 20% spatial holdout fail to test performance on new ground? | Withheld terrain may be unrepresentative of the training area, allowing favorable aggregate performance to conceal failures in the area relevant to a drilling decision. |
| What should be checked before making a drilling or investment commitment? | Verification should examine the spatial split procedure, withheld-area performance, calibration, and independent ground truth. |
| What target should conditional simulation and random forest estimate for a valid comparison? | Both models should estimate the same continuous quantity: total REO grade within eligible drill cells, under one laboratory convention, mineralogical basis, and spatial support. |
| What conditions would make the grade simulation drill-eligible? | It would be drill-eligible only if it yields lower withheld REO MAE than random forest and passes the article’s prescribed empirical prediction-interval coverage gate. |
Also worth reading: The best books for mastering spatial statistics and geospatial mapping: best books for mastering spatial · Kriging vs Random Forest: 34% Fewer Meters, 0.91 AUC: Kriging vs Random Forest: 34% · One rare fossil discovery finally settles the mystery of the Nanotyrannus: One rare fossil discovery finally