Predicting fatigue failure in steel components experimentally is costly, requiring testing across multiple compositions and processing conditions, spurring research on data-driven prediction models. Studies using the NIMS MatNavi steel fatigue dataset often report high point-prediction accuracy but rely on aggregate error metrics, leaving uncertainty about the reliability of individual predictions and whether accuracy is consistent across the fatigue-strength spectrum. This paper is the first to apply conformal prediction to steel fatigue strength, comparing seven interval-construction methods across 50 independent data splits and distinguishing marginal coverage from coverage within specific sub-regions of the predicted property. A gradient-boosting point model achieves an R^2 of 0.976 +/- 0.009 and a mean absolute error of 18.3 +/- 2.3 MPa. Split-conformal prediction provides valid marginal coverage (0.918) but drops to 0.758 in the highest-strength quartile, where design margins are most critical, a pattern also observed with a Gaussian process baseline. Two locally-adaptive methods correct this: a cross-fitted normalised conformal method holds 0.872-0.940 across quartiles at no cost in average width, and Mondrian group-conditional conformal prediction holds the tightest band of any method (0.917-0.939) at a 12% width premium, part of which traces to the more conservative finite-sample quantile level implied by per-group calibration at this sample size. Conformalized quantile regression, by contrast, restores marginal validity but inflates intervals in every quartile without closing the conditional gap. Marginal coverage claims for ML-based fatigue-strength predictions can conceal systematic unreliability precisely where engineering decisions are most risky; therefore, conditional coverage should be routinely assessed alongside marginal coverage.
Distribution-Free Conformal Prediction for Steel Fatigue Strength: Marginal Validity Is Not Enough
Predicting fatigue failure in steel components experimentally is costly, requiring testing across multiple compositions and processing conditions, spurring research on data-driven prediction models.
- Preview

- Year
- 2026
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2608.07589CC-BY-4.0
- TL;DR
- Semantic Scholar