Image Details
Caption: Figure 6.
Performance versus training-set size under importance-weighted evaluation with the tuned recalibration exponent (α ≃ 0.36, ESS target 30%). Conventions match Figure 5. Under shift-corrected evaluation, foundation models lead both metrics across the full range of ntrain.
© 2026. The Author(s). Published by the American Astronomical Society.