The Achilles Heel of AI: Fundamentals of Risk-Aware Training Data for High-Consequence Models
–arXiv.org Artificial Intelligence
AI systems deployed in high - consequence environments -- such as defense, intelligence, and disaster response -- must detect rare, high - impact events under operational constraints. Traditional annotation strategies that emphasize volume over value fail in these s ettings, introducing redundancy, noise, and risk. This paper presents smart - sizing, a strategic approach to training data curation that prioritizes informational value, label diversity, and model - guided selection. We introduce Adaptive Label Optimization ( ALO) as an implementation framework for smart - sizing, combining pre - labeling, human - in - the - loop feedback, disagreement analysis, and marginal utility - based stopping rules. Through empirical experiments, we show that models trained on only 20 - 40% of a curated dataset match or exceed the performance of full - data baselines, particularly in rare class recall and edge - case generalization. We also demonstrate how systematic labeli ng errors embedded in both training and validation sets can produce misleading model evaluations, underscoring the need for internal audit mechanisms and data governance. Smart - sizing reframes annotation from a static task to a feedback - driven protocol aligned with operational goals. It offers a measurable path to improving AI robustness while reducing labeling cost, enabling teams to label what matters -- and stop when it no longer helps.
arXiv.org Artificial Intelligence
May-22-2025
- Country:
- North America > United States (0.28)
- Genre:
- Research Report > New Finding (1.00)
- Industry:
- Government > Military (0.68)
- Transportation > Air (0.46)
- Technology: