However, categorical features present a unique challenge as they require embedding a typically vast vocabulary into a smaller vector space for further calculations.
Optimizing proper loss functions is popularly believed to yield predictors with good calibration properties; the intuition being that for such losses, the global optimum is to predict the ground-truth probabilities, which is indeed calibrated.