Unsupervised Conformal Inference: Bootstrapping and Alignment to Control LLM Uncertainty
Pang, Lingyou, Huang, Lei, Lin, Jianyu, Wang, Tianyu, Horiguchi, Akira, Aue, Alexander, Priebe, Carey E.
Reliable uncertainty quantification (UQ) for large language models (LLMs) is needed for trustworthy AI. An assertive yet baseless claim can swiftly spread and cause damage, but for most practitioners, frontier models arrive only as black-box APIs with no access to gradients, exact log probabilities, or hidden states [18]. Hence deployment teams must make keep-or-discard decisions from samples alone. In black-box deployments, LLM uncertainty must be inferred from the sampled outputs themselves. Query-only signals include: (i) semantic-entropy methods that quantify dispersion across equivalence classes of responses and are effective for hallucination detection [8, 11]; (ii) self-consistency, which uses agreement among independently sampled answers as a proxy for confidence [30, 31]; and (iii) geometry-based measures computed from response embeddings--e.g., local density or Gram-volume statistics--that correlate with quality and robustness [17, 22]. Because these signals require neither logits nor gradients, they are natural conformity scores for our unsupervised conformal calibration; in parallel, conformal wrappers for language modeling and factuality control are emerging [20, 23].
- Country:
- North America > United States
- New York (0.04)
- Pennsylvania > Allegheny County
- Pittsburgh (0.04)
- California > Yolo County
- Davis (0.04)
- Europe
- Monaco (0.04)
- Middle East > Cyprus
- Asia > Japan
- Honshū
- Tōhoku > Iwate Prefecture
- Morioka (0.04)
- Chūbu > Toyama Prefecture
- Toyama (0.04)
- Tōhoku > Iwate Prefecture
- Honshū
- North America > United States
- Genre:
- Research Report (0.65)
- Technology: