Improving MLLM Historical Record Extraction with Test-Time Image
Archibald, Taylor, Martinez, Tony
–arXiv.org Artificial Intelligence
W e present a novel ensemble framework that stabilizes LLM-based text extraction from noisy historical documents. W e transcribe multiple augmented variants of each image with Gemini 2.0 Flash and fuse these outputs with a custom Needleman-Wunsch-style aligner that yields both a consensus transcription and a confidence score. W e present a new dataset of 622 Pennsylvania death records, and demonstrate our method improves transcription accuracy by 4 percentage points relative to a single-shot baseline. W e find that padding and blurring are the most useful for improving accuracy, while grid-warp perturbations are best for separating high-and low-confidence cases. The approach is simple, scalable, and immediately deployable to other document collections and transcription models.
arXiv.org Artificial Intelligence
Sep-15-2025
- Country:
- Europe (0.68)
- North America > United States
- Pennsylvania (0.24)
- Genre:
- Research Report > Experimental Study (0.46)
- Industry:
- Technology: