Investigating the Association Between Text-Based Indications of Foodborne Illness from Yelp Reviews and New York City Health Inspection Outcomes (2023)
Shaveet, Eden, Su, Crystal, Hsu, Daniel, Gravano, Luis
–arXiv.org Artificial Intelligence
Foodborne illnesses are gastrointestinal conditions caused by consuming contaminated food. Restaurants are critical venues to investigate outbreaks because they share sourcing, preparation, and distribution of foods. Public reporting of illness via formal channels is limited, whereas social media platforms host abundant user-generated content that can provide timely public health signals. This paper analyzes signals from Yelp reviews produced by a Hierarchical Sigmoid Attention Network (HSAN) classifier and compares them with official restaurant inspection outcomes issued by the New York City Department of Health and Mental Hygiene (NYC DOHMH) in 2023. We evaluate correlations at the Census tract level, compare distributions of HSAN scores by prevalence of C-graded restaurants, and map spatial patterns across NYC. We find minimal correlation between HSAN signals and inspection scores at the tract level and no significant differences by number of C-graded restaurants. We discuss implications and outline next steps toward address-level analyses.
arXiv.org Artificial Intelligence
Oct-21-2025
- Country:
- North America > United States > New York
- Bronx County > New York City (0.05)
- New York County > New York City (0.04)
- Richmond County > New York City (0.06)
- North America > United States > New York
- Genre:
- Research Report
- Experimental Study > Negative Result (0.34)
- New Finding (0.67)
- Research Report
- Industry:
- Technology: