PETLP: A Privacy-by-Design Pipeline for Social Media Data in AI Research
Oh, Nick, Vrakas, Giorgos D., Brooke, Siân J. M., Morinière, Sasha, Duke, Toju
–arXiv.org Artificial Intelligence
We introduce PETLP (Privacy-by-design Extract, Transform, Load, and Present), a compliance framework that embeds legal safeguards directly into extended ETL pipelines. Central to PETLP is treating Data Protection Impact Assessments as living documents that evolve from preregistration through dissemination. Through systematic Red-dit analysis, we demonstrate how extraction rights fundamentally differ between qualifying research organisations (who can invoke DSM Article 3 to override platform restrictions) and commercial entities (bound by terms of service), whilst GDPR obligations apply universally. We demonstrate why true anonymisation remains unachievable for social media data and expose the legal gap between permitted dataset creation and uncertain model distribution. By structuring compliance decisions into practical workflows and simplifying institutional data management plans, PETLP enables researchers to navigate regulatory complexity with confidence, bridging the gap between legal requirements and research practice.
arXiv.org Artificial Intelligence
Oct-17-2025
- Country:
- Europe
- Denmark > Capital Region
- Copenhagen (0.04)
- Germany (0.04)
- Ireland (0.04)
- Middle East > Cyprus (0.04)
- Netherlands > North Holland
- Amsterdam (0.04)
- United Kingdom > England
- Oxfordshire > Oxford (0.04)
- Denmark > Capital Region
- North America > United States
- Pennsylvania (0.04)
- Europe
- Genre:
- Overview (1.00)
- Research Report
- Experimental Study (1.00)
- New Finding (0.67)
- Industry:
- Government > Regional Government
- Europe Government (0.47)
- Information Technology > Security & Privacy (1.00)
- Law (1.00)
- Government > Regional Government
- Technology: