The AI Community Building the Future? A Quantitative Analysis of Development Activity on Hugging Face Hub
Osborne, Cailean, Ding, Jennifer, Kirk, Hannah Rose
–arXiv.org Artificial Intelligence
Open model developers have emerged as key actors in the political economy of artificial intelligence (AI), but we still have a limited understanding of collaborative practices in the open AI ecosystem. This paper responds to this gap with a three-part quantitative analysis of development activity on the Hugging Face (HF) Hub, a popular platform for building, sharing, and demonstrating models. First, various types of activity across 348,181 model, 65,761 dataset, and 156,642 space repositories exhibit right-skewed distributions. Activity is extremely imbalanced between repositories; for example, over 70% of models have 0 downloads, while 1% account for 99% of downloads. Furthermore, licenses matter: there are statistically significant differences in collaboration patterns in model repositories with permissive, restrictive, and no licenses. Second, we analyse a snapshot of the social network structure of collaboration in model repositories, finding that the community has a core-periphery structure, with a core of prolific developers and a majority of isolate developers (89%). Upon removing the isolate developers from the network, collaboration is characterised by high reciprocity regardless of developers' network positions. Third, we examine model adoption through the lens of model usage in spaces, finding that a minority of models, developed by a handful of companies, are widely used on the HF Hub. Overall, activity on the HF Hub is characterised by Pareto distributions, congruent with OSS development patterns on platforms like GitHub. We conclude with recommendations for researchers, companies, and policymakers to advance our understanding of open AI development.
arXiv.org Artificial Intelligence
Jun-5-2024
- Country:
- Asia
- Bangladesh > Dhaka Division
- Dhaka District > Dhaka (0.04)
- China (0.04)
- India (0.04)
- Middle East > Jordan (0.04)
- Bangladesh > Dhaka Division
- Europe
- France (0.28)
- Germany > Berlin (0.04)
- Italy (0.04)
- Netherlands > North Holland
- Amsterdam (0.04)
- United Kingdom
- England
- Cambridgeshire > Cambridge (0.04)
- Greater London > London (0.04)
- Oxfordshire > Oxford (0.14)
- Scotland > City of Edinburgh
- Edinburgh (0.04)
- England
- North America > United States
- California > San Francisco County
- San Francisco (0.14)
- Hawaii (0.04)
- Massachusetts > Middlesex County
- Cambridge (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.04)
- Missouri > St. Louis County
- St. Louis (0.04)
- New York
- Monroe County > Rochester (0.04)
- New York County > New York City (0.04)
- California > San Francisco County
- South America (0.04)
- Asia
- Genre:
- Research Report
- Experimental Study (0.94)
- New Finding (0.94)
- Research Report
- Industry:
- Government (1.00)
- Information Technology > Security & Privacy (1.00)
- Technology:
- Information Technology
- Artificial Intelligence
- Machine Learning > Neural Networks
- Deep Learning (1.00)
- Natural Language
- Chatbot (1.00)
- Large Language Model (1.00)
- Vision (0.67)
- Machine Learning > Neural Networks
- Communications > Social Media (0.88)
- Artificial Intelligence
- Information Technology