Learning curves theory for hierarchically compositional data with power-law distributed features

Cagnetta, Francesco, Kang, Hyunmo, Wyart, Matthieu

May-13-2025–arXiv.org Machine Learning

Recent theories suggest that Neural Scaling Laws arise whenever the task is linearly decomposed into power-law distributed units. Alternatively, scaling laws also emerge when data exhibit a hierarchically compositional structure, as is thought to occur in language and images. To unify these views, we consider classification and next-token prediction tasks based on probabilistic context-free grammars -- probabilistic models that generate data via a hierarchy of production rules. For classification, we show that having power-law distributed production rules results in a power-law learning curve with an exponent depending on the rules' distribution and a large multiplicative constant that depends on the hierarchical structure. By contrast, for next-token prediction, the distribution of production rules controls the local details of the learning curve, but not the exponent describing the large-scale behaviour.

artificial intelligence, machine learning, natural language, (16 more...)

arXiv.org Machine Learning

May-13-2025

arXiv.org PDF

Add feedback

Country:
- North America
  - Canada (0.04)
  - United States > Maryland
    - Baltimore (0.04)
- Europe
  - Switzerland (0.04)
  - United Kingdom > England
    - Cambridgeshire > Cambridge (0.04)
  - Italy > Friuli Venezia Giulia
    - Trieste Province > Trieste (0.04)
- Asia > South Korea
  - Seoul > Seoul (0.04)

Genre:
- Research Report (0.64)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Grammars & Parsing (1.00)
  - Representation & Reasoning
    - Rule-Based Reasoning (0.81)
    - Expert Systems (0.81)
  - Machine Learning > Neural Networks
    - Deep Learning (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found