Exposing Assumptions in AI Benchmarks through Cognitive Modelling
Rystrøm, Jonathan H., Enevoldsen, Kenneth C.
–arXiv.org Artificial Intelligence
Cultural AI benchmarks often rely on implicit assumptions about measured constructs, leading to vague formulations with poor validity and unclear interrelations. We propose exposing these assumptions using explicit cognitive models formulated as Structural Equation Models. Using cross-lingual alignment transfer as an example, we show how this approach can answer key research questions and identify missing datasets. This framework grounds benchmark construction theoretically and guides dataset development to improve construct measurement. By embracing transparency, we move towards more rigorous, cumulative AI evaluation science, challenging researchers to critically examine their assessment foundations.
arXiv.org Artificial Intelligence
Sep-25-2024
- Country:
- North America > United States
- New York > New York County
- New York City (0.05)
- Illinois > Cook County
- Chicago (0.04)
- Georgia > Fulton County
- Atlanta (0.04)
- California > Orange County
- Irvine (0.04)
- New York > New York County
- Europe > United Kingdom
- England > Oxfordshire > Oxford (0.14)
- Asia > China
- Hong Kong (0.04)
- North America > United States
- Genre:
- Research Report
- Experimental Study (0.66)
- New Finding (0.48)
- Research Report
- Industry:
- Education (0.68)
- Technology: