AI models flub these intelligence tests. Can you fare any better?
That's a major factor in how well models do on the most famous puzzle-based benchmark, ARC-AGI. These problems require you to infer abstract, general rules from a set of examples. Models do better on ARC puzzles when they receive each grid not as an image but as a string of numbers that encodes the color of each cell. Research suggests that even when models answer ARC-AGI questions correctly, they often do so using byzantine and non-generalizable rules, whereas humans draw on simple visual concepts. Despite these disadvantages, models have gotten quite good at ARC-AGI over the past year, but some puzzles--such as the one printed here--still stump them.
Aug-26-2026, 09:00:00 GMT
- Country:
- North America > United States > Illinois (0.14)
- Genre:
- Research Report (0.49)
- Industry:
- Media (0.70)
- Leisure & Entertainment > Games (0.47)
- Technology: