Tech

AI benchmarks reveal infant learning efficiency outpaces current models

New research from Meta and leading universities shows that while AI excels at pattern recognition, it struggles to replicate the rapid, physical, and social learning mechanisms of human infants.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: WIRED · original
AI Isn’t Smarter Than a Baby—Yet
EgoBabyVLM Challenge highlights gaps in vision language model capabilities

Researchers from Meta, Stanford University, the University of Tokyo, and France’s École Normale Supérieure have introduced the EgoBabyVLM Challenge, a benchmark designed to evaluate how vision language models interpret the world from an infant’s perspective. The test requires artificial intelligence systems to process approximately 1,000 hours of video footage captured from cameras strapped to the heads of infants and toddlers. The results indicate that current cutting-edge models fail to match the rapid and efficient learning capabilities demonstrated by babies, who learn through messy, multimodal, and physical interactions rather than curated datasets.

Unlike modern AI models that consume vast amounts of training data and significant energy, infants learn to make sense of their environment with remarkable efficiency. Babies identify new objects after seeing them only once or twice, relying on fleeting observation and physical interaction. The EgoBabyVLM Challenge highlights that this learning process involves a kaleidoscopic view where parents discuss invisible objects, use gestures, and reference past or future events, alongside rich tactile experiences. Experts suggest that future AI architectures must incorporate these social cues and physical reasoning mechanisms to achieve similar efficiency, moving beyond pure pattern recognition.

Michael Frank, a cognitive scientist at Stanford University involved in the challenge’s development, noted that the results indicate language alone is insufficient for AI to understand the world as humans do. The test builds on the 2023 BabyLM challenge, which showed that transformer-based AI models could learn language syntax with data volumes comparable to a 10-year-old’s intake. That earlier finding challenged Noam Chomsky’s theories on hardwired syntax, demonstrating that transformers are adept at finding patterns in data. However, the new benchmark reveals significant limitations when it comes to understanding the physical world.

Joshua Tenenbaum, a cognitive scientist at MIT, observed that while transformers excel at pattern recognition, they fail to acquire common sense regarding the physical world, social dynamics, or theory of mind. In 2024, researchers demonstrated that a basic vision language model could learn simple concepts, such as identifying a ball, using data from a single infant’s head-mounted camera. While this represents progress, it remains distinct from the sophisticated reasoning capabilities that children develop by the age of two.

Early work by Frank and colleagues showed that models designed to learn causality and visual-temporal relationships using baby-head video data were more effective at understanding object dynamics than standard approaches. These models were able to learn about the dynamics of different objects, laying a foundation for physical reasoning. The authors of the EgoBabyVLM paper suggest that borrowing ideas from cognitive science and neuroscience, such as designing models that pay attention over longer periods and interpret social cues, could enable progress toward more humanlike learning algorithms.

Continue reading

More from Tech

Read next: France Enacts Strict Ban on Unsolicited Telemarketing Calls
Read next: OpenAI expands Daybreak cybersecurity programme with new model tiers
Read next: AI models map 766 genes in schizophrenia genetic architecture