NeurIPS 2026 · Atlanta

The BabyVLM Workshop: Developmentally Plausible Multimodal Systems

A workshop I helped organise on whether vision can make language learning as sample-efficient in machines as it appears to be in children.

The BabyVLM Workshop, held at the Georgia World Congress Center alongside NeurIPS 2026,1Event site. starts from a gap that is easy to state and hard to close. A child learns language from something like 100 million words over the first years of life. A frontier language model is trained on well over a trillion. That is roughly a ten-thousand-fold difference in exposure, and children arrive at robust, grounded language anyway. The workshop asks a pointed version of the question: does the difference come down to the fact that children learn from more than text? A toddler hears words while seeing faces, objects, and actions, and the workshop's premise is that this multimodal signal, and vision in particular, may be part of what makes infant learning so efficient.

Aims

Rather than treat data efficiency as a purely architectural problem, BabyVLM frames it developmentally. The call invited work along three lines: developmentally plausible pretraining, where the training signal is constrained to something an infant could plausibly receive; developmentally aligned evaluation, where benchmarks measure the kinds of competence children acquire rather than adult exam performance; and longitudinal egocentric learning, which uses recordings taken from a child's own point of view. The framing is a deliberate sibling to the text-only BabyLM effort,2BabyLM Challenge. extended to the harder case where pixels and words arrive together.

The shared task

Attached to the workshop is the BabyVLM Challenge, a competition on developmentally plausible, sample-efficient vision-language modelling. Participants pretrain models on infant-scale datasets, roughly the volume and character of what a young child actually experiences, and are evaluated on developmental benchmarks rather than the large web-scale suites usual in the field. The submission deadline fell on 8 September 2026, with notifications later that month, and the challenge results are set to be presented at CVPR 2027, giving the multimodal side of the effort a natural home in a vision venue.

My role

I was a co-organiser, working with colleagues including Paula Buttery at Cambridge, Leshem Choshen, Boqing Gong, and Aaron Mueller. Much of my contribution sat around the shared task and the developmental framing that ties the workshop and challenge together. What I find compelling about the exercise is that it inverts the usual scaling instinct. Instead of asking what more data buys, it asks what a fixed, child-sized budget can be made to yield when the input is grounded in a scene. That is a good constraint to reason under, and it keeps the cognitive question honest rather than letting it dissolve into a race for larger corpora.