Suchir Salhan

I research how AI systems learn and gain capabilities.

PhD Candidate in Computer Science · University of Cambridge · Gonville & Caius College

About me
Suchir Salhan click to read more

Suchir Salhan is a PhD candidate at the University of Cambridge +

Lecturing at Cambridge
Lecturing Li18 Computational Linguistics, Cambridge. (click me)

My work is guided by a single central question: how do intelligent systems learn? I ask it across minds and machines at once, tracing the same puzzle through human and machine learning, language and intelligence, cognition and computation. Language is where these threads meet: it offers rich structure (syntax, semantics, morphology, phonology) alongside decades of developmental theory, cross-linguistic variation, and behavioural evidence against which learning can be measured.

I’ve had a long fascination with the intersection of language and computation—how humans have developed the capability to acquire natural language to communicate, learn, and reason, despite the diversity of linguistic systems; and how we might build machines that can do the same. I arrived in Cambridge in 2020 to pursue a BA and MEng in Computer Science & Linguistics at Gonville & Caius College, Cambridge, where I earned a “starred First” and a Distinction.

I got into research early, and slowly the two worlds stopped competing. This is what my work is about — minds and machines: cognition, language and learning on one side; AI, agents and small models on the other; and, between them, questions of interpretability, representation and behaviour. My master’s thesis pulled me into BabyLM, and the BabyLM cinematic universe is expanding. I co-organise BabyVLM. At first glance, a small language model looks like a weaker version of a large one. I think that misses the point. So rather than asking only “How good is the model?”, I became interested in a different question.

On the road
A snapshot from research life. (click to shuffle 🔀)

Get in touch

Say hello 👋

Always happy to hear from students, collaborators, and anyone curious about small language models — drop me a line.

Department of Computer Science & Technology, University of Cambridge · Gonville & Caius College
Research Interests
The Computer Laboratory, Cambridge click to read more

My research sits at the intersection of Machine Learning, Cognitive Science, and Linguistics, with a particular focus on developing Small Language Models. +

My Research Agenda

BabyLM poster at EMNLP 2025
Human-scale modelling — what sequence length is best for BabyLM? EMNLP 2025.

How do models learn?

How does language shape learning?+

What representations emerge?+

Can we trust small AI?+

Addressing these questions are directly beneficial for the AI community and stakeholders using AI in their systems. My work helps stakeholders build mergeable, efficient small models and agentic pipelines that are cheaper to deploy than frontier systems. More generally, I aim to contribute to the AI community by understanding and building a foundation for multilingual, fairer AI that reaches beyond English into under-served languages and build more human-centered AI that is culturally-adaptive, safe, fair and trustworthy.

Research

Research Themes

  • Cognitively-Inspired Design & Evaluation — developmentally plausible models & evaluation.
  • Pretraining & Interpretability — learning dynamics of small LMs via Pico.
  • Multilinguality — human-scale, cross-lingual, bilingual competence via Beetle.
  • Tokenization — information-driven subword segmentation (ByteSpan).
  • Alignment & Interaction — pedagogical alignment, student proxies.
  • Cognitive Science & Linguistics — grammar induction, cross-lingual phonology.

Explore my Research & Publications

Research Impact

BabyLM at EMNLP
Human-scale AI — BabyLM at EMNLP.

I study how small language models learn languages, using a collection of 330 open-source models trained in different bilingual and monolingual settings, known as BEETLE. +

Explore Beetle

Two Outstanding Paper Awards at EMNLP 2025
Two Outstanding Paper Awards at the BabyLM Workshop, EMNLP 2025 (Suzhou).
News

Latest

2026
Three EMNLP 2026 main-conference papers — Beetle, LangMAP and Cross-Lingual Alignment — accepted at EMNLP 2026 (Budapest).
January 2026
BabyBabelLM accepted to EACL 2026 (Main Conference, Rabat); two oral papers at the 23rd Old-World Conference in Phonology (OCP23), hosted at Gonville & Caius College. See my reflections on hosting OCP23.
September 2025
🎉 Eight accepted papers at EMNLP and NeurIPS workshops — including two Outstanding Paper Awards at BabyLM 2025.
March 2025
Released PicoLM, the Cambridge Small Language Model & Learning Dynamics Framework.
Talks

Invited Talks & Seminars

  • X-PPL-26 — Crosslinguistic Perspectives on Processing & Learning, Bayonne — 2026
  • Language & Multimodal Processing Group, University of Copenhagen (with Dr Diana Galvan-Sosa) — 2026
  • Sheffield NLP — Understanding the Human-Scale AI Frontier — 2026
  • 13th Conference on the Mental Lexicon, McGill University — Invited Keynote — 2025
All talks
Pico demo at EMNLP 2025
PicoLM — hypothesis-driven small-model research.

Work with me

Work with me.

At Cambridge?

I enjoy working with students and collaborators from a wide range of disciplines.

If you're a computer scientist

If you're a linguist

If you're interested in cognitive science

If you're interested in education

If you're interested in AI more broadly

And if you're working somewhere in between?

Find out more about research that I have mentored →

Theoretical linguistics poster, CogInterp NeurIPS 2025
Representations & interpretability — CogInterp @ NeurIPS 2025.
poke an image to explode it ·
Teaching

Teaching & Supervision

I lecture and supervise several courses in Cambridge. +

I also organise the Cambridge Natural Language Processing Seminars (NLIP), which take place every Friday at midday in the Computer Lab and are open to the public. Please feel free to attend if you are interested!

Teaching & resources