Publications
Publications
Peer-reviewed conference, journal, and workshop papers, plus working papers. Author names in bold indicate my contribution; ✦ marks equal contribution.
Peer-Reviewed Conference & Journal Publications 12
1
Richard Diehl-Martinez, David Demitri Africa, Yuval Weiss, Suchir Salhan, Ryan Daniels, Paula Buttery
EMNLP 2025 Systems Demonstration · Suzhou, China
2
Jaap Jumelet, Abdellah Fourtassi, Akari Haga, Bastian Bunzeck, Bhargav Shandilya, Diana Galvan-Sosa, Faiz Ghifari Haznitrama, Francesca Padovani, François Meyer, Hai Hu, Julen Etxaniz, Laurent Prévot, Linyang He, María Grandury, Mila Marcheva, Negar Foroutan, Nikitas Theodoropoulos, Pouya Sadeghi, Siyuan Song, Suchir Salhan, Susana Zhou, Yurii Paniv, Ziyin Zhang, Arianna Bisazza, Alex Warstadt, Leshem Choshen
EACL 2026 (Main Conference) · Rabat, Morocco
3
Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Jaap Jumelet, Tal Linzen, Aaron Mueller, Suchir Salhan, Raj Sanjay Shah, Alex Warstadt, Ethan Gotlieb Wilcox
BabyLM Workshop 2026 @ EMNLP · Budapest, Hungary
4
A Computational Operationalisation of Competing Maturational Theories of Syntactic Development via Statistical Grammar Induction
Mila Marcheva, Suchir Salhan & Weiwei Sun
CogSci 2026 (Main Conference), Proceedings of the Annual Meeting of the Cognitive Science Society · Rio de Janeiro, Brazil
5
Fermín Moscoso del Prado Martín, Suchir Salhan
6
Fermín Moscoso del Prado Martín, Suchir Salhan
Proceedings of the Society for Computation in Linguistics (SCiL) · Presentation at ACL 2026, San Diego, USA
7
Structured Exposure Pretraining in Bilingual Language Models for Modelling L2 Language Processing
Suchir Salhan, Catherine Arnett, James Michaelov & Paula Buttery
8
LangMAP: A Language-Adaptive Approach to Tokenization
Clara Meister, Suchir Salhan, Andrzej Szablewski, Pietro Lesci, Paula Buttery & Tiago Pimentel
9
Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?
EJ Zhou, Suchir Salhan, Catherine Arnett
10
Raising the Stakes: A Pedagogical Reframing of Automarking in Low-Stakes Assessment
Aoife O'Driscoll, Lily Goulder, Laura Barbenel, Suchir Salhan, Gabrielle Gaudeau, Andrew Caines & Paula Buttery
11
Emergence and Stability of Cross-Lingual Representations in Language Models
Suchir Salhan, Laura Barbenel, Lily Goulder, Aoife O'Driscoll, Andrew Caines, Catherine Arnett & Paula Buttery
12
xBLiMPs: Cross-Lingual Syntactic Minimal Pairs for Language Model Interpretability and Evaluation
Suchir Salhan, Nuria Bosch Masip, Yury Makarov, Rigel Cierniak, Laura Barbenel, Lily Goulder, Aoife O'Driscoll, Andrew Caines, Paula Buttery, Theresa Biberauer & Catherine Arnett
Peer-Reviewed Workshop Publications 14
13
Salhan, S.A.✦, Diehl-Martinez, Richard, Goriely, Zebulon & Buttery, Paula (2024)
CoNLL 2024 BabyLM Challenge (Paper Track) · Miami, Florida, USA (Nov 2024)
14
Zebulon Goriely, Suchir Salhan, Pietro Lesci, Julius Cheng, Paula Buttery
ICML 2025 Tokenization Workshop (TokShop) · Vancouver, Canada (August 2025)
Recent dynamic tokenisation methods operate directly on bytes and pool their latent representations into patches. This bears similarities to computational models of word segmentation that determine lexical boundaries using spikes in an autoregressive model's prediction error. Inspired by this connection, we explore whether grouping predictable bytes — rather than pooling their representations — can yield a useful fixed subword vocabulary. We propose a new information-driven subword tokeniser, ByteSpan, that uses an external byte-level LM during training to identify contiguous predictable byte sequences and group them into subwords. Experiments show that ByteSpan yields efficient vocabularies with higher morphological alignment scores than BPE for English. Multilingual experiments show similar compression and Rényi efficiency for 25 languages.
15
Fermín Moscoso del Prado Martín, Suchir Salhan
ACL 2025 Main Conference (Poster) · Vienna, Austria (August 2025)
16
David Demitri Africa, Suchir Salhan, Yuval Weiss, Paula Buttery, Richard Diehl-Martinez
5th Workshop on Multilingual Representation Learning (MRL), EMNLP 2025 · Suzhou, China
Named-entity recognition (NER) in low-resource languages is usually tackled by finetuning very large multilingual LMs, an option that is often infeasible in memory- or latency-constrained settings. We ask whether small decoder LMs can be pretrained so that they adapt quickly and transfer zero-shot to languages unseen during pretraining. To this end we replace part of the autoregressive objective with first-order model-agnostic meta-learning (MAML). Tagalog and Cebuano are typologically similar yet structurally different in their actor/non-actor voice systems, and hence serve as a challenging test-bed. Across four model sizes (11 M – 570 M) MAML lifts zero-shot micro-F1 by 2–6 pp under head-only tuning and 1–3 pp after full tuning, while cutting convergence time by up to 8%. Gains are largest for single-token person entities that co-occur with Tagalog case particles si/ni, highlighting the importance of surface anchors.
17
Suchir Salhan, Hongyi Gu, Donya Rooein, Diana Galvan-Sosa, Gabrielle Gaudeau, Andrew Caines, Zheng Yuan, Paula Buttery
BabyLM Workshop, EMNLP 2025 · Suzhou, China
18
Bianca-Mihaela Ganescu, Suchir Salhan, Andrew Caines, Paula Buttery (Supervised MPhil Advanced Computer Science Thesis)
BabyLM Workshop, EMNLP 2025 · Suzhou, China
19
Yuan Gao, Suchir Salhan, Andrew Caines, Paula Buttery, Weiwei Sun
BabyLM Workshop, EMNLP 2025 · Suzhou, China
20
Suchir Salhan, Richard Diehl-Martinez, Zebulon Goriely, Paula Buttery
BabyLM Workshop, EMNLP 2025 · Suzhou, China
21
Pedagogical Alignment of LLMs Requires Diverse Cognitively-Inspired Student Proxies
Suchir Salhan, Andrew Caines, Paula Buttery
NeurIPS First Workshop on CogInterp: Interpreting Cognition in Deep Learning Models, 2025 · San Diego, California, USA
22
Theoretical Linguistics Constrains Hypothesis-Driven Causal Abstraction in Mechanistic Interpretability
Suchir Salhan, Konstantinos Voudouris
NeurIPS First Workshop on CogInterp: Interpreting Cognition in Deep Learning Models, 2025 · San Diego, California, USA
23
Glints of Gold or Troubling Waters? Can a School of Merged Monolingual Goldfish Models Swim in Bilingual Seas?
Suchir Salhan, EJ Zhou, Laura Barbenel, Aoife O'Driscoll, Lily Goulder, Lucas Resck, Catherine Arnett & Paula Buttery
EACL Workshop on Multilingual Multicultural Evaluation (MME), Non-Archival Full Paper, 2026 · Rabat, Morocco
24
Do Monolingual Language Models Learn Cross-Lingual Universal Conceptual Representations?
Suchir Salhan, EJ Zhou & Paula Buttery
ICLR Workshop on Unifying Concept Representation Learning (UCRL) & ICLR Workshop on Representational Alignment (Re-Align), 2026 · Rio de Janeiro, Brazil
25
L1 Influence in L2 Language Models: A Human-Centric Approach
Laura Barbenel, Lily Goulder, Aoife O'Driscoll, Suchir Salhan, Catherine Arnett, Andrew Caines & Paula Buttery
Computational Developmental Linguistics (CDL) Workshop @ ACL 2026
26
Repertoires, Not Scores: Instability as Signal in Cultural Evaluation of LLMs
Suchir Salhan, Filip Trhlik, Diana Galvan-Sosa & Paula Buttery
Culture x AI Workshop, ICML 2026 · Seoul, Korea
Working Papers 6
27
Salhan, S.A.✦ (2023)
Cambridge Occasional Papers in Linguistics, Volume 15, Article 3: pp. 55–110. ISSN: 2050-5949
28
Multimodal Language Modelling across Languages and Cultures: Grounding Strategies for Nouns and Verbs
Salhan, Suchir, Liu, Fangyu & Collier, Nigel (2022 / preprint)
Research Project, Language Technology Lab, Department of Theoretical and Applied Linguistics, University of Cambridge
29
The Distribution of Phonemes across Languages: Chance, Costs, and Integration across Linguistic Tiers
Fermín Moscoso del Prado Martín, Suchir Salhan (2026)
23rd Old-World Conference in Phonology (OCP23), Gonville & Caius College (Accepted Oral)
30
Convergent Equilibria in Cross-Lingual Phoneme Surprisal Distributions: Statistical and Simulation-Based Analysis
Suchir Salhan, Fermín Moscoso del Prado Martín (2026)
23rd Old-World Conference in Phonology (OCP23), Gonville & Caius College (Accepted Oral)
31
EJ Zhou, Suchir Salhan
5th Workshop on Multilingual Representation Learning (MRL), EMNLP 2025 · Suzhou, China
32
Linguistics in the Age of Language Models: What can Cognitively-Inspired Language Models offer to Linguistic Theory?
Salhan, S.A. (2025)
Position Paper in Cambridge Occasional Papers in Linguistics (CoPiL), Accepted, Volume 17