BriGap

BriGap is a venue for linguists and NLP scientists to meet: what fruitful interactions can we have? How do we build upon each other’s work?

Program of the event

Venue: Université Paris Cité, campus des Grands moulins, room 247E and 237C, Halle aux farines, Paris, France
Date: 11 July, 2026

09:00-09:10 Welcome and introduction
09:10-10:10 Keynote talk by Raquel Fernández (Universiteit van Amsterdam)
Title: Modelling Multimodality in Human Cognition with Machine Learning Models
Abstract: Linguistic communication is both inherently multimodal and interactive. Conceptual knowledge encoded in language is grounded in our sensory-motor experience, while in face-to-face dialogue we use speech in tandem with non-verbal signals such as gaze and gestures. In this talk, I will present our work on using machine learning models trained on body movements, vision, and language to study fundamental questions regarding multimodality in human cognition. I will first focus on behaviour: I will discuss our approach to co-speech gesture representation learning, showing that the resulting gesture embeddings exhibit properties that support theoretically motivated hypotheses. I will then move to describing our experiments on modelling brain activity with pre-trained vision-language models, which add to existing evidence for the intricate relationship between language and perception in the human brain.
10:10-10:30 A graph-based analysis of semantic types and coercion in contextualized word embeddings (Long Chen, Deniz Ekin Yavas) [abstract]
10:30-10:50 Cross-linguistic Geometry of Adjective Representations in Multilingual Transformers: Semantic Class, Gradability, and Positional Effects (Tancredi Monterosso) [abstract]
10:50-11:30 coffee break/poster session

Implementing Disjunctive Anaphora ‘a la Dependent Type Semantics (Hinari Daido, Daisuke Bekki) [abstract]
Inferring Formal Grammars from Syntactically Annotated Corpora (Ekaterina Voloshina, Krasimir Angelov) [abstract]
Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests (Manar Ali, Judith Sieker, Sina Zarrieß, Hendrik Buschmeier) [abstract]
Using the Mimi codec for metalinguistic representations (Artem Saloev, Erin Pacquetet, Nicolas Ballier) [abstract]
Towards Benchmarking Old Church Slavonic Lemmatization (Usman Nawaz, Marianna Napolitano, Iris Karafillidis, Liliana Lo Presti, Marco La Cascia) [abstract]
Processing Effects of Code-Switching in Humans and LLMs (Marina Sokolova, Natalia Moskvina, Nayara Mirio e Silva) [abstract]
Implicatures: a Dataset and Experiments on a Language Model (Gustavo Cilleruelo Calderón, Alexandra Birch, Emily Allaway) [abstract]
11:30-11:50 A Formal Model of Lexical Negation in Discrete Communication (Mikołaj Piotr Golecki, Timothée Bernard) [abstract]
11:50-12:10 Polar Questions in SPA–TTR: Linking Dialogue, Acquisition, and Neurosemantics (Jonathan Ginzburg, Shiyun Dong, Robin Cooper, Andy Luecking, Staffan Larsson) [abstract]
12:10-12:30 From Execution to Exploration: Bridging the Usability Gap in Formal Natural Language Inference (Koharu Saeki, Daisuke Bekki) [abstract]
12:30-13:30 lunch break
13:30-13:50 Artificial Language Learning Paradigm Reveals Pragmatic Blind Spots in Vision-Language Models (Yan Cong, Julia Rayz) [abstract]
13:50-14:10 Misalignments in Common Ground as a Bridge Between Pragmatic Theory and LLM Evaluation (Judith Sieker, Sina Zarrieß) [abstract]
14:10-14:30 Transformers Learning Contrafactives: The Importance of Data Distributions (David Strohmaier, Simon Wimmer) [abstract]
14:30-15:10 coffee break/poster session

Neural DTS: Integrating Hyperbolic Classifiers into Natural Language Inference Systems (Honoka Kobayashi, Hinari Daido, Daisuke Bekki) [abstract]
Neural Wani: Toward Accelerating the Automated Theorem Prover wani for Dependent Type Theory (Nanako Miyagawa, Hinari Daido, Daisuke Bekki) [abstract]
Diagnosing Compositional Generalization in Transformers on ReCOGS with Compositional Graph Similarity (Bruno Leite Franco, Edson Emilio Scalabrin) [abstract]
Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks (Chit-Fung Lam, Elaine Uí Dhonnchadha) [abstract]
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty (Omar Momen, Sina Zarrieß) [abstract]
Beyond surprisal: Capturing N400 and P600 effects for metaphor via semantic, pragmatic, and predictive computational models (Veronica Mangiaterra, Paolo Canal, Chiara Barattieri di San Pietro, Valentina Bambini) [abstract]
15:10-16:10 Keynote talk by Tal Linzen (New York University)
Title: Formal symbolic models for LLMs: pretraining, evaluation and post-training
Abstract: Formal symbolic models—ideal versions of the computational problems involved in learning about and interacting with the world—are central in linguistics and cognitive science. These models make it possible to generate unlimited amounts of synthetic data of variable complexity, as well as define verifiably correct outcomes for each instance of the problem. I will discuss three studies that leverage these properties of formal models. First, I will show that by pretraining transformer LLMs on formal languages before training them on natural languages, we can make training more compute and data efficient overall. Second, I will introduce context-free language recognition as an evaluation task for LLM. I will show that the complexity of the grammar reliably predicts the model’s accuracy on this task, and that even the strongest reasoning models available struggle to perform this task as the complexity of the language increases. Finally, I will demonstrate how symbolic Bayesian models can be used to evaluate and improve LLM abilities to update their probabilistic beliefs when interacting with users.
16:10-16:20 Closing remarks