Levent Sagun
About
I am a Research Scientist at FAIR in Paris. I study how large models fail and whether the methods used to measure these failures actually measure what we think they do. I am particularly interested in how benchmarks and technical metrics are situated within, and help shape, AI governance frameworks.
Much of my current work centers on some kind of representation problem. I look at how datasets, benchmarks and alignment methods represent people, values and social concepts. I'm interested in how design choices determine whose views enter into a system, and how differences between communities are characterized and who or what gets left out.
A new line of upcoming work I am interested in is on the tensions between personalization and homogenization. In particular, I study how preference optimization may compress/erase plural values into a narrow range of outputs, and how this loss of diversity (or collapse) could be detected and measured. This work is still at an early stage and is mostly unpublished. If you are interested in discussing it please reach out.
My earlier work was on the mathematical foundations of deep learning including optimization, loss landscapes, over-parameterization and inductive bias. I received my Ph.D. in Mathematics from the Courant Institute at NYU and later worked at EPFL and ENS Paris as a postdoctoral fellow in the Simons Collaboration on Cracking the Glass Problem.
Research
The papers below fall loosely into three groups, with considerable overlap between them. A recurring question is how technical systems represent disagreement and when benchmarks, datasets or alignment methods instead obscure or collapse it.
Representation, participation and datasets
On who gets represented in a dataset and how: Collection, curation and participation choices shape what a model can learn and what an evaluation can actually show. Several of these focus specifically on speech data and queer representation, a domain in which standard collection pipelines often break down.
- Towards Participatory Speech Dataset Curation: A Queer Case Study and Conceptual Framework. Brooklyn Sheppard, Anaelia Ovalle, Adina Williams and Levent Sagun Interspeech 2026
- Queer Inclusion in Speech Datasets: An Audit and Taxonomy of Practical Tensions. Brooklyn Sheppard, Anaelia Ovalle, Adina Williams and Levent Sagun Interspeech 2026
- The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models. Anaelia Ovalle, Krunoslav Lehman Pavasovic, Louis Martin, Luke Zettlemoyer, Eric Michael Smith, Kai-Wei Chang, Adina Williams and Levent Sagun FAccT 2025
- On the Lack of Queer Voices in Diverse Speech Datasets. Brooklyn Sheppard, Anaelia Ovalle, Adina Williams and Levent Sagun SSaLM 2025 and Speech AI for All Workshop at CHI 2025
- On the Role of Speech Data in Reducing Toxicity Detection Bias. Samuel Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun and Marta R. Costa-jussà NAACL 2025 - Dataset
Measurement and validity
On the validity of benchmarks: do metrics actually measure what they claim to measure? Some propose new ways to test validity directly, others identify where an existing measure breaks down or is used beyond the conditions under which it remains meaningful.
- Automated Bias Blind Spots: Examining Stereotype Associations in VLM-as-a-Judge Paradigms. Rachel Hong, Anaelia Ovalle, Megan Ung, Evangelia Spiliopoulou, Levent Sagun, Adina Williams and Candace Ross FAccT 2026
- Issues in Measuring the Fairness of Social Representation in Synthetic (Speech) Data. Arjun Subramonian, Brooklyn Sheppard and Levent Sagun Synthetic Data Workshop at Aarhus Decennial Conference 2025
- A Differentiable Rank-Based Objective for Better Feature Learning. Krunoslav Lehman Pavasovic, David Lopez-Paz, Giulio Biroli and Levent Sagun ICLR 2025
- Reassessing the Validity of Spurious Correlations Benchmarks. Samuel J. Bell, Diane Bouchacourt and Levent Sagun Preprint 2024
- On Generated vs Collected Data. Levent Sagun, Kartik Ahuja, Elvis Dohmatob and Julia Kempe Global AI Cultures Workshop at ICLR 2024 - pdf
- Weisfeiler and Leman Go Measurement Modeling: Probing the Validity of the WL Test. Arjun Subramonian, Adina Williams, Maximilian Nickel, Yizhou Sun and Levent Sagun Preprint 2023
- Fairness Indicators for Systematic Assessments of Visual Feature Extractors. Priya Goyal, Adriana Romero Soriano, Caner Hazirbas, Levent Sagun and Nicolas Usunier FAccT 2022
Model failure modes
These papers look at specific ways large models fail: them relying on spurious signals, being brittle to superficial changes, amplifying bias already present in training data or generating inconsistent reasoning across languages.
- LLM Knowledge Is Brittle: Truthfulness Representations Rely on Superficial Resemblance. Patrick Haller, Mark Ibrahim, Polina Kirichenko, Levent Sagun and Samuel J. Bell COLM 2026
- Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models. Chantal Shaib, Vinith M. Suriyakumar, Levent Sagun, Byron C. Wallace and Marzyeh Ghassemi NeurIPS 2025, Spotlight
- Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages. Anaelia Ovalle, Candace Ross, Sebastian Ruder, Adina Williams, Karen Ullrich, Mark Ibrahim and Levent Sagun Multilingual Representation Learning Workshop at EMNLP 2025
- An Effective Theory of Bias Amplification. Arjun Subramonian, Samuel J. Bell, Levent Sagun and Elvis Dohmatob ICLR 2025
- Simplicity Bias Leads to Amplified Performance Disparities. Samuel J. Bell and Levent Sagun FAccT 2023
The complete publication list also includes my earlier work on optimization, loss landscapes, over-parameterization and inductive bias: see CV or Google Scholar
Mentoring and Teaching
I work with Ph.D. students, postdoctoral researchers and interns on fairness, representation, measurement and the social/institutional questions raised by AI systems.
Ph.D. interns
- Nasanbayar Ulzii-Orshikh, 2026
- Chantal Shaib, 2025
- Brooklyn Sheppard, 2024
- Arjun Subramonian, 2022 and 2024
- Elia Ovalle, 2023
- Sam Bell, 2021
- Berfin Şimşek, 2020
Ph.D. students
- Nicole Osayande, 2026-today
- Krunoslav Lehman Pavasovic, 2024-2025
- Stéphane d'Ascoli, 2019-2022
Postdoctoral researchers
- Sam Bell, 2022-2023
Teaching
I have taught courses in probability, statistics, machine learning and data science at NYU's Courant Institute and Center for Data Science. I have also given invited lectures and short courses on deep learning and AI at EM Normandie, ENS-PSL and Les Houches.
Community and Organizing
I organize workshops and research communities that bring different disciplinary and intellectual traditions into conversation.
-
Co-Chair,
CRAFT,
FAccT 2026
A program for critical and community-driven engagement with computing and AI, centered on power, governance and collective imagination.
-
Co-Organizer,
Society and Responsible AI Seminar Series, FAIR, 2020–present
An interdisciplinary seminar series on the social and political dimensions of AI research.
-
Co-Organizer,
Communication Across Communities in ML Research and Practice,
FAccT 2022
A workshop on communication across technical, social-scientific and affected communities.
-
Co-Organizer,
Science and Engineering of Deep Learning,
ICLR 2021
A workshop on scientific practice, disciplinary values, and the social consequences of deep-learning research.
-
Co-Organizer,
Science Meets Engineering in Deep Learning,
NeurIPS 2019
A workshop connecting theoretical understanding with the engineering and experimental practice of deep learning. Workshop report.
-
Co-Organizer,
Theoretical Advances in Deep Learning,
IMBM at Boğaziçi University, 2019
A multi-day workshop on the mathematical foundations of deep learning.
Contact
Email: leventsagun@{gmail or meta}.com