Research Library

Explore EleutherAI papers across language modeling, interpretability, alignment, policy, multimodal research, security, and computing.

198 publications

Newest first

2026

No Single Tokenizer Feature Reliably Predicts Downstream Language Model Performance

EMNLP Findings

Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?

EMNLP

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

EMNLP Findings

Bergson: An Open Source Library for Data Attribution

EMNLP System Demonstrations

Beetle: Structured Exposure Pretraining in Bilingual Language Models for Modelling L2 Language Processing

EMNLP

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

EMNLP

Agent Memory Is a Surface for Endogenous Authorization Laundering

arXiv

Where Does Social Reasoning Come From? Capability Provenance in Language Models

COLM

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

arXiv

What Helps Agentic Lean Provers? A Trace-Level Attribution Study

Math4AI Workshop @ ICML

Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems

Agents in the Wild @ ICML

What AI Governance Needs from Mechanistic Auditing

TAGIR @ ICML

Quantifying the Effect of Test Set Contamination on Generative Evaluations

Foundations of Deep Generative Models @ ICML

🎙️

Don't Just "Fix it in Post'': A Science of AI Must Study Learning Dynamics

Oral at ICML

ICML

Automated Attribution Graph Interpretation via Probe Prompting

Mechanistic interpretability workshop @ ICML

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

ICML

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

ICML

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

ICML

L1 Influence in L2 Language Models: A Human-centric Approach

Computational Developmental Linguistics Workshop at ACL

Weight Tying Biases Token Embeddings Towards the Output Space

ACL Findings

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

ACL

Challenges to Grassroot Organization Engagement with AI Policy

FAccT

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

NeurIPS Datasets and Evaluations

LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold

NeurIPS

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

NeurIPS Datasets and Evaluations

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

NeurIPS

Attention-Guided Audio Compression for Multimodal LLM

Low-Resource Audio Codec @ ICASSP

The Spatial Blindspot of Vision-Language Models

I Can't Believe It's Not Better @ ICLR

Sparse Autoencoders Trained on the Same Data Learn Different Features

ICLR

Evaluating SAE interpretability without explanations

ICLR

🥈

Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs

Best Paper Runner-Up at BioSafe Workshop

ICLR

How Open Must Language Models be to Enable Reliable Scientific Inference?

NeurIPS Datasets and Benchmarks

Queer NLP: A Critical Survey on Literature Gaps, Biases and Trends

arXiv

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

arXiv

2025

Token Entanglement in Subliminal Learning

Mech Interp Workshop @ NeurIPS

Mitigating Emergent Misalignment with Data Attribution

Mech Interp Workshop @ NeurIPS

Benchmark Disaggregation Improves Interpretability of Training Dynamics

CogInterp Workshop at NeurIPS 2025

The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity

AI4Music @ NeurIPS

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

NeurIPS Datasets and Benchmarks

More of the Same: Persistent Representational Harms Under Increased Representation

NeurIPS

Explaining and Mitigating Cross-Linguistic Tokenizer Inequalities

NeurIPS

Persona-Vector Routing: A Lightweight, Interpretable Guardrail for Mitigating LLM Hallucinations

Workshop on LLM Persona Modeling @ NeurIPS

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

arXiv

Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures

NeurIPS Evaluations and Benchmarks

CCS-Lib: A Python Package to Elicit Latent Knowledge from LLMs

Journal of Open Source Software

RWKV-7 "Goose" with Expressive Dynamic State Evolution

COLM

RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale

COLM

Binary Sparse Coding for Interpretability

arXiv

Open Problems in Mechanistic Interpretability

Survey Certification

TMLR

🏆

A Traditional Approach to Symbolic Piano Continuation

First Place at the Symbolic Music Generation competition

MIREX @ ISMIR

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

ISMIR

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

ACL

Steering Language Model Refusal with Sparse Autoencoders

Actionable Interpretability Workshop @ ICML

Write Code that People Want to Use

CodeML Workshop @ ICML

Evaluating Morphological Alignment of Tokenizers in 70 Languages

Tokenization Workshop @ ICML

🏆

BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization

Best Paper

Tokenization Workshop @ ICML

🎙️

Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?

Oral at TiFA Workshop, ICML

ICML

Automatically Interpreting Millions of Features in Large Language Models

ICML

A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety

Columbia Convening on AI Openness and Safety

When AI Co-Scientists Fail: SPOT—a Benchmark for Automated Verification of Scientific Research

arXiv

🏆

The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Best Paper at NAACL

NAACL

KMMLU: Measuring Massive Multitask Language Understanding in Korean

NAACL

Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon

ICLR

PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs

ICLR

Composable Interventions for Language Models

ICLR

Bridging the Data Provenance Gap Across Text, Speech, and Video

ICLR

Aria-MIDI: A Dataset of MIDI Files for Symbolic Music Modeling

ICLR

Mechanistic Anomaly Detection for "Quirky" Language Models

arXiv

Estimating the Probability of Sampling a Trained Neural Network at Random

arXiv

Examining Two Hop Reasoning Through Information Content Scaling

arXiv

Does Transformer Interpretability Transfer to RNNs?

AAAI

Beyond Release: Access Considerations for Generative AI Systems

arXiv

Transcoders Beat Sparse Autoencoders for Interpretability

arXiv

Slowing Learning by Erasing Simple Features

arXiv

Converting MLPs into Polynomials in Closed Form

arXiv

Can We Partially Rewrite Transformers in Natural Language?

arXiv

Matrix-Driven Identification and Reconstruction of LLM Weight Homology

NeurIPS

Humanity's Last Exam

arXiv

Robin: a Suite of Multi-Scale Vision-Language Models and the CHIRP Evaluation Benchmark

arXiv

Towards Best Practices for Open Datasets for LLM Training

arXiv

2024

🔦

RedPajama: an Open Dataset for Training Large Language Models

Spotlight at NeurIPS

NeurIPS Datasets and Benchmarks

LLM Circuit Analyses Are Consistent Across Training and Scale

NeurIPS

Consent in crisis: The rapid decline of the AI data commons

NeurIPS Datasets and Benchmarks

A Walsh Hadamard Derived Linear Vector Symbolic Architecture

NeurIPS

Understanding Gradient Descent through the Training Jacobian

arXiv

The Foundation Model Development Cheatsheet: A Review of Responsible AI Tools & Resources

Survey Certification

TMLR

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models

arXiv

Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources

Northern European Journal of Language Technology

Language Model Crossover: Variation through Few-Shot Prompting

ACM Transactions on Evolutionary Learning and Optimization

🔦

Linear Representations of Sentiment in Large Language Models

Spotlight at Mechanistic Interpretability Workshop, ICML 2024

BlackBox NLP

Refusal in LLMs is an Affine Function

arXiv

Re-Evaluating Evaluation for Multilingual Summarization

EMNLP

Lina-Speech: Gated Linear Attention is a Fast and Parameter-Efficient Learner for text-to-speech synthesis

arXiv

Balancing Label Quantity and Quality for Scalable Elicitation

arXiv

From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models

Survey Certification

TMLR

Eliciting Latent Knowledge from Quirky Language Models

COLM

Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

COLM

The Case for Co-Designing Model Architectures with Hardware

ICPP

Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU Clusters

arXiv

Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?

ACL

Cold Compress: A Toolkit for Benchmarking KV Cache Compression Approaches

arXiv

Self-Directed Synthetic Dialogues and Revisions Technical Report

arXiv

🔦

Stay on topic with Classifier-Free Guidance

Spotlight at ICML 2024

ICML

Social Choice Should Guide AI Alignment: On Dealing with Diverse Human Feedback

ICML Position Paper Track

🎙️

On the Societal Impact of Open Foundation Models

Oral at ICML

ICML Position Paper Track

Neural Networks Learn Statistics of Increasing Complexity

ICML

Grokking Group Multiplication with Cosets

ICML

Beyond Open vs. Closed: Emerging Consensus and Key Questions for Foundation AI Model Governance

Carnegie Endowment for International Peace

🎙️

A Safe Harbor for AI Evaluation and Red Teaming

Oral at ICML

ICML Position Paper Track

Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers

ICML

GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression

arXiv

Improving Black-box Robustness with In-Context Rewriting

TMLR

Simple and Scalable Strategies to Continually Pre-train Large Language Models

TMLR

Lessons from the Trenches on Reproducible Evaluation of Language Models

arXiv

HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models

COLING

Towards a Framework for Openness in Foundation Models: Proceedings from the Columbia Convening on Openness in Artificial Intelligence

Columbia Convening on Openness in Artificial Intelligence

Democratic Governance of AI Systems and Datasets

Think7 Policy Brief

OpenFold: Retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization

Nature Methods

Attributing Mode Collapse in the fine-tuning of Large Language Models

Workshop on Mathematical and Empirical Understanding of Foundation Models @ ICLR

YaRN: Efficient Context Window Extension of Large Language Models

ICLR

Sparse Autoencoders Find Highly Interpretable Features in Language Models

ICLR

ReLoRA: High-Rank Training Through Low-Rank Updates

ICLR

Quality-Diversity through AI Feedback

ICLR

🎙️

OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text

Oral at MATH-AI Workshop, NeurIPS

ICLR

LLeMA: An Open Language Model for Mathematics

ICLR

Holographic Global Convolutional Networks for Long-Range Prediction Tasks in Malware Detection

AISTATS

Musically Aware Automatic Piano Transcription Using Synthetic Pretraining

arXiv

Investigating the Effectiveness of HyperTuning via Gisting

arXiv

The OpenELM Library: Leveraging Progress in Language Models for Novel Evolutionary Algorithms

Genetic Programming Theory & Practice

Suppressing Pink Elephants with Direct Principle Feedback

arXiv

Comparative Study of Large Language Model Architectures on Frontier

arXiv

2023

🔦

Eliciting Language Model Behaviors using Reverse Language Models

Spotlight at SoLaR Workshop, NeurIPS

SoLaR Workshop @ NeurIPS

Detecting Backdoors with Meta-Models

Backdoors in Deep Learning @ NeurIPS

BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

NeurIPS Datasets and Benchmarks

🔦

The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs

Spotlight at NeurIPS

NeurIPS

🔦

Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors

Spotlight at NeurIPS

NeurIPS

Neural MMO 2.0: A Massively Multi-task Addition to Massively Multi-Agent Learning

NeurIPS Datasets and Benchmarks

LEACE: Perfect linear concept erasure in closed form

NeurIPS

Emergent and Predictable Memorization in Large Language Models

NeurIPS

Prompting Multilingual Large Language Models to Generate Code-Mixed Texts: The Case of South East Asian Languages

Workshop on Computational Approaches to Linguistic Code-Switching @ EMNLP

trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback

EMNLP

RWKV: Reinventing RNNs for the Transformer Era

EMNLP (Findings)

StarCoder: may the source be with you!

VentureBeat’s 2024 Best Enterprise Implementation of Generative AI (Software)

TMLR

2023 Open Source Generative AI Survey Report

Linux Foundation Research

Utilizing Weak Supervision to Generate Indonesian Conservation Datasets

Workshop on Southeast Asian Language Processing @ AACL

Current Status of NLP in South East Asia with Insights from Multilingualism and Language Diversity

IJCNLP-AACL

Representation Engineering: A Top-Down Approach to AI Transparency

arXiv

Towards Meta-Models for Automated Interpretability

arXiv

Reclaiming the Digital Commons: A Public Data Trust for Training Data

AIES

ARB: Advanced Reasoning Benchmark for Large Language Models

arXiv

Recasting Self-Attention with Holographic Reduced Representations

ICML

🎙️

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Oral at ICML

ICML

HyperTuning: Toward Adapting Large Language Models without Back-propagation

ICML

Continual Pre-Training of Large Language Models: How to (re)warm your model?

Workshop on Efficient Systems for Foundation Models @ ICML

Challenges and Applications of Large Language Models

arXiv

GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration

ACL 2023 Demos Track

Crosslingual Generalization through Multitask Finetuning

ACL

BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting

ACL

Evaluating the Social Impact of Generative AI Systems in Systems and Society

arXiv

A Technical Report for Polyglot-Ko: Open-Source Large-Scale Korean Language Models

arXiv

Role-Play with Large Language Models

Nature

Can Transformers Learn to Solve Problems Recursively?

arXiv

Eliciting latent predictions from transformers with the tuned lens

arXiv

🥈

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Outstanding Paper Finalist

TMLR

Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

ICLR 2023

🎙️

ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics

Oral

MATH-AI Workshop @ NeurIPS

Masked inverse folding with sequence transfer for protein representation learning

Protein Engineering Design and Selection

🏆

SantaCoder: don't reach for the stars!

Best Paper

Deep Learning for Code (DL4C) Workshop @ ICLR

2022

What Language Model to Train if You Have One Million GPU Hours?

EMNLP (Findings)

🔦

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Featured Paper at NeurIPS Datasets and Benchmarks Track

NeurIPS Datasets and Benchmarks

🏆

LAION-5B: An open large-scale dataset for training next generation image-text models

Outstanding Datasets and Benchmarks Paper at NeurIPS

NeurIPS Datasets and Benchmarks

🎙️

EleutherAI: Going Beyond "Open Science" to "Science in the Open"

Oral at Broadening Research Collaborations Workshop, NeurIPS

Workshop on Broadening Research Collaborations @ NeurIPS

BLOOM: A 176b-parameter open-access multilingual language model

TMLR

VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance

ECCV

Fooling MOSS with Pretrained Language Models

CIKM

Robust Preference Learning for Storytelling via Contrastive Reinforcement Learning

arXiv

Musical audio samples generated from joint text embeddings

The Journal of the Acoustical Society of America

Data Governance in the Age of Large-Scale Data-Driven Language Technology

FAccT

Datasheet for the Pile

arXiv

You reap what you sow: On the Challenges of Bias Evaluation Under Multilingual Settings

Workshop on Challenges & Perspectives in Creating Large Language Models @ ACL

GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Workshop on Challenges & Perspectives in Creating Large Language Models @ ACL

PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts

ACL 2022 System Demonstrations

🔦

Multitask Prompted Training Enables Zero-Shot Task Generalization

Spotlight at ICLR 2022

ICLR

Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets

TACL

MP-NeRF: A massively parallel method for accelerating protein structure reconstruction from internal coordinates

Journal of Computational Chemistry

2021

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Workshop on Data-Centric AI @ NeurIPS

Cut the CARP: Fishing for zero-shot story evaluation

arXiv

An empirical exploration in quality filtering of text data

arXiv

Towards a Model-Theoretic View of Narratives

Workshop on Narrative Understanding @ NAACL

Fabula Entropy Indexing: Objective Measures of Story Coherence

Workshop on Narrative Understanding @ NAACL

The Hard Problem of Aligning AI to Human Values

The State of AI Ethics Report, Volume 4

Scaling Scaling Laws with Board Games

arXiv

Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm

arXiv

Multiversal Views on Language Models

arXiv

2020

2019