Agent Memory Is a Surface for Endogenous Authorization Laundering
Cerruti, Okamoto, and Erol
arXiv, 2026
Work on an AI research project under the mentorship of experienced researchers
Applications have closed. Please come back next year!
The Summer of Open AI Research is a 5-week, fully online research program sponsored by EleutherAI.
We invite people with little research experience to contribute to open science under the mentorship of experienced researchers. Participants work on one of the projects below and are credited on work that may result in publication.
The program is fully online and coordinated through the EleutherAI Discord.
The current project list spans interpretability, AI safety, AI for science, information retrieval, computer vision, and generative modeling.
Anyone may apply. Applicants are considered based on their ability to contribute to a project.
Yes. A central goal of SOAR is to give people outside academia their first research experience.
The commitment differs by project. Estimated weekly hours are listed with each project.
No. Here, open AI means research conducted openly and collaboratively.
Pranava Kumar · MIT CSAIL Kellis Lab
This project studies whether open-weight reasoning models can arrive at correct answers for the wrong reasons, and whether interpretability tools can help detect that failure mode. Participants will create matched reasoning problems in clean, subtly hinted, and misleadingly hinted versions, then compare final answers, written reasoning, activations, and sparse-autoencoder features across conditions. The goal is to test whether internal evidence can distinguish genuine reasoning from shortcut-driven reasoning when surface behavior is misleading.
Participants should have Python experience, basic machine learning knowledge, and familiarity with PyTorch or Hugging Face. Participants should be comfortable with basic data analysis, Git/GitHub workflows, and careful reading of technical material. Careful experimental thinking is especially important.
Aman Kumar, Mike Smith · IUCAA, India
This project studies the spectral geometry of foundation-model representations, focusing on effective rank, intrinsic dimensionality, and semantic compression in embedding spaces. Participants will investigate whether independently trained vision and language models converge toward similarly low-dimensional semantic structures despite differences in architecture, training objective, and dataset. The project extends recent work on representational convergence, multimodal semantics, and the geometry of learned representations.
Participants should have Python experience and introductory machine learning knowledge. Participants should be comfortable with Linux command-line tools, Git/GitHub workflows, and reading and writing technical documentation or research notes.
Kshitij Duraphe, Mike Smith, Shashwat Sourav · UniverseTBD
This project extends an award-winning NeurIPS 2025 workshop paper that asks whether larger foundation models converge on shared internal representations despite never being trained on astronomy data. The project extends this work by applying mechanistic interpretability to astronomy: studying when physical quantities such as galaxy morphology, redshift, stellar mass, and metallicity emerge inside models, how they change layer by layer, and whether astronomy-specific and general-purpose models decompose galaxies into similar features. Participants will investigate representations within and across models, and will ultimately produce an open-source feature dictionary that working astronomers can use to inspect what their models are keying on.
Participants should have strong Python experience beyond notebooks, including writing modules, debugging, and working with research code. Participants should be comfortable with PyTorch training loops, mixed precision, model-loading libraries such as Hugging Face or timm, Linux, Git/GitHub, SSH-based remote workflows, and scientific Python tools such as NumPy, SciPy, Matplotlib, and pandas. Astronomy background is not required.
Elena Golimblevskaia, Gonçalo Paulo · Fraunhofer HHI
Several recent methods aim to recover hierarchical structure from LLM activations: Matryoshka SAEs and Temporal SAEs impose nested reconstruction bottlenecks, while Temporal Feature Analysis extracts hierarchies via post-hoc clustering of predictive codes. None has been quantitatively evaluated for whether recovered hierarchies form coherent parent-child structures on real LLMs. This project builds a coverage-based diagnostic for hierarchy quality, applies it across these methods on Gemma-2-2b, and uses controlled PCFG experiments to isolate which properties of natural language cause hierarchy-recovery methods to fail.
Participants should have Python and PyTorch experience, familiarity with sparse autoencoders, and comfort with the Hugging Face ecosystem. Participants should be comfortable using SAELens or TransformerLens, training a small transformer from scratch, and doing basic data analysis in Python.
Gonçalo Paulo · EleutherAI
This project replicates results from the SAE multiple-seeds paper and investigates why they differ from later results in another study. Participants will compare features across SAE training seeds, starting with TopK SAEs trained on small models, and extend the comparison methodology to activation overlap, CKA, or SVCCA. If time permits, the project may move on to larger models or more recent SAE architectures.
Participants should have Python and PyTorch experience. Familiarity with sparse autoencoders or SAE training libraries is useful but not required.
Gonçalo Paulo · EleutherAI
This project reproduces experiments from Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations, focusing on activation oracles, paraphrasing, and interventions using the models released by Anthropic. The goal is to get natural language autoencoders working in an open setup, benchmark them on simple tasks, and test whether verbalization-based interventions can causally steer model behavior.
Participants should have Python and PyTorch experience. Familiarity with transformer internals, activation patching, or mechanistic interpretability is useful but not required.
Gonçalo Paulo · EleutherAI
This project asks whether subliminal learning is partly determined by the data order a model sees during pretraining. It uses the PolyPythia suite, which includes Pythia-style models trained with controlled variations in random seed and data order, to test whether animal preferences transmitted through subliminal learning can be traced to pretraining data order.
Participants should have Python and PyTorch experience. Familiarity with model fine-tuning libraries such as PEFT or Unsloth is useful but not required.
Shashwat Sourav, Mike Smith · Washington University in St. Louis
This project asks when embeddings of the same astronomical object are reliably aligned across surveys. Participants will train a model to translate one survey's embedding into another survey's embedding space, then use conformal prediction to put a statistically valid uncertainty region around the translated embedding. Small regions indicate survey agreement; large regions indicate unreliable alignment and possible need for follow-up. Multimodal Universe HATS datasets will help scale the work beyond a fixed paired-embedding benchmark.
Participants should have Python and PyTorch experience. Participants should be comfortable with machine learning and deep learning concepts. A statistics background is a plus but not required.
Manjari Narayan · The Surrogate Science Project
This project audits the scientific reasoning capabilities of AI systems for frontier science and high-stakes scientific decisions. It builds on prior work showing flawed reasoning in open-source models on retrospective definitive experiments in drug toxicity. The project asks whether such benchmarks are useful for interrogating interpretability methods and their claims to be causal, and may also build stronger benchmarks for evaluating or falsifying AI forecasting abilities.
Each fellow takes one interpretability method (linear probes, activation steering, or interchange interventions) and produces a focused study on the existing benchmark. Activation steering and interchange intervention tracks are open to fellows who bring prior experience.
Participants should be able to show open-source Python contributions, experimental design skill, or strong test-driven development for AI workflows. Participants should be comfortable using AI coding assistants to produce high-quality work. ARENA or similar coursework is relevant if backed by completed work.
Suvajit Majumder · Optym
Motivated by recent work on subliminal learning, especially "Subliminal Effects in Your Data: A General Mechanism via Log-Linearity," this project aims to develop a general framework for understanding in-context variants of subliminal learning: subliminal prompting and subliminal chain-of-thought. The goal is to take a model-organism and red-teaming approach, identifying ways to systematically accumulate or enhance artifacts such as entangling numbers so that they meaningfully modify model preferences later in a conversation.
Participants should have experience with model inference APIs and basic familiarity with probing model internals. Participants should be comfortable running or hosting models on GPUs.
Avyukth R. Nilajagi · SPAR, Pitt
This project builds on exploratory work from the SPAR '26 Spring cohort, using ideas from subliminal learning and prompt-induced behavioral shifts. The goal is to develop a cheaper and more predictive heuristic for ranking inoculation prompts, prompts designed to steer model behavior toward or away from undesirable traits such as sycophancy, before fine-tuning. Participants will construct contrastive behavioral datasets, compute prompt-conditioned preference shifts, and analyze those shifts with PCA to recover latent behavioral directions in a model's preference space.
Participants should have strong Python experience, experience working with LLMs, and basic linear algebra and statistics for machine learning. Participants should be comfortable working independently and reading research code. Familiarity with interpretability tools such as TransformerLens and fine-tuning APIs is a plus but not required.
Mika Okamoto, Ansel Erol · Georgia Institute of Technology
When an LLM agent says it will follow a rule, does it actually execute compliant actions, or just produce compliant-sounding text? This project addresses the gap between stated reasoning and concrete tool-use behavior in LLM agents, extending prior work on unstable LLM compliance from chatbots to agents. Participants will investigate factors affecting agentic complicance and the extent to which stated intent reflects actual behavior. Time permitting, the project will extend to coding agents.
Participants should have Python experience, basic familiarity with LLM APIs such as OpenAI or LiteLLM, and familiarity with agentic coding patterns such as tool use or function calling. Prior LLM-agent experience, experimental design or data analysis experience, and familiarity with AI safety or LLM evaluation literature are pluses but not required.
Nahid Alam · Oreon Labs
Build a controllable, photorealistic drone-scene editing and simulation system inspired by ChatSim. The project focuses on preparing drone video data, adapting scene-editing components, and producing a demo that can modify drone footage in realistic ways.
Participants should have Python experience and basic computer vision knowledge. Participants should be comfortable using modern coding tools. Experience with video processing, segmentation, tracking, or multimodal models is a plus but not required.
Enrico Shippole · Teraflop AI
This project develops domain-specific encoders for information retrieval tasks. The goal is to build an information retrieval benchmark, fine-tune domain-specific encoder models, and evaluate scaling behavior for embedding models across permissive and synthetic data.
Participants should have Python and PyTorch experience, and should be familiar with sentence-transformers and training encoder or decoder models. Participants should be comfortable with large-scale training workflows such as SLURM, DDP, vLLM, or SGLang. A strong mathematical background is a plus but not required.
Ioana Ciucă, Aman Kumar, Shashwat Sourav, Mike Smith, Matthieu Le Lain · Stanford University
Vision-language models could help astronomers explore large scientific datasets, but most open VLM pipelines are built on image-text pairs. This project extends AstroLLaVA into a reproducible multimodal alignment pipeline for astronomy, covering spectra, light curves, multi-wavelength image cutouts, and selected X-ray event data. Participants will construct modality-language supervision from catalog labels, redshifts, object classes, metadata, cross-matches, and scientific descriptions, then train and evaluate a LLaVA-style architecture with modality-specific encoders connected to an open language model. The aim is to release a data pipeline, training code, and an evaluation harness as a reusable, open-source framework for future researchers.
Participants should have Python experience, machine learning or deep learning fundamentals, and familiarity with PyTorch or a similar framework. Experience with scientific datasets and Git is useful. Familiarity with Hugging Face, VLMs, or embeddings is a plus but not required.
Christian Zhou-Zheng, Ronald McClellan Jr. · Stanford University
Singing voice synthesis (SVS) is the task of generating a human singing voice from digital input. It is the technology behind Hatsune Miku, Kasane Teto, and other VOCALOIDs and "virtual singers." State-of-the-art SVS models are either closed-source (VOCALOID, SynthV) or use older architectures and train on limited data (DiffSinger, NNSVS). This project aims to bring open-source SVS into the modern era. Participants will develop modern open-source SVS methods by scaling up data pipelines and applying modern architectures and modeling paradigms. The project aims to release a full data pipeline, pretrained acoustic model, and pretrained vocoder.
Data side:
Model side:
Participants should have Python experience and basic experience running machine learning experiments. Participants should be comfortable with PyTorch and signal processing. Audio or music-domain experience is a plus but not required.
Cerruti, Okamoto, and Erol
arXiv, 2026
Okamoto and Erol
arXiv, 2026
Lopardo, Frentzen Salim, Srivastava, Jiang, Sharma, and Arnett
EMNLP Findings, 2026
Rane, Vatsa, Pethe, Aktolun, Li, and Singh
Low-Resource Audio Codec @ ICASSP, 2026
Alam, Murali, Bharadwaj, Liu, Chung, Sharma, A, Kiran, Tam, and Vegesna
I Can't Believe It's Not Better @ ICLR, 2026
Imran and Chatterjee
Workshop on LLM Persona Modeling @ NeurIPS, 2025
Zhou-Zheng, Backsund, Chan, Coventry, Eslami, Goel, Han, Soomro, and Wei
MIREX @ ISMIR, 2025
Rolling decisions may arrive earlier.
Mentors may specify preparation work before the main event.
Each cohort can share its progress with the other cohorts.
“Just after SOAR ended, I went on to be a MATS fellow in a very theory-heavy research stream. I think having participated in SOAR made the difference between me looking like 'aspiring AI/ML researcher with lots of math background' instead of 'physicist bandwagoning to AI'.”
“As someone from industry, this was a great opportunity for me to connect with folks looking to contribute to the AI ecosystem as well as gain a better understanding of academia and ongoing research.”
“SOAR was my first experience in an AI research program. Working remotely with experienced mentors pushed me to think like a researcher, from reading literature critically to designing experiments systematically to knowing when to ask the right questions. I'd recommend SOAR to any early-stage researcher who wants to bridge the gap between coursework and real research.”
“SOAR as a program helped me tremendously as an early career researcher. Not only did it help me with research and introduce me to previously unknown research areas, but it also greatly helped my academic development. My mentor was really kind and helpful, introducing me to research beyond the program's topics. Highly recommend for students or people who intend to go into research.”
“SOAR gave me the opportunity to get my foot in the door and actually contribute to a publication-worthy AI research project. It connected me with talented mentors who helped build the foundations for my research journey.”
EleutherAI
Stanford University
Constellation Institute
EleutherAI
Algoverse
EleutherAI