You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I'm an undergraduate student at the University of Mumbai working on AI safety and mechanistic interpretability. As far as I know, I'm currently the only person at my college working seriously in this area. We don't have a club, reading group, or faculty member working on AI safety, so for the last year I've basically been learning and doing research on my own.
I recently got a sole-authored paper accepted at the MI4MedFM workshop at MICCAI 2026 in Strasbourg on October 1. I really want to attend, but the main problem is getting there. My university doesn't provide travel funding for undergraduates, and I don't have a lab or PI who can support the trip.
My paper looks at whether multimodal models actually use clinical text when making predictions, or whether they can get the right answer using shortcuts. I used activation patching, counterfactual experiments and sparse autoencoders to test this. One of the things I found was that some models kept their predictions even after I changed the clinical text to contradict the image. So the model could give the correct answer while not really using the information we expected it to use.
I started with medical models because it gave me a concrete problem where I could learn mechanistic interpretability properly. But I don't see medicine as the end goal. I want to use these methods on more general AI safety problems, and I've already started doing that through my work on LLMs.
Accepted Paper: https://openreview.net/forum?id=0WdP9Rg0P1
Code: https://github.com/gauravhadavale07/reading-between-the-lesions
LLM Preprint: Causal Auditing of Latent Affect in Language Models (This work applied these methods to general-purpose LLMs)
Workshop Information: MI4MedFM 2026 (Focuses on mechanistic interpretability, sparse autoencoders, and deployment safety)
Personal Portfolio Website
I don't want this to just be a trip where I attend a workshop and come back.
First, I want to write up what I learn and put it online, probably on the EA Forum or LessWrong. I'm not thinking of a normal workshop summary. I'd like to write something like "how I got into mechanistic interpretability as an undergraduate and what I wish I'd known a year ago." Especially for students in India, there isn't much practical information about where to start, what courses to take, what papers to read, or how to find funding.
I'm also planning to give a talk at my college about AI safety and interpretability and try to start a small club or reading group. Right now, most people at my college probably don't even know this field exists. I'd like to change that.
After that, I'd like to help start an AI safety reading group in Mumbai. I'm already connected to the EA India network through Aditya Prasad, and I think getting the first 5–10 people together would be realistic.
For my own research, I want to start applying the techniques I've learned to more general models and questions that are closer to alignment research, especially things like misgeneralization or deceptive behavior. I don't have the exact project figured out yet. That's actually one of the reasons I want to attend the workshop. I think talking to people who are already doing this work would help me figure out what is worth working on instead of trying to guess everything myself.
I'd also like to help other Indian students who are interested in AI safety find the resources, courses, grants and fellowships that I had to search for myself.
It's just me.
I didn't have a local lab or mentor in this field, so I taught myself most of this from papers, open-source code and online resources. I implemented the activation patching, probing, counterfactual ablations and SAE experiments myself.
A lot of the work was done on Google Colab and my personal laptop because I don't have access to GPUs through my university.
I also cold-emailed a professor in Denmark asking for feedback on the paper. I made the revisions based on that feedback and eventually got the paper accepted.
I have also had a sole-authored scientific poster accepted at ECR 2026.
Over the past year I've completed BlueDot's Future of AI course, applied to MATS Winter 2027 in Neel Nanda's stream, and applied to BlueDot's Technical AI Safety Project Sprint. I'm also in contact with Aditya Prasad at AI Safety India / Groundless AI about moving my research more toward alignment and helping build the community in India.
I've raised $0 in the last 12 months. I applied for MICCAI's travel awards but didn't get one. There were only around 15–25 awards available. The government travel schemes I looked at also generally require postgraduate enrollment, which I don't have.
So this would be the first grant I've received.
The methods I used in the paper are directly related to mechanistic interpretability and technical AI safety: activation patching, sparse autoencoders, counterfactual interventions and studying internal representations.
The question I was trying to answer was basically: if a model gives us the right answer, do we actually know why it gave that answer?
In my experiments, I found cases where the model seemed to keep making the same prediction even when the information we expected it to rely on was changed. That's interesting to me because it's a small, concrete example of a broader AI safety problem: we can't necessarily assume that a model is doing what we think it is doing just because the output looks good.
I used medical models because they gave me a manageable way to learn these methods and actually finish a research project. I'm now trying to move toward general-purpose models. My LLM preprint is one example of that, and I've also applied to MATS and BlueDot's Technical AI Safety Project Sprint.
I think attending MI4MedFM would help me make that transition much faster.
I'm paying for food and accommodation myself.
The remaining costs are approximately:
Flights from Mumbai to Strasbourg via Frankfurt: ~$800
MICCAI student registration: ~$185
Schengen visa, VFS charges and travel insurance: ~$180
Local transport: ~$35
Total: ~$1,200
I don't necessarily need the full amount. If I can get around $900 or more, I can work out the rest myself.
The main reason is that I don't have another obvious way to get there.
I'm an undergraduate in India, working on AI safety without a lab, PI or university funding. Despite that, I managed to teach myself the field, do the experiments, write a paper and get it accepted at an international workshop.
The research is already done. The paper is accepted and the code is public. This isn't funding a project that might or might not happen. It's basically funding the opportunity to actually be there, present the work, meet researchers in the field and learn from them.
I'm also hoping that the benefit doesn't stop with me. There aren't many undergraduates in India working on technical AI safety, and I know how difficult it can be to figure out where to start when you don't have anyone around you working in the field. I'd like to use what I learn at the workshop to help other students avoid some of that same difficulty.
For me, this trip would be a pretty important step from doing AI safety research mostly by myself to actually becoming part of the community I'm trying to contribute to.