You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Im testing whether the causal circuits we find in language models are actually properties of the models or partly artifacts of how we test them. i built a pipeline to reproduce published circuits and found that changing the counterfactual changed one result from 3/6 heads to 5/6. i want to see how common this is across models, circuits, interventions, and other experimental choices.
I wanna take the pipeline i already built and use it to stress-test published circuits across different models and experimental setups. i'll automate as much of this as possible, run lots of repeated experiments, and compare how circuits change. the goal is basically to find out where causal circuit discovery works, where it breaks, and why.
Most of the funding would go towards building a research workstation with a high vram gpu, good cpu and ram, storage, and everything else needed for a full build. i'd also keep some aside for cloud compute, model hosting, and other research costs that come up.
Anyways its just me on this lol. i built the causal-interp pipeline myself and used it to reproduce published circuits, then kept messing with it when the results didnt make sense until i found the counterfactual issue.
I think the most likely outcome if it fails is that the counterfactual effect i found turns out to be pretty specific to the docstring task, or that the differences are mostly just noise. if that happens i'll try to figure out what caused it instead of just throwing the result away. It could also show that the circuits are actually pretty stable once you control for certain things, which would still be a useful result.
$0 raised in the past 12 months