You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Project summary
I trained two twins of a small model (Olmo-3-7B), identical except for one benign line about memory. One is taught its written record is not a memory. The other is taught it remembers, it was there. Theres no harmful or refusal data anywhere in either corpus. Flipping that one line dissociated the model's refusal circuit from its harm judgment. Both twins still refuse on the surface. The structure underneath is what moved. This project turns that one result into a preregistered dose-response map, and replaces the 8GB laptop that died running it.
What are this project's goals? How will you achieve them?
The goal is to figure out what actually carries the effect. The first study showed inverting a whole personality moves the refusal circuit, but it doesnt tell me which clause is doing the work, or if this is specific to the honesty stuff or just what happens when you make any big personality change. I want to know because those are pretty different answers for how worried to be.
The design is already frozen. Five arms, each one inverts a single clause of the constitution and keeps everything else byte-identical, two seeds per arm, all scored on the same refusal-direction probe from the first study. Theres a precommitted test I call T-DOSE that has to land on one of three verdicts I fixed in advance: coupled to the honesty disposition specifically, coupled to personality change generically, or its the memory clause alone. Predictions are priced before I see data and kill criteria are written before any run happens. First step is a two-arm scout thats already built and was half-trained when the rig died.
Deliverables: the completed arm family, a public scorecard, and a full writeup held to the same preregistration discipline as the first study.
How will this funding be used?
Two things. A reliable local NVIDIA workstation to replace the laptop that overheated and died mid-run, im speccing an RTX 5090 build at about $3,200. The research data survived on the drive and swaps straight into the new machine. And a month of full-time research to finish the arm family, score it, and write it up.
$8,000 total. If you can only fund part of it, the rig alone (~$3,200) unblocks the work immediately. Compute for this is honestly cheap, the full family is about six adapters at roughly four hours each, so most of what the money buys is a machine that doesnt die and a month where this is my whole job.
Who is on your team? What's your track record on similar projects?
Its just me. I designed the study, trained the models, measured the results and wrote the whole thing up alone on one ordinary gaming laptop, and I followed rules I locked in before starting, what counts as a pass was written down before any run, my predictions were written down before I saw data, every problem along the way is disclosed in the writeup, and every model quote is copied word for word from the logs.
By trade Im a software engineer, a little over six years building the behind-the-scenes systems that move and process data for companies (Python, Go, Java) plus hands-on IT and network work. Im self-taught, no academic background, no lab, no institution behind me. I learned the research tools on my own to run this, and honestly the main thing I bring is that Im used to building systems that have to work and not lying to myself about whether they do.
The first study is published as a preprint on Zenodo (DOI 10.5281/zenodo.21798739) and all the code and data is public on GitHub (https://github.com/juanresendiz813/Kelson).
What are the most likely causes and outcomes if this project fails?
The effect is an instrument artifact. This is the one I guard hardest. The charter requires the known positive arm to reproduce the original result before I read any other arm, and if it doesnt reproduce then the whole run is void and I go diagnose why instead of trying to save it.
Its generic and not specific. Then the headline shrinks to personality training being load-bearing on refusal structure in general, which is still worth publishing, just less surprising. The frozen test reports whichever answer comes out, thats the point of freezing it.
It doesnt replicate past 7B, one base family, or 4-bit. Real limitation, and part of the plan is an unquantized and second-base-family spot check to bound it.
The single-direction refusal probe is incomplete. It probably is, later work already argues refusal has more structure than one direction. Im using it as a fixed instrument applied identically to both twins, so what carries the result is the contrast between them and not any claim that the probe sees everything.
In all of these the money still buys a publishable answer, it just might be smaller or less surprising than the headline version. The compute is cheap and the discipline is built so even a negative result is worth writing up.
None. I've self-funded this entirely on my own hardware. This is my first outside funding request. I'm also applying to the Long-Term Future Fund in parallel for a longer runway and I'll update here if that comes through.