You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
i want to run one experiment that could kill my own research direction
i found that when a language model makes something up it looks different on the inside while it is generating, truthful answers settle early in the network and fabricated ones keep shifting right until the last layer reading that cross layer trajectory separates true from false at 0796 auroc on truthfulqa across Pythia and llama 2 7b
the problem is i have not proven this beats the obvious simpler thing
a plain linear probe on a single layer
single layer probes for truthfulness are established prior work
if a probe does just as well then my whole approach is a probe with extra steps and people should know that
i have not run the comparison because i do not have the compute
the goal is a clean answer to one question, does cross layer trajectory carry signal beyond a single layer probe
i will train single layer linear probes at every layer of llama 3 8b on the same data and the same splits i used for my trajectory features. Same evaluation same seeds same everything and then compare directly
three outcomes and i will publish all three:
trajectory clearly wins and the direction is real
they tie and i say so publicly and stop claiming novelty
probe wins and i write that up as a negative result
i am pre registering this before i run it which is the whole point, nobody publishes the version where their own idea loses
about 15000 dollars of gpu compute
llama 3 8b activation extraction across all layers on truthfulqa and a held out set
probe training at every layer
seed repeats so the comparison is not one lucky run
no salary no equipment nothing else
i am in class 11 and doing this around school either way
the money only removes the compute wall
just me, Parth 16 based in new delhi working outside any university
2 published papers on where knowledge sits inside transformers, Springer CCIS @ HBAI workshop at IJCAI 2026
an oral at AISB 2026 in the UK
arxiv endorsed by Mor Geva
Main finding so far is that non western factual knowledge resolves roughly 30 to 35 percent deeper inside the network than western knowledge, models are not just trained on less of it they store it somewhere else.
I also collaborate with the Kellis lab at MIT CSAIL on embeddings work
and I built project sadak a pothole reporting platform in india that has produced 360 plus reports and 11 confirmed government repairs
disclosure that matters
i am also building a commercial tool on this same signal called veritas
the ablation and its result get published either way regardless of what it does to the product
what are the most likely causes and outcomes if this project fails
most likely cause is that the answer is simply no
a single layer probe does as well and the trajectory adds nothing real
that is not a failure of the project it is the project working
the failure mode i actually worry about is a muddy result
small gap that could be noise where i cannot tell either way
if that happens i will say it is inconclusive rather than round it up in my favour
second cause is that it does not transfer to llama 3
i validated on pythia and llama 2 7b and a newer model may behave differently
one emergent ventures grant supporting the underlying interpretability research and funding my travel to the UK to present my paper
There are no bids on this project.