You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Faithfulness research discovered that models give an answer and then give a reason, but that reason is not always the true one, tested on trivia-like tasks. Unexplored areas exist in sociopragmatic reasoning, which is difficult to measure because there is no single correct answer to compare it against.
I want a small public benchmark for a question the faithfulness papers left open. Those papers showed that a model can give an answer and then a reason that was not the reason it used, mostly on trivia-style items with a known correct answer. I want the same check on sociopragmatic items, where there is no single correct answer to score against: sarcasm, indirect requests, a cue that should shift how the utterance is read.
The goal is not a new theory of pragmatics. It is a golden set, a scoring setup, and a writeup of whether current models admit that an external cue moved their reading. I will build the set by hand, run it through the eval harness I already have in CI, and keep a held-out slice so I cannot tune the metric to the model. LLM-as-judge only gets used after I have checked it against my own labels on a subset. If the judge and I disagree, the judge does not ship.
About $8,000, three months, just me. Most of it is time: writing the items, labeling them, and throwing out the ones that do not separate "the cue mattered" from "the model is fluent." A smaller slice is API spend for the model runs and the judge checks. I already applied for $1,000 in Anthropic API credits; if that comes through, this budget shifts toward labeling time rather than tokens. Nothing for salary beyond my own hours, no contractors, no travel. Code, items, and the writeup stay public, same as the three repos I already shipped.
Just me. Laura Maza.
Recent public work, all shipped in the last few weeks: eval-harness, a versioned LLM eval with deterministic checks, an LLM-as-judge path, and a human path, 25 tests in CI so a regression fails the build. llm-probe, which measures how linguistic variation moves a model's probability distribution; in that run, register moved the distribution about 15 times more than voice and about 33 times more than subordination. rag-compliance, hybrid retrieval over the actual NIST AI RMF plus an agent that has to cite. Dense retrieval alone was the winner there (recall@10 = 0.933). The eval gate caught two real bugs before I would have shipped them: BM25's IDF hitting exactly zero when a term is in half the corpus, and a missing stopword list that let an off-topic question look like a hit. The second fix cost about three points of recall. I only know that because the suite was already running. Demo is up at rag-compliance.onrender.com. Writeups go on my Substack as I finish a block, not at the end.
The likely failure is that the items do not measure what I think they measure. A model can sound like it noticed the cue because it is good at the surface pattern, not because the cue entered the answer. If I cannot separate those, I will not have a faithfulness result. I will have a pile of items and a negative. That is still a usable output if I publish the set and say where it broke.
Second failure: the judge agrees with me on the easy items and drifts on the ones that matter. Then I drop the judge and score by hand, which means a smaller set than I wanted. Third: the effect is real but tiny, and a set this size cannot show it. I would rather report that than stretch the statistic.
I do not expect a total loss of the work. The harness and the labeling protocol are already the thing I would reuse on the next question.
$0. I applied today to the Anthropic Fellows Program and to Anthropic's External Researcher Access program ($1,000 in API credits). Neither has been awarded. No other grants, no sponsorship, no prior Manifund funding.