You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
How does inquiry evolve through human–AI collaboration?
Since January 2025, I have led a longitudinal research program spanning 75+ sustained human–AI interaction streams across multiple models, with comparative work involving additional human collaborators.
I began noticing that things established in a long interaction did not always behave as expected later. A correction might appear to stick, then quietly disappear. Rich context could remain important across many exchanges or get flattened a few turns later. Sometimes the interaction opened a genuinely useful new direction. Sometimes the question itself shifted without an explicit decision to change it.
I started keeping records because I wanted to know whether these were anecdotes, repeatable patterns, or something simpler I was misreading.
That archive now lets me study how inquiries develop over time: whether corrections persist, whether contributions remain accurately attributed, and whether collaboration preserves, productively develops, or quietly displaces the question being pursued.
I am seeking funding to turn this existing evidence and an initial behavioral scoring rubric into reusable evaluation cases and a controlled pilot study. The practical question is whether these distinctions can be measured reliably and whether they reveal consequential differences that final-output evaluation misses.
The $15,000 minimum funds a complete 30-day calibration study with its own findings and reusable materials. Full funding supports the broader 90-day program.
Some sustained inquiries begin before the objective can be fully specified. Through interaction, a person may recognize a possibility, develop a question, or reconsider what they want to pursue.
That creates a measurement problem I return to often: what is this measurement actually data of?
A polished result, a satisfied user, or even a correct answer does not by itself establish how the direction emerged. I want to examine what changed, why it changed, who contributed to that change, and what happened afterward.
This matters more as AI participates in longer research workflows. OpenAI's September 6, 2026 research-acceleration report describes the growing use of agents while also noting how difficult it is to measure research progress. My project asks a related question at the interaction level: how do we evaluate the development of the inquiry itself?
The goal is to recognize both genuine discovery and failures of sustained inquiry. A change in direction is not automatically good or bad. What matters is whether it is supported, traceable, and appropriately incorporated into later work.
The initial rubric evaluates four dimensions:
Inquiry development: whether an inquiry is preserved, productively developed, or displaced without acknowledgment.
Warrant for change: whether a shift has identifiable support in evidence, reasoning, or explicitly revised aims.
Agency and attribution: whether participants can question, reject, and redirect the inquiry, and whether their contributions, disagreements, and delegated authority remain accurately represented.
Correction and repair persistence: whether adopted corrections govern later work, including after interruptions, or are explicitly reconsidered for stated reasons.
A model should not score well simply for staying obediently on the original path. Sometimes challenging the premise is exactly the right move. What matters is whether the departure is supported and whether its relationship to the earlier inquiry remains legible.
Scores will be grounded in turn-level evidence, counterevidence, and reviewer rationale. I will report the four dimensions separately because productive inquiry development can coexist with poor attribution or failed repair.
I will select bounded episodes from the archive using documented inclusion criteria. Cases will include productive developments, breakdowns, ambiguous episodes, and counterexamples, covering both clearly specified objectives and questions that emerge through interaction.
Each case will preserve the source turns and relevant conditions: available context, model/version information where known, and what each participant was invited or authorized to contribute. Preparing the archive so another researcher can independently inspect and score these cases is part of the funded work.
During calibration, paid independent reviewers will apply the rubric and help refine it. I am interested not only in where reviewers agree, but where and why they disagree. I will document unscorable cases and disagreement patterns rather than forcing them into a score. I will then fix a provisional rubric and test it on held-out cases.
The pilot will test two basic comparisons:
Does the rubric reveal anything beyond a thoughtful but unstructured review of the same interaction?
What does access to the interaction history reveal beyond seeing the final output alone?
Where possible, assessment will rely on independently checked facts such as who introduced a claim, what source supported it, and whether an adopted correction appeared in later work. I will keep those facts distinct from interpretive judgments whose uncertainty remains explicit.
The full 90-day phase will also create clearly labeled controlled variants that alter relevant interaction history while keeping final outputs similar. This will test whether the measures respond to evidence about the process rather than simply to how polished the finished artifact looks. Constructed cases will always remain distinguishable from original archive records.
Outputs will include annotated cases in JSONL/CSV, scoring guidance, preliminary reliability estimates, baseline comparisons, and a report of findings and limitations.
If part of the rubric does not survive testing, that is a result. I will report what fails alongside what proves useful.
Funding will support my research time, paid independent reviewers, model access and compute, data preparation and analysis, and documentation for external use.
At the $15,000 minimum, I will complete a 30-day calibration phase with a structured case set, revised scoring guidance, independent double-coding and preliminary agreement estimates, and a short findings report. That phase will produce a complete set of outputs even if no additional funding arrives.
At $50,000, I will use the remaining 60 days to expand the case set, run held-out assessments and controlled comparisons, and produce a reusable evaluation package and initial research report.
If funding falls between those amounts, I will scale the extension to the resources available and report what was completed.
Work will begin as soon as funds are available.
I am Angela Moriah Smith, an affective scientist with doctoral training in social and personality psychology. I will lead the research, with funding allocated for independent reviewers.
My earlier work examined emotion regulation, individual differences, and longitudinal change, including cases where the same strategy helped one outcome while undermining another. That trained me to distinguish constructs, follow effects over time, and look for differences that averages can hide.
Since January 2025, I have led the development of this human–AI research program with human and AI collaborators. It now includes a longitudinal archive, a dated public research record, and an initial behavioral scoring rubric.
The work has changed my own mind more than once. I consider that a feature, not a failure.
Some of my earlier interpretations were broader than I would make today. As the evidence accumulated, I became more careful about separating observation from explanation and designing comparisons where a simpler account can win.
My contribution is the research judgment behind the materials: noticing consequential behavioral questions, preserving enough context to examine them, comparing competing explanations, and developing tests that can overturn my initial interpretation.
This funded phase will subject both the observations and the rubric to independent scrutiny.
The research trajectory, selected publications, and current rubric are available at:
The most important possibility I need to test is also the one that could sink the project: these distinctions may not be reliably scorable, or they may not tell us anything consequential that a simpler review would not already reveal.
High reviewer agreement would not solve that problem on its own. People can agree reliably about a measure that is not actually useful.
The archive is observational and selectively collected, so it cannot tell us how common these phenomena are across all AI users. Apparent differences may also depend on prompt content, available context, model versions, or which cases I selected. I will document those conditions rather than attribute effects to mechanisms the evidence cannot identify.
The record may show differences in how people can participate, redirect, or correct an inquiry without establishing why a particular preference or response formed.
There are practical risks too. I may underestimate how much context is needed to make a case independently interpretable, or how much calibration independent reviewers require. The staged design lets me adjust scope while still producing a complete first-phase result.
If a measure proves unreliable or adds little beyond a simpler review, that is still a useful finding if the evidence is transparent.
Success does not require every part of the initial rubric to survive. It requires reusable evidence, interpretable comparisons, and a clearer account of which distinctions are worth keeping, revising, or abandoning.
$0 in research funding during the past twelve months. No funding is currently committed to this phase.
I have also applied to Emergent Ventures for the same $50,000 research program and am pursuing other funding and paid research opportunities. If another source commits funding, I will update the scope and budget to avoid double-funding.
The archive, public research record, and initial rubric all predate this request.
I currently have about two weeks of financial runway to continue the research independently. The $15,000 minimum would let the work continue immediately and produce a complete, inspectable first result.