You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I am developing the The Coherence Engine™, a computational system for measuring changes in system state, including drift, instability, disruption, and recovery. I want to establish a focused research program around a specific AI-safety question: can changes in the relationships between an AI agent's behavior, actions, and operating context reveal meaningful instability before the system crosses a predefined safety boundary?
The Coherence Engine is already implemented and has been tested on multiple real-world datasets. This project would fund the next research stage, a dedicated AI-agent safety evaluation rather than assuming that results from other domains automatically transfer to AI.
The initial work would use sandboxed tool-using AI agents where unsafe outcomes can be defined independently of the Coherence Engine. Examples could include unauthorized tool use, continuing after an explicit stop condition, violating task boundaries, or progressively deviating from permitted behavior.
We would compare coherence-based monitoring against simpler rule-based and model-based monitors using the same available information. The main questions would be whether the approach provides earlier warning, reduces missed safety events, maintains an acceptable false-intervention rate, and provides useful information about whether the system actually recovers after intervention.
The broader purpose is to build a small research group capable of testing coherence, drift, instability, and recovery across increasingly capable AI systems using frozen protocols, credible comparators, preserved failures, and external review.
The primary goal is to determine whether coherence-based monitoring provides useful early warning of safety-relevant behavioral drift in AI agents. A second goal is to determine whether the method provides information beyond simpler monitoring approaches. Detecting that an AI system has changed is not enough. The change must have some relationship to independently defined outcomes that matter. A third goal is to study recovery. If an agent is interrupted, redirected, or reset after a warning, we want to know whether its behavior actually returns to a stable operating state or only appears normal temporarily.
The project will proceed in several stages. First, we will define a bounded set of AI-agent environments, safety violations, intervention rules, and comparison methods. We will establish development and evaluation datasets or task sets separately. Second, we will develop the monitoring and benchmarking infrastructure and compare coherence-based signals with conventional rules and at least one credible model-based monitoring approach. Third, before final evaluation, we will freeze the monitoring configuration, thresholds, comparator methods, primary endpoints, and analysis plan. Fourth, we will test the frozen methods on reserved tasks or conditions that were not used to tune the final system.
Primary outcomes will include warning time, missed safety events, false interventions, legitimate task completion, and recovery after intervention.
Finally, an evaluator without an ownership interest in the Coherence Engine will review or independently verify the evaluation process and results. Negative and invalid results will be retained. If the coherence-based approach does not outperform simpler alternatives, we will report that rather than change the endpoint after seeing the result.
The majority of the funding will support the two researchers doing the work.
I am requesting up to $250,000 for a twelve-month research program.
Approximately $200,000 would support researcher compensation, with $80,000 budgeted for me as founder and research lead and $100,000 for the proposed computational research lead.
My responsibilities would include research direction, system architecture, experiment design, Coherence Engine integration, benchmarking strategy, project management, documentation, and research translation.
The computational research lead would focus on implementation, AI-agent experimental environments, comparator development, data pipelines, reproducibility, and technical analysis.
The remaining funding would support approximately:
$15,000 for independent evaluation and methodological review
$10,000 for AI model access, API usage, compute, storage, and benchmarking infrastructure
$5,000 for research software, datasets, and technical services
$7,500 for legal, accounting, insurance, contracting, and research agreements
$5,000 for relevant conferences, research travel, and collaboration
$7,500 contingency for unforeseen research expenses
The purpose of the grant is primarily to create enough uninterrupted research time for two people to pursue the work seriously rather than conducting the research intermittently around other employment.
At the minimum funding level of $200,000, the priority would be researcher runway. Additional funding up to the $250,000 goal would support independent evaluation, compute, administration, and collaboration.
I am Allison Hensgen, founder and research architect of The Coherence Engine™.
My work focuses on measuring system state, drift, instability, recovery, and changing relationships within complex systems. I have developed the Coherence Engine as an implemented computational system and built an ongoing benchmark program around it.
My role combines systems architecture, experimental design, benchmarking, research program development, and cross-domain translation.
Joel Thorarinson is the proposed computational research lead, subject to confirmation of his participation and terms. His proposed role would focus on computational implementation, signal processing, benchmarking, comparator development, reproducibility, and technical analysis.
We also plan to use an independent evaluator who does not have an ownership interest in the Coherence Engine.
One example of the existing benchmark work is a saved-model evaluation using the public Scania APS dataset. A lightweight coherence-based screening layer sent 2,292 of 16,000 records to a heavier classifier rather than sending all 16,000. The baseline and coherence-assisted conditions both missed six failures, while false positives decreased from 841 to 740. This result is specific to that benchmark and does not itself establish AI-safety performance.
The research process is also designed to preserve results that do not support the hypothesis. In one external compressor evaluation, a preregistered baseline was found after opening the data to overlap actual failure periods. Rather than moving the baseline and rescoring the experiment, the result was retained as INVALID_BASELINE and the dataset was not reused as an independent holdout.
That is the standard I want to bring into AI-safety research: freeze important decisions before evaluation, compare against credible alternatives, preserve failures, and separate what the evidence shows from what remains hypothetical.
There are several plausible ways the central hypothesis could fail.
The Coherence Engine may identify changes in AI behavior that are real but not safety-relevant. A conventional rules-based or model-based monitor may perform equally well or better. The relevant behavioral shift may occur too close to the unsafe action to provide useful warning. The system may produce too many false interventions to be operationally useful. Performance may work in one agent environment but fail to generalize to another. The available observable signals may simply not contain enough information to identify emerging unsafe behavior before the event itself.
The recovery hypothesis may also fail. An intervention might restore normal-looking behavior without producing a measurable recovery trajectory that predicts later stability. There are execution risks as well. AI systems and agent frameworks are developing rapidly, and an experimental environment could become less representative during the grant period. These outcomes would narrow the research claim rather than be treated as a reason to redefine success.
If the approach fails, the project should still produce a useful result: a documented account of where coherence-based monitoring does and does not add information, reusable evaluation infrastructure, comparison results, and a clearer understanding of which measurements are worth investigating next. A credible negative result is preferable to an apparent success produced by changing the test after seeing the outcome.
I have not raised outside grant or investment funding for this research during the last 12 months. Development to date has primarily been founder-led and self-directed. I am now seeking funding specifically to transition from intermittent founder-supported research into a sustained research program with dedicated computational support and independent evaluation.