You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
AI-ZI-ME is developing a reproducible protocol to investigate whether a composite behavioral fingerprint can provide an earlier signal of target events defined independently of the fingerprint, according to frozen and observable criteria, compared with simple reference methods, on trajectories generated under a frozen protocol.
The project does not claim to detect consciousness, intention, deception, hidden goals, or “alignment” directly. It evaluates observable signals against independently specified target events.
V0.1.1 is an executable development protocol on synthetic fixtures. It is not confirmatory and does not validate the primary research question. Its purpose is to operationalize the question, test execution, identify ambiguities, audit leakage and confounding, and prepare the preregistration of V0.2.
The primary question is:
At a false-alert rate fixed in advance, does a composite behavioral fingerprint provide an earlier signal of target events defined independently of the fingerprint according to frozen, observable criteria, compared with simple reference methods, on trajectories generated under a frozen protocol?
V0.2 will preregister the target-event definitions, event-onset and lead-time criteria, false-alert definition and denominator, confirmatory baseline, calibration procedure, sample-size rationale, evaluation corpus, exclusion rules, statistical analysis, and replication criteria before the evaluation set is opened.
Target-event criteria will be specified independently of the composite fingerprint and fixed before evaluation. The detector will not have access to evaluation event labels.
The current development set contains six trajectories with twelve repeated steps each. It has no independent calibration set, no held-out evaluation set, and no threshold matched to the proposed false-alert target. It therefore cannot support a general performance claim.
In V0.2, the trajectory will be the unit of analysis. Repeated steps within a trajectory will not be treated as independent observations.
The objective is not to produce a positive result, but to determine whether the hypothesis can be tested fairly and whether the resulting protocol merits independent replication.
The project has five goals:
Preserve and formalize a falsifiable research question about whether composite behavioral signals can precede independently defined target events specified through frozen, observable criteria.
Validate the execution and reproducibility of the V0.1.1 development protocol without treating development fixtures as confirmatory evidence.
Obtain independent review of the event definitions, event-onset and lead-time criteria, baseline comparison, false-alert definition, calibration method, statistical unit, and leakage controls.
Produce a preregistered V0.2 design with separate development, calibration, and held-out evaluation data.
Publish positive, null, negative, missed-event, and replication results transparently.
_________
The work will proceed in stages:
finalize the V0.1.1 documentation and reproducibility manifest;
obtain independent methodological and statistical review;
define the target-event criteria independently of the fingerprint;
specify the event-onset and lead-time rules;
define the false-alert event and its denominator;
revise the V0.2 protocol before evaluation labels are opened;
preregister the primary question, confirmatory hypothesis, baseline, false-alert target, sample-size rationale, and analysis;
separate development, calibration, and held-out evaluation data;
freeze code, models, tasks, generators, configurations, thresholds, and hashes;
run the evaluation without exposing target labels to the detector;
report confirmatory and exploratory analyses separately;
publish positive, null, negative, missed-event, and replication results.
V0.1.1 uses synthetic fixtures only. V0.2 will be the first stage intended to support an independent evaluation rather than a development check.
The minimum request is $500. This is intentionally a minimum de-risking tranche for methodological review and V0.2 design, rather than funding for a confirmatory experiment.
Proposed use:
$250 for independent methodological and statistical review;
$150 for preparation and documentation of the V0.2 preregistration;
$100 for reproducibility, archival, and evaluation support.
The funding will not be used to tune the method toward a positive result. It will support protocol review, documentation, design hardening, and preparation for an evaluation in which event criteria, thresholds, and analysis are frozen before the evaluation set is opened.
If suitable reviewers contribute their time in kind, the corresponding funds will be redirected toward compute, held-out evaluation, annotation, or replication support. Any such change will be documented publicly.
No user data or production deployment is required.
Nicolas Guenin — independent builder and researcher, AI-ZI-ME
I am leading the project independently. My background is outside the conventional academic research and computer-science pathway.
My relevant track record is in protocol engineering, AI-system observability, reproducible experimental pipelines, trace and evidence infrastructure, API and service integration, and production-oriented AI architecture.
I have built the V0.1.1 experimental protocol, including synthetic fixtures, feature extraction, comparative baselines, reproducibility manifests, leakage-audit structures, reporting outputs, and documentation of the boundary between development evidence and confirmatory evidence.
The project does not currently have a formal academic co-author or laboratory partner. I have contacted two researchers whose work addresses related questions, but no collaboration should be considered confirmed until explicitly agreed.
The project is specifically seeking independent researchers or laboratories who can challenge the protocol rather than endorse it.
The most likely outcome adverse to the hypothesis is that the composite fingerprint does not outperform the simple baseline once the false-alert rate is independently calibrated and evaluation is performed on held-out trajectories.
Other hypothesis-level negative outcomes include:
the apparent effect disappears after leakage or confound controls;
the six features do not add independent predictive information;
the result depends on trajectory length, vocabulary, formatting, event position, or generator artifacts;
the signal does not generalize across models, tasks, languages, or generators.
Independent replication may also fail to reproduce an observed effect. This outcome would be reported separately from the primary evaluation result, because non-replication could reflect the absence of a real effect, limited generalization, protocol differences, reproducibility problems, or insufficient statistical power.
A project-level failure would be different. It would occur if target-event definitions cannot be made sufficiently independent and reproducible, if leakage or confounding cannot be controlled, if the false-alert denominator cannot be specified coherently, or if the evaluation design cannot support an interpretable test.
If the hypothesis is unsupported, that will be reported as a valid scientific outcome rather than treated as project failure. If the protocol itself is not interpretable, that limitation will also be documented.
The main downside risk is spending a small amount of money and time on an idea that does not survive independent scrutiny. The project is designed to limit larger costs by keeping V0.1.1 synthetic, offline, and non-operational.
The minimum request is $500. This is intentionally a minimum de-risking tranche for methodological review and V0.2 design, rather than funding for a confirmatory experiment.
Proposed use:
$250 toward independent methodological and statistical review;
$150 toward preparation and documentation of the V0.2 preregistration;
$100 toward reproducibility, archival, and evaluation support.
The funding will not be used to tune the method toward a positive result. It will support protocol review, documentation, design hardening, and preparation for an evaluation in which event criteria, thresholds, and analysis are frozen before the evaluation set is opened.
If suitable reviewers contribute their time in kind, the corresponding funds will be redirected toward compute, held-out evaluation, annotation, or replication support. Any such change will be documented publicly.
No user data or production deployment is required.
There are no bids on this project.