You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I’m building and testing SGOS, a model-agnostic governance layer designed for high-stakes, AI-assisted decisions.
The core idea is straightforward: just because an AI model is capable enough to suggest an action doesn’t mean it should automatically have the authority to execute it. SGOS sits on top of the model as a separate decision-authority layer. Depending on the supporting evidence and how reliable the situation is, it can greenlight the action, flag it for human review, or withhold the action.
The Next Step: IRV-1
We are moving into the IRV-1 stage, which is a preregistered validation program for the SGOS Reference Core. The main goal here is to see if this setup actually makes decisions safer compared to much simpler baselines, and honestly document exactly where it breaks.
I want this evaluation to be strictly falsifiable. If we get negative, null, inconclusive, or limitation findings, we are going to report them exactly as they are—no spinning failures into successes. We are also separating founder-led engineering evidence from future independent reproduction and external review to keep the findings clean.
To Be Clear
IRV-1 isn’t trying to prove the entire SGOS architecture works perfectly. It’s a tight, focused test to answer one specific question: can an external authority layer stop unjustified AI actions without making the system overly defensive and paralyzed?
The main goal is to find out whether SGOS provides real decision-safety value beyond simpler alternatives.
I will do this through IRV-1, a preregistered validation program. The evaluation will compare SGOS with simpler governance baselines such as always-proceed, confidence-only, evidence-only, coverage-only, and a strong simple baseline selected before final evaluation.
The tests will focus on false-proceed, false-withhold, proceed rate, calibration, missing evidence, distribution shift, and controlled failure cases. Where appropriate, comparisons will be made at similar proceed rates so SGOS cannot look safer simply by withholding more actions.
The Reference Core will remain fixed during the evaluation. Thresholds, baseline-selection rules, and statistical procedures will be set before the final evaluation set is used.
A second goal is reproducibility. I plan to move from founder-led engineering evidence toward independent reproduction, external technical/statistical review, locked environments, and auditable artifacts.
The desired outcome is a clear answer about when this approach helps, when it does not, and what its main failure modes are.
Funding would support the parts of IRV-1 that are difficult to do reliably as a self-funded, founder-led project.
The main uses would be:
- independent reproduction of the existing evidence package;
- external technical and statistical review;
- compute and reproducibility infrastructure;
- controlled evaluation and failure testing;
- data and research-related costs;
- dedicated founder research time and project coordination.
I would prioritize reproducibility and independent review before expanding the broader SGOS architecture.
If I receive overlapping support from another funder, I will disclose it and adjust the scope or budget so the same costs are not funded twice.
Who is on your team? What's your track record on similar projects?
I am currently the founder and research lead for SGOS, and the project is founder-led.
I designed the SGOS research direction, built the current Reference Core, and produced the existing EPP-1 engineering evidence package and reproducibility artifacts. That work established an initial technical baseline, but it was produced internally and should not be treated as independent validation.
The next stage is deliberately structured to add independence. I plan to recruit an independent reproduction operator and external technical/statistical reviewers who were not involved in developing the SGOS Reference Core. Reviewers will be selected for relevant experience in AI evaluation, reproducibility, calibration or statistical inference, with conflict-of-interest disclosure and role separation where practical.
My track record so far is therefore strongest in building the research system, defining the validation protocol, preserving reproducibility evidence, and documenting limitations. IRV-1 is intended to test that work under stronger external scrutiny rather than relying on founder-led evidence alone.
The project could fail in several informative ways. SGOS may not outperform simpler baselines once proceed rates are matched. It may reduce false-proceed decisions only by causing too many false-withholds, or it may not behave robustly under missing evidence or distribution shift.
Independent reproduction could also uncover implementation or integration limitations that reduce the strength of the original engineering evidence.
If this happens, I will preserve the negative, null, or limitation result and document the failure mode rather than treating it as a success. A failed IRV-1 would still provide useful evidence about the limits of this decision-authority approach and which assumptions do not hold.
I have raised $0 in committed external funding for SGOS in the last 12 months. The project has been self-funded so far.
I currently have several pending grant applications for overlapping research work, but none should be counted as funding raised unless an award is actually made. If I receive support from more than one funder, I will disclose the overlap and adjust the relevant budgets or scopes to avoid double-funding the same expenses.