You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
My CHRONOS project is an open-source effort to build a structured methodology for evaluating frontier AI systems using reproducible evidence rather than benchmark scores alone.
Each evaluation specimen preserves the original evidence—including raw screenshots, transcripts, logs, and model outputs—before separating fact extraction and reporting into independent layers. This allows anyone reviewing the work to inspect the original evidence and independently evaluate the conclusions instead of relying solely on my interpretation.
The project is already public, with a GitHub repository, published specimens, methodology documentation, and ongoing public development updates. I also publish standalone LinkedIn posts documenting AI failure analyses, and selected specimens are accompanied by NotebookLM-generated explanatory videos that make documented AI failures easier for a broader audience to understand while preserving the original evidence as the primary research artifact.
Over the next six months, I plan to expand the repository from three published specimens to approximately twelve, refine the methodology based on practical experience, and conduct a small pilot external validation to determine whether other researchers can reproduce the workflow using only the published documentation.
My goal is not to replace existing AI benchmarks, but to complement them with transparent, evidence-based qualitative evaluation that can be independently reviewed, challenged, and improved over time.
Public updates on CHRONOS, including the repository, published specimens, AI failure analyses, and ongoing development, are available through my LinkedIn profile:
https://www.linkedin.com/in/sirajudeen-seethapathy-404b10410
The primary goal of CHRONOS is to develop and test a practical methodology for evidence-based qualitative evaluation of frontier AI systems.
Over the next six months, I plan to expand the repository from three published specimens to approximately twelve by systematically evaluating frontier AI models across a range of reasoning, factuality, and safety-related scenarios. Every specimen will preserve the original evidence before applying the CHRONOS three-layer workflow of Raw Evidence, Fact Extraction, and Reporting.
As the specimen library grows, I will refine the methodology based on practical findings, documenting both improvements and limitations transparently rather than assuming the current version is final.
I also plan to conduct a small pilot external validation by inviting independent reviewers to follow the published documentation and determine whether they can reproduce parts of the workflow using only the publicly available evidence and methodology.
All specimens, methodology updates, and progress reports will be published openly so that other researchers can inspect, critique, and build upon the work. My objective is to determine whether CHRONOS can become a practical, reusable framework that complements existing AI evaluation methods by emphasizing transparency, reproducibility, and preserved evidence rather than aggregate benchmark scores alone.
Unlike capability benchmarks or conventional red-teaming exercises that primarily measure model performance or identify isolated failures, CHRONOS focuses on preserving the complete evidence trail surrounding individual AI behaviours. The methodology is designed so that the original evidence, extracted facts, and final reporting remain independent, allowing future researchers to revisit the same specimen using different analytical frameworks without reproducing the original evaluation.
The funding will allow me to dedicate six months of focused effort to developing and validating CHRONOS.
My minimum funding target of US$3,000 will support the core work needed to continue the project, including API costs for evaluating frontier AI models, essential research expenses, and partial research support so I can consistently produce new evaluation specimens and maintain the public repository.
If the project reaches its full funding goal of US$10,000, I will also purchase a laptop suitable for AI research and documentation. To date, all CHRONOS work—including the public repository, published specimens, methodology development, AI failure documentation, and NotebookLM explanatory videos—has been produced using a Redmi 14C 5G Android phone and free-tier AI tools. Improved hardware would significantly increase my ability to conduct larger evaluations, manage documentation efficiently, and accelerate research output.
The funding will also support a pilot external validation by enabling two or three independent reviewers to reproduce the Layer 2 fact extraction process using only the published Layer 1 evidence and methodology documentation. The objective is to determine whether reviewers can independently produce substantially consistent fact extraction without additional guidance and to document any differences transparently.
Throughout the project, all methodology updates, specimens, and progress reports will continue to be published openly so that the wider community can inspect, critique, and build upon the work.
I am currently the sole researcher and maintainer of the CHRONOS project.
I independently designed the methodology, built the public GitHub repository, published the initial specimen library, documented the workflow, and continue to maintain the project through ongoing evaluations and public updates.
My professional background includes more than 30 years in corporate finance, internal auditing, compliance, inventory control, and quality systems. Those disciplines required maintaining accurate evidence records, separating factual observations from conclusions, and producing audit trails that could be independently reviewed. Those same principles form the foundation of the CHRONOS methodology.
Although CHRONOS is an early-stage independent research project and has not yet received external research funding, it already has a public development history, published evaluation specimens, documented AI failure analyses, and openly available methodology documentation. I also publish standalone LinkedIn posts documenting AI failure cases, including NotebookLM-generated explanatory videos that help broader audiences understand the documented evidence while preserving the original research artifacts.
My practical experience evaluating frontier AI systems comes directly from building and maintaining CHRONOS. The public repository, specimen library, and methodology documentation represent the evidence of that work. One objective of this funding is to move beyond a single-researcher project by inviting independent reviewers to evaluate whether the methodology is understandable, reproducible, and useful using only the publicly available documentation.
CHRONOS is an experimental research methodology, so the primary risk is not that the project stops, but that the methodology proves less useful or less reproducible than expected.
The most significant risk is that independent reviewers may find the workflow too complex, too time-consuming, or insufficiently clear to reproduce using only the published documentation. If that happens, the external validation phase will help identify those weaknesses, and the findings will be documented openly rather than treated as failures to hide.
Another risk is that limited computing resources and hardware may reduce the number of evaluation specimens I can complete within the planned timeframe. The funding is intended to reduce this constraint, but the project will continue even if progress is slower than planned.
Even if CHRONOS does not achieve broad adoption, all completed specimens, methodology documentation, and preserved evidence will remain publicly available as open research artifacts. Other researchers will still be able to inspect, critique, reuse, or build upon the work, ensuring that the project's outputs remain valuable regardless of the final outcome.
This is an independent project that I have self-funded. I have not received any external grants, research funding, or institutional financial support for CHRONOS during the last 12 months.
The public repository, methodology development, published evaluation specimens, AI failure documentation, and supporting educational materials have all been produced using my own resources. This funding would represent the first external support dedicated specifically to advancing the CHRONOS project.