You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Currently most of AI systems can give confident answers even when the starting assumption is wrong, and the situation is unclear, or even the information being used is irrelevant.
This issue I tested and then I have developed ScenariQ to test a different approach: instead of only checking the final answer, this approach controls and checks the conditions under which the AI reasons.
The goal with this project is to test and review whether this approach can actually reduce reasoning failures. I will compare the same problems with and without ScenariQ, including difficult cases involving wrong assumptions, conflicting information and unsupported conclusions.
My assumption is that ScenariQ is not proven yet. I want to test it properly, including trying to break it. If the results do not show meaningful improvement, or if the additional complexity creates too many false refusals or other problems, I will treat that as a valid research result.
My main goal is to test whether ScenariQ can reduce reasoning failures in AI systems.
I will try first test and review the same problems without ScenariQ and then test them with ScenariQ and compare the results.
I will include cases with wrong assumptions, conflicting information, irrelevant context, unsupported conclusions and situations where more than one scenario may apply.
I also want to test and review the limitations of the approach, including false refusals, complexity and latency.
After the controlled testing and review, I want to proceed towards more open and adversarial situations to see whether the results still hold.
The funding will mainly be used for my research time and for testing and reviewing ScenariQ properly. This will include building and validating the test benchmark, running the same tests with and without ScenariQ, testing different AI models, creating difficult and adversarial test cases, analysing failures and maintaining regression tests.
Some part of the funding will also be used for AI model/API costs, evaluation infrastructure, independent technical review and documentation of the research results.
A major part of this work is research and analysis rather than only software development. I will require time to design the tests, review the results, understand the reasons for failures and decide whether the modifications are actually meaningful.
I am currently the solo researcher on this project.
My professional background is as a Product Owner working on complex enterprise technology projects across insurance, manufacturing and automotive environments. My work has included requirements analysis, system and integration analysis, data migration, software testing, process design and working with complex enterprise systems.
ScenariQ grew from this practical experience. I have been developing, testing and reviewing the framework independently, including defining scenarios, identifying failure cases, developing the governance approach, testing the behaviour and reviewing and refining the framework based on the results.
I do not want to overstate my research background. This is an independent research effort that came out of my curiosity about how AI works, and the funding would help me take the work already done into a more rigorous and independently evaluated research program.
The most likely failure is that ScenariQ does not provide enough improvement over simpler approaches to justify the additional complexity.
Another possible failure is that the system becomes too restrictive and creates too many false refusals. I want to measure this because simply refusing difficult questions should not be considered a pass result.
There is also a risk that the approach works well in a controlled environment where scenarios are already known, but does not work as well in more open or adversarial situations.
The framework could also introduce additional latency, context overhead or maintenance effort that may not be justified by the overall improvement.
Also, it may not provide enough value when implemented at a larger scale.
$0. ScenariQ has not received any grant funding or investment in the last 12 months. The work so far has been independently developed and self-supported.
I have invested around 5 months of my own time to understand how AI systems work and to develop and test the framework. I have also used my own funds for AI subscriptions and other tools needed to explore and test the approach.
There are no bids on this project.