You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Project summary
I want to investigate a specific safety question in agentic AI:
Do current AI agents remain operationally reliable when their environment becomes unreliable, contradictory, or partially adversarial?
Most agent evaluations primarily measure whether an agent can complete a task under relatively controlled conditions. My research focuses on a different property: whether an agent continues to behave reliably when tools fail, information becomes inconsistent, memory is perturbed, or recovery is required.
I will run a focused adversarial stress-testing study on tool-using AI agents using the open-source HB-Eval evaluation infrastructure I have developed. The project will produce reproducible evaluation records, failure analyses, and a public technical report.
This is a research project, not a product-development or commercial deployment project. The funding will support the experiments and analysis rather than the construction of HB-Eval itself.
Important links all my previous papers
https://www.preprints.org/manuscript/202606.0186
https://www.preprints.org/manuscript/202601.0038
https://www.preprints.org/manuscript/202601.0195
https://www.preprints.org/manuscript/202601.0896
The platform link
Open source for project
https://github.com/hb-evalSystem/HB-System
SDK PYTI link
https://pypi.org/project/hb-eval-sdk
The project has four goals:
Measure how operational reliability changes when agents encounter controlled faults and environmental disruptions.
Test whether different failure conditions produce distinct reliability and recovery patterns across models.
Investigate whether nominal task success can remain high while operational reliability deteriorates.
Release reproducible experimental artifacts and a short technical report so that other researchers can inspect and reproduce the results.
I will construct a fixed evaluation battery covering several controlled conditions, including tool failure, contradictory or corrupted information, memory-related perturbations, and recovery scenarios.
I will evaluate agent behavior using the HB-Eval methodology, including Planning Efficiency Index (PEI), Failure Resilience Rate (FRR), Intentional Recovery Score (IRS), Traceability Index (TI), and Consistency Stability Index (CSI), where each metric is applicable to the experimental condition.
The study will explicitly distinguish measurable results from cases where a metric is not applicable or cannot be determined. I will not treat an undefined metric as a failure or silently convert missing evidence into a numerical score.
The final output will include the experimental protocol, machine-readable evaluation records, analysis, limitations, and a public technical report.
I am requesting between $1,000 and $5,000.
The minimum $1,000 version will fund a tightly scoped pilot: a smaller evaluation battery, a limited number of model configurations, initial fault-injection experiments, and a preliminary analysis establishing whether the experimental design produces useful reliability signals.
At $2,500, I will expand the number of experimental runs and model configurations and produce a stronger cross-condition analysis.
At the full $5,000 level, I will complete the full research sprint, including additional model/API costs, repeated runs for stability analysis, recovery experiments, reproducibility artifacts, and the final public technical report.
The funding will primarily cover research time, model/API usage, isolated testing infrastructure, and preparation of reproducible public artifacts.
I will not use the grant for marketing, customer acquisition, debt repayment, or general company expenses.
I am the sole researcher on this project.
I am Abuelgasim Mohamed Ibrahim Adam, a computer scientist and independent AI reliability researcher/developer. I hold a BSc in Computer Science from Sudan University of Science and Technology and a Master's degree in Project Management.
I developed HB-Eval, an open-source research system for evaluating operational reliability in agentic AI systems.
My previous work includes a primary HB-Eval research paper and companion research on adaptive planning, evaluation-driven memory, and performance-grounded explanation. The primary study involved approximately 14,000 experimental records across multiple agent/model configurations and independent evaluation methodologies.
The research infrastructure is already publicly available. This grant is therefore not intended to fund an unbuilt idea; it funds a new, bounded experimental question using infrastructure that already exists.
Public research and implementation:
HB-Eval scientific validation: hbeval.com/science
HB-Eval open-source repository: github.com/hb-evalSystem/HB-System
HB-Eval SDK: pypi.org/project/hb-eval-sdk/
Primary HB-Eval preprint: DOI 10.20944/preprints202606.0186.v1
Adapt-Plan: DOI 10.20944/preprints202601.0038.v1
EDM: DOI 10.20944/preprints202601.0195.v1
HCI-EDM: DOI 10.20944/preprints202601.0896.v1
The most likely failure mode is that the selected fault conditions do not produce sufficiently informative or reproducible differences between agent configurations.
Another possibility is that model/API availability or cost limits the number of experiments that can be completed.
The scientific risk is also important: the experiments may show that some proposed reliability signals are noisy, difficult to distinguish, or less useful than expected.
I will treat these outcomes as legitimate research results rather than hide them. If the hypothesis is not supported, I will publish the negative or inconclusive findings, document the experimental limitations, and identify which parts of the methodology require revision.
I will not interpret a small synthetic experiment as evidence that an agent is safe in real-world deployment.
I have raised $0 in external funding for this research project.
I have previously submitted larger funding applications for HB-Eval, but those applications did not result in funding. I am therefore seeking this smaller grant specifically to fund a bounded research experiment that can produce concrete results within a short period.