You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I want to test a small but consequential assumption in AI training: if a cached tool response has the right distribution for each individual call, is reusing it safe for the learning process? My September 2026 preprint gives a counterexample in a two-action model. It does not show that an actual language model was trained incorrectly. This proposal funds the next step: a controlled study of repeated learning and an open audit harness that makes the distinction testable.
I request $12,000 for 12 weeks of part-time independent research. A $6,000, six-week pilot would produce a useful standalone result. The output will be public code, a preregistered evaluation protocol, complete results and a short working paper, including negative results.
What are this project's goals? How will you achieve them?
The question is whether shared stochastic tool outcomes change the behavior learned by a policy when it makes several decisions, and whether a cheaper execution policy can retain acceptable behavior under an equal physical-call budget.
My existing artifact verifies 540 finite-sum configurations and 3,240 estimator evaluations. A separate audit runs 256 scripted rollouts through a pinned TVCache implementation. It establishes an execution path for sharing; it does not refute TVCache's deterministic-output contract or demonstrate an LLM training regression. Code and paper: https://github.com/shi1720/tool-cache-coupling
Weeks 1-2: specify three small sequential decision environments with stochastic tools, including a deterministic-output control. Freeze the environments, reward definitions, seed list and stopping rules before the main runs. Separate development cases from held-out cases.
Weeks 3-6: compare independent execution, full within-group sharing and partial sharing. Use both group-normalized and centered-only estimators. Run 20 seeds for each environment/execution/estimator combination, initially with small policies that fit a modest compute budget. Log cache identity, physical draws, rollout count and update statistics. Evaluate each learned policy in a fresh independent environment. Report paired differences in held-out return, variation across seeds and physical-call cost, rather than treating a single favorable run as evidence.
Weeks 7-10: repeat the comparison with equal physical-call budgets, investigate discrepancies, and package the instrumentation as a reproducible audit harness. A reviewer who did not implement the experiments will be budgeted to check the analysis and reproduce selected results. No reviewer is committed yet.
Weeks 11-12: publish the full artifact and a concise report stating which claims the evidence supports. I will include failure cases and a checklist for identifying when reuse changes the joint distribution consumed by an estimator. I will not claim production prevalence, general safety guarantees or dollar savings from toy experiments.
The pilot ends after week 6 with the protocol, initial results and executable tests. The full project adds equal-budget comparisons, independent review and a more reusable release. Work would start within four weeks of an award after confirming a feasible schedule.
How will this funding be used?
Full request, $12,000: $10,000 for 200 hours of my research time at $50/hour; $1,000 for compute and storage; $1,000 for independent methods/reproducibility review. Pilot, $6,000: $5,000 for 100 hours of research; $500 compute/storage; $500 review. There are no travel costs or institutional overhead. I will use simple policies first, so the project does not depend on expensive frontier-model training. Review is a planned purchase, not an existing partnership.
Who is on your team? What's your track record on similar projects?
I am Shivam Gupta, an independent researcher and applied AI engineer based in Dubai. I hold a Bachelor of Technology in Computer Science and Design from IIIT Delhi. I work across AI product engineering and evaluation, including creating more than 500 expert-reviewed coding tasks, reproducible RL environments and reference implementations at micro1. I also build RepoGym, a framework for verifiable coding-agent environments. These experiences help me turn a narrow research claim into something another engineer can rerun.
I am the sole applicant. My four recent research manuscripts are preprints, not peer-reviewed publications. I use AI tools for coding, analysis and drafting, including assistance preparing this application; I remain responsible for checking claims and outputs. Portfolio: https://shivamgupta.web.app . Code: https://github.com/shi1720 .
What are the most likely causes and outcomes if this project fails?
The effect could vanish with repeated optimization, be limited to contrived environments, or be too small to matter at equal cost. That would still be useful if the protocol and negative results are reproducible. Implementation mistakes are another risk; deterministic controls, a separately checked baseline and external review are intended to catch them. A small-policy result will not establish behavior in frontier models. If the pilot finds no robust effect, I will report that outcome and reassess further spending with the funder rather than expand the claim.
How much money have you raised in the last 12 months, and from where?
I have received no research grants or fellowships. This request has no committed funding. I have pending applications to several funders, including the AI Safety Research Fund, GTR, Mercor and Lightcone Commons, for related research. These are alternative or potentially overlapping sources, not awards. Before accepting overlapping support I will disclose it, reduce or withdraw the overlapping request, and agree distinct deliverables and budgets where appropriate. I will not charge the same hours or expenses to two funders. This is a part-time project alongside my engineering work; employer/client material will not be used without permission.