You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
mtg-kernel is an open research environment for studying whether reinforcement-learning agents learn strategies that transfer to unfamiliar opponents in Magic: The Gathering. Players must make decisions with hidden hands, uncertain draws, interacting rules, and consequences that may appear much later in a game.
The project already includes a deterministic Rust rules kernel, a training and evaluation stack, and an XMage bridge for independent comparison. I am seeking $10,000 for a twelve-week pilot that turns this infrastructure into a documented benchmark configuration and one reproducible study of self-play generalization. The proposed study uses a limited, explicitly supported card and deck scope. It does not depend on achieving expert or professional playing strength.
The research question is whether an improvement against training opponents also improves performance against opponents outside the training population. We will select a bounded comparison, freeze its analysis gates before measurement, and report both positive and negative findings.
For the funded comparison, we will change one declared training-population condition while holding the deck configuration, observation interface, learner architecture and compute budget fixed. Existing MTG learning projects include MTG-Causal-RL and MageZero. This project adds a controlled population-generalization study and reproducible artifacts; it does not claim to be the first MTG learning environment.
The deliverables are:
A documented benchmark configuration identifying the supported cards, decks, observations, and evaluation conditions.
Reproducible baseline and evaluation artifacts, with instructions and redistribution limited to material we can lawfully release.
A report comparing performance against training opponents and held-out opponents, including uncertainty, failures, exclusions, and the limits of the result.
Weeks 1-3 will establish the supported configuration, baseline artifacts, experimental comparison, and reproducibility check. Weeks 4-8 will complete the bounded comparison and held-out evaluation. Weeks 9-12 will prepare and publish the artifacts, instructions, and report.
Evaluation will use matched seeds and balanced player positions, predeclared integer win-count gates, and paired bootstrap comparisons of terminal game results. One same-seed rerun will check bit-identical output-store reproduction under a pinned runtime. The exact experimental arms and sample sizes will be finalized before measurement and within the available compute budget.
The $10,000 target supports:
Research, analysis, and release preparation: $5,000
CPU/GPU training and held-out evaluation: $4,000
Storage and reproducibility: $1,000
Total: $10,000
These are allocations rather than provider quotes. The final workload will be sized against actual resource estimates before paid measurement begins.
At the $5,000 minimum, I would complete a smaller pilot: document one supported configuration, package an existing baseline, and evaluate one prespecified comparison against a smaller held-out opponent set. The smaller budget would allocate $2,500 to research and release work, $2,000 to compute, and $500 to storage and reproducibility. It would omit a broad training search or expanded opponent sweep. Sample size and uncertainty would be reported, and the result would support only the narrower evaluated scope.
Between the minimum and target, the additional budget would expand the prespecified evaluation coverage after the scope is set and before outcomes are examined. I will not claim that a smaller funded pilot completed the full $10,000 plan.
I am also pursuing other funding and cloud-credit opportunities for this pilot. An overlapping award or credit will reduce the unmet request for the same costs. Every expense will be assigned to one funding source. If support exceeds the documented need, I will update the request and agree an appropriate reduction, return, or separately approved scope with the relevant funders.
I am Jack Maiorino, a part-time master's student in Applied Machine Learning at the University of Maryland, College Park. My work includes reinforcement learning for Magic: The Gathering and experiments on debate-style AI oversight. The latter work includes pilot re-analysis, an audit that identified oracle-channel bugs, and a validation harness. Those are research contributions, not evidence that the MTG project has achieved a particular playing strength.
I have worked on the independent MTG research project for approximately two years. It is separate from my course work and is not a UMD-sponsored thesis or capstone.
mtg-kernel already has a Rust rules engine, reinforcement-learning infrastructure, and an XMage comparison bridge. I can devote approximately 16-23 hours per week to the proposed twelve-week pilot alongside my studies. The exact experiment will be finalized before measurement.
Project and research artifacts: mtg-kernel and debate research code and reports. These links identify the work itself, not a funding endorsement.
The main scientific risk is that improved performance against training opponents will not transfer to the selected held-out opponents. That would be a useful negative result about the evaluated method, provided the comparison is valid and reported clearly.
The main engineering risk is that the chosen card and deck scope contains rules or evaluation defects that make some comparisons unreliable. We will document the supported scope, use the independent comparison path where applicable, and report failures and exclusions rather than treating invalid games as evidence of playing strength. A defect may require narrowing the configuration or the report's claims.
The main resource risk is that simulation or training costs limit the planned evaluation. The workload will be estimated before measurement, and the experiment will be sized to the available budget. A smaller study will be labeled as such. The project will still publish usable artifacts and a candid account of what was completed and what remains unresolved.
mtg-kernel has received no cash awards or formal support to date. I submitted a $10,000 application to Emergent Ventures and an application for Lambda cloud credits on September 6, 2026; no award decisions have been made. I am also pursuing AWS cloud credits, and none has been awarded through these applications.
A separate debate-style AI oversight project has a $10,000 award through Manifund. Those funds are restricted to that separate project and do not support mtg-kernel. This is a factual disclosure of other project funding, not an endorsement or credential for the present proposal.