You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
DysonSETI is a project that trains on massive amounts of data from astronomical catalogs to create a general astronomical anomaly model. The pilot contains 14.1 million Gaia-WISE-2MASS sources for multimodal analysis; the project's first moonshot target is Dyson spheres and infrared excess anomaly sources. I will first flag and train this model on all available stellar source types, and then on anomalous stellar types like YSOs and T Tauri stars, etc. Then I will run the pipeline on large stellar catalogs to even find unknown-unknown anomaly clusters in datasets, for finding new types of sources or finding technosignatures. Even if the pipeline yields no verified candidates, it can still be used as a reusable anomaly finder in astronomy; I will publish model weights and so on. It will be Dysonian SETI's turboSETI tool at least. False-positive mitigation is the model's priority, as it is critical and a bottleneck in anomaly searches.
The project's goals are listed below:
1) Build a reusable multimodal astronomical anomaly model using Gaia, 2MASS, and WISE.
2) Develop several independent discovery methods. These include physical/statistical infrared-excess modeling, contextual ML, deep SED models, class-conditional anomaly detection, and OOD/open-world methods. Every discovery channel provides different advantages.
3) Detect both individual outliers and unknown anomaly populations.
4) Make false-positive mitigation part of the discovery system. Candidates will be tested against YSOs, evolved and dusty stars, binaries, variability, galaxies, AGN, blending, bad crossmatches, survey systematics, background contamination, and so on. A recent pilot version shows that a high fraction of detected anomalies are caused by background contamination or observational confusion rather than genuinely anomalous stellar types, so my priority will be to mitigate false positives in that channel.
5) Validate the models using calibration sets, synthetic Dyson-like infrared injections, and retraining. This is critical, as we can’t just rely on black-box models when trying to find anomalies.
6) Scale beyond the current 14.1-million-source pilot to much larger stellar catalogs (possibly 200-300 million sources with Gaia DR4). And I will continue to scale my model as open-source datasets evolve; for example, Gaia DR5, which is expected to be released before 2030, will be a genuinely substantial improvement compared to the Gaia DR4 DysonSETI model.
7) Release the code, model weights, calibration results, and anomaly products openly to support both citizen science and further research by other scientists.
8) Publish the results as two peer-reviewed papers: one methodology paper and one search paper within one year.
DysonSETI can usefully absorb funding from $10,000 to $54,000.
The first $10,000 would fund the current Gaia DR3 stage.
Approximate minimum compute budget:
400 L40S hours = $396
2,300 A100 hours = $3,197
800 H100 hours = $2,312
High-memory CPU preprocessing = $782
Storage = $992
Additional reruns and price margin = $2,321
Minimum compute request: $10,000
Funding above this would support the Gaia DR4-scale rebuild. My current estimate is approximately $30,000 in additional compute, mainly for large-scale H200/H100 training and storage.
This makes $40,000 the maximum useful compute budget.
Separately, I am requesting $14,000 for nine months of housing and basic living expenses in Ankara. I currently live in a noisy dormitory with limited infrastructure, including unreliable Wi-Fi, which makes research difficult.
Approximate nine-month budget:
Rent: $6,552
Deposit: $728
Utilities and internet: $969
Food: $2,700
Transportation and phone: $288
Moving/setup: $1,000
Inflation and emergency margin: $1,763
Research environment: $14,000
So:
Minimum funding: $10,000
Maximum compute funding: $40,000
Research environment: $14,000
Full funding goal: $54,000
Partial funding would still be immediately useful.
I am currently the sole investigator.
I am sixteen years old and conduct my research independently, without a university lab or research supervisor. I formulate the research questions, build the pipelines, analyze the data, write the manuscripts, and seek criticism directly from researchers in the field via cold-mailing.
I published a sole-authored SETI target-selection paper in the Publications of the Astronomical Society of the Pacific.
I also developed Stellar J-Harvesting, a proposed novel technosignature based on stellar angular-momentum extraction, conducted its first observational search using Kepler and Gaia data, and published my results on arXiv; the paper is currently under peer-review at Acta Astronautica. The search produced two initially interesting outliers, but I rejected them after false-positive analysis and reported a non-detection and requested follow-up observations.
I also have two sole-authored contributions accepted for presentation at the 2026 International Astronautical Congress.
DysonSETI is already underway. I have built and run the current 14.1-million-source Gaia, WISE, and 2MASS pilot. Now the project is waiting for funding to scale.
One possibility is that the ML models do not outperform simpler physical and statistical methods. Astronomical catalogs contain strong nuisance structure, mostly instrument artifacts and uncertainties. If the learned models mostly rediscover these effects, they may remain supporting tools. But I think we can mitigate this problem by making the system artifact-aware and conservative. Of course, we can't know without trying it.
A second possibility is that the pipeline works technically but finds no convincing technosignature candidates. I consider that very scientifically acceptable. The goal is not to manufacture candidates. If every interesting source eventually receives an astrophysical or observational explanation, that is preferable to preserving false positives. If the most advanced algorithm can't find candidates, then that's a strong conclusion.
A third possibility is that scaling is more computationally expensive than expected. I left a margin of error in the budget, but this is still a possibility. I will reduce this risk by profiling models on smaller data shards.
Even in these failure cases, the project can still produce a reusable astronomical anomaly model, rare-source or cluster discoveries, validation results, code, model weights, a strong upper limit on occurrence, and open anomaly catalogs. Everything will be released as open source. Of course, the upside is unrealistic but a real possibility: answering humanity's deepest question, “Are we alone?”
This project has not received any funding.