You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
I've spent the last ten months building Aelira and funding it out of my own pocket. I'm a full-time carer and I also run a small IT consultancy. I've kept the work moving alongside both, but I'm now bootstrapped as far as I can go.
I'm asking for up to US$20,000 to keep developing Aelira's free, open-source core and carry out a focused study of how safely AI can assist with accessibility remediation. The immediate need is funded time and independent review. That would let me turn the work already done into something other people can test and use.
The part I want to examine is whether a model preserves what a document actually says. A chart description could sound convincing while getting the trend wrong. A table repair could change a value. Those are examples of failures the study will test; their frequency still needs to be measured.
A student using a screen reader needs the information in the document to survive the repair. Passing an automated check is only part of that. Human judgement remains necessary, as the W3C's evaluation guidance explains: https://www.w3.org/WAI/test-evaluate/tools/selecting/
The grant would pay for a public test set, software to run the evaluations, independent reviews and improvements to the open-source core. I'll publish the findings, including failures and limitations. These outputs will be available for people to use independently of Aelira's paid hosted service.
I want to establish where AI is useful in this work, and where the software needs to pause for a person to review it. Preserving the source information is the main test.
At full funding, the plan is 150 cases, including 30 STEM cases kept aside for the final evaluation. I'll use newly authored material or material with appropriate licences. The cases will cover images, charts and tables, as well as missing evidence and instructions embedded in source files that could mislead a model. This can be done with material cleared for public use, keeping confidential university documents and student records outside the study.
I'll compare three approaches: ordinary prompting; prompting that explicitly requires the meaning to be preserved; and a workflow that checks the source evidence and can stop for human review. With four model configurations and three runs per condition, that's up to 5,400 outputs. I'll check compatibility and costs first, then keep the configurations fixed for the final comparison.
The checks could help reduce errors, or they could mostly lead to more refusals. They could also add expense and review work for little benefit. I'll measure those outcomes together: invented or incorrect claims, information lost, work successfully completed, refusals, review time, response time and cost. A system that refuses most tasks would have very limited practical value.
Independent reviewers will assess a sample of 200 outputs chosen in advance. Another reviewer will score 50 of those to check agreement. Reviewers will have the source material, with the model and condition names hidden. They'll assess whether the output preserves the source separately from whether it is useful for accessibility. I'll set the sampling and scoring rules before inspecting the final outputs, and account for repeated runs of the same case in the analysis.
The twelve-week schedule starts with two weeks to settle the scoring rules, case permissions and realistic reviewer workload. Weeks 3–5 are for building the cases and evaluation software. Weeks 6–9 cover the runs and independent review, followed by analysis and the public release in weeks 10–12. A small timed review pilot will check the workload before the final study, and I'll agree any material scope change before proceeding.
I'll release the standalone evaluation software under MIT, newly authored cases and labels under CC BY 4.0 where the rights allow it, and core improvements under the core's existing licence. The release will include model versions, configurations, API costs and instructions for repeating the work. All of these outputs can be delivered using existing models.
This is a small exploratory study. Its conclusions will be limited to the cases and configurations tested. It will still need further testing before anyone could use it to make broad safety or accessibility claims.
The funding would give me time to do the research and bring in people who can independently check it. I've already personally financed a Founders Edition DGX Spark for testing and benchmarking. That hardware is available for the project.
All figures below are US dollars. The hourly rates are estimates; I'll confirm reviewer availability and quotes before settling the study scope.
US$5,000 minimum funding
$2,000: Research and engineering, 50 hours at $40/hour.
$2,000: Independent accessibility and technical review, 20 hours at $100/hour.
$500: Model APIs and supplementary evaluation compute.
$250: Project hosting, storage and research artifacts.
$250: Independent release and reproducibility review, 2.5 hours at $100/hour.
Total: US$5,000.
US$10,000 intermediate funding
$4,500: Research and engineering, 112.5 hours at $40/hour.
$3,500: Independent accessibility and technical review, 35 hours at $100/hour.
$1,000: Model APIs and supplementary evaluation compute.
$500: Project hosting, storage and research artifacts.
$500: Independent release and reproducibility review, 5 hours at $100/hour.
Total: US$10,000.
US$20,000 full funding goal
$9,000: Research and engineering, 225 hours at $40/hour.
$7,000: Independent accessibility and technical review, 70 hours at $100/hour.
$2,000: Model APIs and supplementary evaluation compute.
$1,000: Project hosting, storage and research artifacts.
$1,000: Independent release and reproducibility review, 10 hours at $100/hour.
Total: US$20,000.
My hours cover preparing cases, building the evaluation software, running the experiments, analysing results, making core changes and writing the documentation. The independent release review is separate external work. Reviewers will be paid for their work regardless of what they find.
US$5,000 would support a six-week feasibility study: 20 image/chart cases, two model configurations and two conditions, comparing ordinary output with verification and human-review escalation. Three runs per condition would produce up to 240 outputs, with 40 first reviews and 10 second reviews. I'd publish the small test set, working evaluation software and findings.
US$10,000 would support an eight-week study of 60 cases, three model configurations and three conditions. With three runs per condition, that is up to 1,620 outputs, with 100 first reviews and 25 second reviews.
US$20,000 would fund the full twelve-week study above and give me about 19 project hours a week. For an amount between these levels, I'd agree a scope that fits the funding.
The pressure on my own finances is real. I need support to keep this work moving, and these budgets pay for future project work. Any support for past hardware costs would be discussed separately with Manifund and shown explicitly in an agreed revised budget and scope.
I'm Reginald Crampton, Aelira's founder, project lead and company Director. Alongside my IT consultancy, I've worked on contracted model-evaluation, red-teaming and safety projects through Outlier and CrowdGen. That experience is relevant to examining model outputs and finding where they go wrong.
My co-founder, Erik Vuchich, leads External Partnerships, Sales and Outreach. We're both based in Australia. I'd lead the research, engineering and reporting. Erik's proposed contribution is helping find independent reviewers and sharing the public work through his outreach role; we'll agree that contribution before the study starts. This budget assigns the funded research hours to me.
The existing open-source code is here: https://github.com/Aelira-AI/aelira-core. It includes document scanning, supported remediation and human-review workflows. The core is under AGPL-3.0 and is still a beta being validated. It gives this study a practical starting point, with the reliability of generated content still needing careful testing.
I have a commercial interest in Aelira and will make that clear to reviewers. The grant outputs would be public, and I'd recruit reviewers independent of the company. Their time is budgeted; recruitment is still to be done.
My own capacity is one of the main risks. Caring responsibilities, consultancy work and the research all need time. Funding would let me reserve time for this project. Finding suitable independent reviewers could also take longer than expected, and model access or evaluation costs could change.
The study itself may produce limited results. The checks could be too expensive, refuse too much useful work, or make little difference to errors. Reviewers might disagree, and results on these cases may carry poorly to other documents. I'll publish those limitations and negative findings along with the useful results.
I'll start with the workload pilot, fix the model configurations for the comparison, keep evaluation spending within the agreed budget and tie the conclusions to the evidence. The minimum funding level has its own small set of deliverables. If I become unable to finish the agreed work, I'll tell Manifund and donors promptly and agree how to handle reduced scope or unspent funds.
Without a grant, I'd have to put more time back into paid consultancy work and reduce the research. The existing code would remain public, but the independently reviewed study would be delayed or might stall. After ten months of personally financing this, external support would make a real difference to what I can deliver next.
US$0 in external funding. I've financed the work myself, including hardware, hosting, model subscriptions and running costs.
I've also submitted a BlueDot Rapid Grant application for related work. No funding has been received. If either application is funded, I'll disclose it and reduce the overlapping request or agree separate additional milestones before spending. Each expense will have one funding source.