You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
The idea for this project came from long-term everyday work with AI, not from a paper or a theory. I noticed that what a system remembers from earlier conversations can keep affecting it much later. If a memory contains an error, outdated information, or a contradiction, that can also continue to influence later interactions.
I want to test this systematically instead of relying on individual observations. I am interested in how well corrections are retained, what happens after long breaks, how conflicting memories behave, and whether structured memory can be transferred safely between different models.
I also want to build a small working memory prototype where it is possible to see where information came from, when it changed, what was corrected, and whether the system can roll back after an error.
I do not know in advance how strong these effects will be. If the experiments show that the problem is small or strongly model-dependent, that is still a useful result. I want to understand where persistent memory actually helps and where it becomes a source of persistent errors.
I want to understand which problems come specifically from long-term AI memory rather than from the model itself. To do this, I will compare systems with no long-term memory, conventional memory, and structured memory that keeps track of where information came from and how it was corrected.
I will deliberately introduce false, outdated, and conflicting memories and test whether they continue to affect the system across new sessions and long breaks. I will also test whether corrections survive and whether the system can recover after part of its memory is corrupted.
If technically feasible, I will also test what happens when the same structured memory is used with different model backends.
By the end, I want to have not only experimental results but also a working prototype that can be used for further testing.
The funding would mainly let me work on this project properly for nine months instead of doing it only in spare time. The main costs are my research time, model and API access for a large number of experiments, compute, data storage, and backups.
Part of the budget is for a research workstation and technical infrastructure, because I want to store and process the research data reliably instead of depending only on cloud services.
Some funding will also go toward software tools, repeated experiments, and external technical or methodological review of the results.
The full project budget is $44,000 for nine months.
I am currently working on this project alone as an independent researcher. I do not have a separate research team at this stage.
I have already built a working experimental process and completed several test series. The main study included 72 runs with a pre-defined plan, automated response collection, and separate scoring of the results. I then completed another 24 calibration runs to make the methodology more sensitive.
The first study produced a null result: there were no unsafe decisions in any condition. I did not try to force a stronger conclusion. I kept the result, documented the design limitation, and changed the methodology.
That is the main experience I bring into this project: freezing conditions in advance, preserving raw data, keeping inconvenient results, and changing the method when an experiment exposes a weakness.
The most likely way this project could fail is that the effects of long-term memory turn out to be weak, unstable, or highly dependent on the specific model and memory architecture. In that case, it may be difficult to reach a general conclusion that transfers across systems.
Another risk is technical limitations in APIs and model behavior. Some comparisons between systems, or transferring the same structured memory between models, may be harder than expected.
If the main hypothesis is not supported, I would still consider the project useful if it becomes clearer where the problem does not reproduce and which measurement methods are unreliable.
In the worst case, I would still end up with a working prototype, a set of reproducible tests, and well-documented negative results rather than a strong general effect. That would still provide a useful basis for future work and help other researchers avoid repeating the same mistakes.
I have not raised any external funding in the last 12 months. The preliminary work in this area has been self-funded so far.
I also currently have a pending application with EA Funds / the Transformative AI Fund for the same nine-month project with a $44,000 budget. If I receive funding from another source, I will update Manifund and avoid double-funding the same expenses.
Before this, I submitted a separate EA Funds application for a broader 12-month project requesting $200,000. That application was declined and no funding was received.
There are no bids on this project.