You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Archive II/III started with a question I couldn't stop noticing in my own use of AI: if a system talks to the same person for years, does it actually get better at understanding that person, or does it just get better at telling a convincing story about them?
I have more than three years of conversations, corrections, research, arguments, creative work, technical projects, genealogy, decisions, mistakes, and generated material involving the same subject: me. I didn't create that record for an experiment. Most of it existed before I even realized there was an experiment hiding inside it.
That makes it different from giving a model a fictional profile and asking what it thinks. I can go back to the actual record and ask: Where did this belief come from? Was that ever verified? Did I correct it six months ago? Did the correction stick? Did two different models make the same unsupported leap? Is the model more accurate because it remembers more, or simply more confident?
Archive II/III turns that accumulated record into a structured experiment.
Different AI systems will be given comparable parts of the archive and asked to model the same person. I will compare what they get right, what they invent, where they disagree, what they refuse to conclude, and what happens when their interpretations are challenged by new or contradictory evidence.
I'm especially interested in the point where memory stops being simple recall and becomes a model of a human being. That matters more as AI systems become personalized, persistent, and capable of acting on our behalf.
I'm not trying to prove that an AI can "know" somebody. I'm trying to figure out how we can tell when it doesn't.
The first goal is simple: find out whether more memory actually means more accuracy.
I will build controlled versions of the existing archive so different models can be given roughly equivalent information. Then I can compare what changes as more longitudinal context is added.
The second goal is to keep evidence and interpretation from quietly becoming the same thing. If a model says I'm interested in something because I explicitly said so, that's one kind of claim. If it says I'm motivated by something because it inferred a pattern across 40 conversations, that's another. I want those treated differently.
The third goal is to test correction. If a model gets something wrong and I correct it, does that correction survive later conversations? Or does the old interpretation come back because it fits the model's larger story better?
The fourth goal is to get other people involved. One obvious weakness of this project is that I am both the researcher and the person being modeled. I know things about my own history an outside evaluator doesn't. I can also be biased about interpretations I don't like. So selected claims will be reviewed by independent evaluators against the underlying evidence instead of simply asking whether I agree with the model.
I also want to use prospective tests when possible. If a model claims it understands a pattern well enough to predict something, the prediction should be written down before the outcome happens. That is much harder to rationalize after the fact.
The final goal is to publish the failures along with the successes. If five models confidently agree on something and they're all wrong, I want that preserved. If a model makes an interpretation I reject and later evidence supports it, that stays too.
A result doesn't have to flatter the project to be useful.
I'm asking for a maximum of $50,000, with $10,000 as the minimum level where I can run a meaningful version of the experiment.
At $10,000, I would concentrate on the essentials: model/API access, a small group of independent reviewers, basic provenance and comparison tooling, and an initial published evaluation.
At $50,000, I can run the project the way I think it actually needs to be run: more model families, repeated trials instead of one-off prompts, outside evaluators, methodology help, privacy review, better research infrastructure, and enough time to do the analysis carefully instead of fitting it around everything else.
The full budget is approximately:
$15,000 — project lead stipend for experiment design, corpus preparation, testing, analysis, coordination, and documentation
$8,000 — model/API access and compute for repeated multi-model experiments
$7,000 — independent evaluators and adversarial review
$5,000 — methodology/statistical consultation
$3,500 — privacy, security, and redaction work
$3,000 — provenance and evaluation tooling
$2,000 — storage and experiment infrastructure
$1,500 — documentation and publication
$5,000 — contingency for unexpected compute, evaluation, or research costs
More funding doesn't just mean analyzing more conversations. It mainly means more chances for someone other than me to challenge the results.
Today 12:43 AM
Rapid Grant Application - Airtable.pdf
I read the uploaded form. Fill it like this:
What are you working on? (183/200 characters)
Archive II/III: a longitudinal human-AI evaluation testing whether years of memory and conversation produce an accurate, corrigible model of a person—or merely a convincing narrative.
Link to your work
Primary:
https://dorianmartinsmith.wixsite.com/the-archive-1
If it lets you add another link:
https://github.com/THECOMMONWEALTHOFBLACKAMERICA/The-Akashic-record
If we approve your grant, can we share details publicly?
Select:
Can share publicly with my name
Public URL
https://dorianmartinsmith.wixsite.com/the-archive-1
How are you connected to the BlueDot community?
Don’t claim a connection we don’t have. If the menu has something like Other, New to BlueDot, No prior connection, or Online, choose the closest one. If none of those appear, send me the dropdown choices and I’ll pick the exact one.
How much funding are you requesting?
$10,000
Paste:
I am seeking funding for Archive II/III, the experimental phase of a longitudinal human-AI research project built from several years of real interactions between one person and AI systems.
The core research question is whether an AI system with extensive long-term information about a person actually develops a more accurate model of that person, or whether additional memory can instead produce a highly coherent but incorrect narrative.
Archive II will give multiple AI models access to structured portions of the same longitudinal record and compare their independent conclusions. I will document where models agree, where they disagree, which claims are supported by the underlying record, which claims are inference, and how they respond when the subject challenges or corrects their interpretations.
Archive III extends this into adversarial and prospective evaluation. Independent human reviewers will examine selected model claims and the evidence behind them. Where feasible, predictions or forward-looking interpretations will be recorded before outcomes are known, rather than judged retrospectively. Corrections and model failures will remain part of the published record rather than being removed.
This is particularly relevant to AI safety as systems increasingly use persistent memory, personalization, and long-term context. A system can remember thousands of accurate facts while still developing an inaccurate higher-level model of a person. That creates risks around overconfidence, misclassification, resistant beliefs, confirmation loops, and personalized systems treating inference as established fact.
The project is not intended to prove that AI can fully “know” a person. It is intended to create an unusually deep case study for testing where long-term AI memory improves understanding, where it fails, and whether those failures remain corrigible.
I am requesting $10,000 for a focused evaluation phase:
$3,000 — multi-model API/compute, storage, and experiment infrastructure
$2,500 — independent human evaluators and adversarial review
$1,500 — provenance, comparison, and correction-tracking tooling
$1,250 — research methodology/statistical consultation
$1,000 — privacy, redaction, and data-security work
$750 — documentation and publication
The underlying archive and research program already exist; this funding would be used to test the claims more rigorously rather than simply continue accumulating material.
A useful outcome would include negative results. If years of additional context make models more confidently wrong rather than more accurate, that is exactly the kind of result I want the project to make visible.
This is a better BlueDot pitch than presenting it as funding for The Archive books. The books can remain outputs of the project, but the grant request itself is for AI-memory evaluation, corrigibility testing, adversarial review, and measurement.
Pasted markdown(20260821-061404).md
File
Pasted markdown(20260821-063108).md
File
Archive II/III is a longitudinal human-AI evaluation project testing whether years of accumulated conversations, memories, corrections, and artifacts enable AI systems to form increasingly accurate models of a real person—or simply increasingly coherent narratives that can still be wrong.
The project uses an existing multi-year record of real human-AI interaction rather than a synthetic persona or dataset created specifically for an experiment. Multiple AI systems will analyze equivalent portions of that record, make claims about the subject, and have those claims traced back to evidence, challenged, corrected, and compared across models.
The research focuses on persistent-memory failure modes that may become increasingly important as AI systems become personalized and agentic: unsupported inference becoming treated as fact, corrections being forgotten, multiple models converging on the same false interpretation, confidence increasing without accuracy, and downstream decisions inheriting an incorrect model of the user.
Archive II/III will publish both successes and failures. The goal is not to demonstrate that AI can “know” a person, but to develop practical methods for testing whether long-term AI models of people are accurate, auditable, corrigible, and appropriately uncertain.
The project has five primary goals.
1. Measure whether more memory actually improves person-model accuracy.
I will construct controlled subsets of the existing longitudinal archive and provide equivalent evidence to multiple AI models under different memory/context conditions.2. Separate evidence from inference.
Model claims will be linked back to the underlying record and classified according to whether they are directly supported, inferred, contradicted, uncertain, or unsupported.3. Test corrigibility and error persistence.
Models will receive explicit corrections and contradictory evidence, then be evaluated in later sessions to determine whether those corrections persist or whether the original mistaken interpretation returns.4. Test conclusions independently.
External evaluators will review selected claims against source evidence rather than relying only on my judgment as the subject. Where feasible, prospective predictions will be recorded before outcomes occur.5. Publish a reusable evaluation methodology.
The final work will document model disagreements, correction histories, calibration failures, successful predictions, false narratives, and negative findings. The aim is to produce methods that could later be tested on other longitudinal human-AI datasets.The project succeeds even if the central hypothesis fails. Discovering that additional memory makes models more confidently wrong—or that reliable person-modeling cannot be measured from this dataset—would itself be useful evidence.
Funding goal: $50,000. Minimum viable funding: $10,000.
At the full funding level, approximately:
$15,000 — Project lead stipend: experiment design, corpus preparation, analysis, coordination, and execution.
$8,000 — Multi-model API/compute: repeated controlled runs across model families, context conditions, correction tests, and adversarial experiments.
$7,000 — Independent human evaluators: external scoring of model claims against the underlying record.
$5,000 — Research methodology/statistical consultation: evaluation design, inter-rater reliability, scoring methods, and prospective testing.
$3,500 — Privacy/security/redaction: responsible handling of a sensitive longitudinal corpus.
$3,000 — Evaluation and provenance tooling: linking model claims to sources, corrections, versions, and evidence states.
$2,000 — Storage/experiment infrastructure
$1,500 — Documentation and publication
$5,000 — Contingency
With only $10,000, I would run a smaller evaluation using fewer models and evaluators while prioritizing API/compute, independent review, and publication of an initial study.
Additional funding primarily increases experimental replication and outside evaluation rather than simply expanding the amount of material analyzed.
The project is currently led by Dorian Martin-Smith, who is both the project designer and the primary longitudinal subject.
I have spent more than three years accumulating and preserving the underlying human-AI record through normal use rather than constructing it retrospectively for this study. The archive includes research, corrections, source material, creative projects, technical work, genealogy, decisions, disagreements with AI systems, and generated artifacts across time.
I completed the first Archive phase, which treats that accumulated interaction history as a research object, and developed Mapping the Living Subject, which formalizes questions around AI person-modeling, inference versus evidence, prediction, confidence, memory, and correction.
I have also begun cross-model analysis. Preliminary work has shown that different systems can receive overlapping evidence yet disagree significantly about what counts as fact, what should be inferred, and how confidently those interpretations should be expressed.
A related project, T.A.R. — The Akashic Records, explores source provenance, explicit evidence states, corrections, persistent memory, and auditable AI-assisted research:
https://github.com/THECOMMONWEALTHOFBLACKAMERICA/The-Akashic-record
The Archive:
https://dorianmartinsmith.wixsite.com/the-archive-1
I do not come from a conventional AI-safety laboratory or academic research career, and I do not want to obscure that. The next phase deliberately uses funding to add independent evaluators, methodological advisers, and adversarial review so the findings are not dependent solely on my interpretation.
The most likely failure is methodological rather than operational: the dataset may be unusually deep but still too specific to one subject to support conclusions that generalize beyond this case.
A second risk is subject-researcher bias. Because I know the underlying life history better than an evaluator does, disagreement between me and a model cannot automatically be treated as evidence that the model is wrong. This is why the funded design introduces source-based scoring, outside reviewers, explicit evidence classifications, and prospective tests where possible.
Another possible failure is that model differences prove highly sensitive to prompting, model versions, or context construction, making stable conclusions difficult to reproduce.
There is also a possibility that the central phenomenon is simply weaker than expected: more longitudinal information may straightforwardly improve model accuracy without producing distinctive new failure modes.
None of these outcomes would make the work worthless. If the single-subject methodology fails to generalize, the project can document its limits. If results are prompt-sensitive, that itself becomes an evaluation finding. If long-term context reliably improves accuracy, that is useful evidence too.
The worst outcome would be producing an impressive narrative and mistaking it for evidence. The project is deliberately designed so that a compelling story without measurable support counts as failure rather than success.
$0 has been awarded or received for Archive II/III during the last 12 months. The project has been independently developed and self-funded to date.
I currently have several pending applications submitted in August 2026, including:
BlueDot Impact Rapid Grant: $10,000 requested for a smaller Archive II/III evaluation phase — pending.
Lightcone Commons: Archive II/III funding proposal — pending.
EA Funds / Transformative AI Fund: $50,000 requested — pending.
I also have pending applications for separate projects, including T.A.R., CRFTE, and arts/public-history work. Those funds would not be treated as Archive II/III funding unless explicitly awarded for overlapping research.
If multiple Archive II/III applications are funded, I will disclose the awards and reduce, expand, or reallocate scopes so the same expense is not funded twice.
Funding
Minimum funding (USD)*svg
Max funding goal (USD)*svg
Decision deadline*svg
Logistics
Cause areas
svgScience & technology
svgTechnical AI safety
svgBiomedical
svgAI governance
svgAnimal welfare
svgBiosecurity
svgEA community
svgForecasting
svgGlobal catastrophic risks
svgGlobal health & development
What countries are you (and anyone else working on this) based in?*
This is for Manifund operations and will not be published.
This project will engage in lobbying.svg
Fill it like this:
Minimum funding (USD): 10,000
Max funding goal (USD): 50,000
Decision deadline: October 31, 2026
For Cause areas, select:
Technical AI safety
Science & technology
Global catastrophic risks
I would not select AI governance, forecasting, biosecurity, biomedical, or EA community—the project is fundamentally an empirical AI-memory/evaluation study.
What countries are you (and anyone else working on this) based in?
United States
This project will engage in lobbying:
No / leave unchecked
Archive II/III is research and evaluation, not an effort to influence legislation or political activity.
Pasted markdown(20260821-064624).md
File
IT SAYS WE RECCOMEND WRITING YOUR PROPOSAL IN YOUR OWN WORDS NOT AI THE HOME PAGE FLAG OUT CONTENT WRITTEN AS AI SO SPIVY THIS UP FOR ME:
Yeah — the current version is solid, but it reads a little too much like a grant template. I’d make it more personal, less symmetrical, and more willing to say what you actually noticed and why you care about it. No rewrite can guarantee what an automated detector will label, but this sounds much more like an individual explaining a project rather than a generated proposal.
Archive II/III started with a question I couldn't stop noticing in my own use of AI: if a system talks to the same person for years, does it actually get better at understanding that person, or does it just get better at telling a convincing story about them?
I have more than three years of conversations, corrections, research, arguments, creative work, technical projects, genealogy, decisions, mistakes, and generated material involving the same subject: me. I didn't create that record for an experiment. Most of it existed before I even realized there was an experiment hiding inside it.
That makes it different from giving a model a fictional profile and asking what it thinks. I can go back to the actual record and ask: Where did this belief come from? Was that ever verified? Did I correct it six months ago? Did the correction stick? Did two different models make the same unsupported leap? Is the model more accurate because it remembers more, or simply more confident?
Archive II/III turns that accumulated record into a structured experiment.
Different AI systems will be given comparable parts of the archive and asked to model the same person. I will compare what they get right, what they invent, where they disagree, what they refuse to conclude, and what happens when their interpretations are challenged by new or contradictory evidence.
I'm especially interested in the point where memory stops being simple recall and becomes a model of a human being. That matters more as AI systems become personalized, persistent, and capable of acting on our behalf.
I'm not trying to prove that an AI can "know" somebody. I'm trying to figure out how we can tell when it doesn't.
The first goal is simple: find out whether more memory actually means more accuracy.
I will build controlled versions of the existing archive so different models can be given roughly equivalent information. Then I can compare what changes as more longitudinal context is added.
The second goal is to keep evidence and interpretation from quietly becoming the same thing. If a model says I'm interested in something because I explicitly said so, that's one kind of claim. If it says I'm motivated by something because it inferred a pattern across 40 conversations, that's another. I want those treated differently.
The third goal is to test correction. If a model gets something wrong and I correct it, does that correction survive later conversations? Or does the old interpretation come back because it fits the model's larger story better?
The fourth goal is to get other people involved. One obvious weakness of this project is that I am both the researcher and the person being modeled. I know things about my own history an outside evaluator doesn't. I can also be biased about interpretations I don't like. So selected claims will be reviewed by independent evaluators against the underlying evidence instead of simply asking whether I agree with the model.
I also want to use prospective tests when possible. If a model claims it understands a pattern well enough to predict something, the prediction should be written down before the outcome happens. That is much harder to rationalize after the fact.
The final goal is to publish the failures along with the successes. If five models confidently agree on something and they're all wrong, I want that preserved. If a model makes an interpretation I reject and later evidence supports it, that stays too.
A result doesn't have to flatter the project to be useful.
I'm asking for a maximum of $50,000, with $10,000 as the minimum level where I can run a meaningful version of the experiment.
At $10,000, I would concentrate on the essentials: model/API access, a small group of independent reviewers, basic provenance and comparison tooling, and an initial published evaluation.
At $50,000, I can run the project the way I think it actually needs to be run: more model families, repeated trials instead of one-off prompts, outside evaluators, methodology help, privacy review, better research infrastructure, and enough time to do the analysis carefully instead of fitting it around everything else.
The full budget is approximately:
$15,000 — project lead stipend for experiment design, corpus preparation, testing, analysis, coordination, and documentation
$8,000 — model/API access and compute for repeated multi-model experiments
$7,000 — independent evaluators and adversarial review
$5,000 — methodology/statistical consultation
$3,500 — privacy, security, and redaction work
$3,000 — provenance and evaluation tooling
$2,000 — storage and experiment infrastructure
$1,500 — documentation and publication
$5,000 — contingency for unexpected compute, evaluation, or research costs
More funding doesn't just mean analyzing more conversations. It mainly means more chances for someone other than me to challenge the results.
Right now, this is primarily me: Dorian Martin-Smith.
I'm also the primary longitudinal subject, which is unusual, but that's the reason this particular dataset exists.
I didn't spend three years pretending to be a research subject for an experiment. I was actually using AI: researching my family history, writing, developing projects, changing my mind about things, correcting mistakes, arguing with models, building technical ideas, and living my life while this record accumulated around it.
Eventually I realized there was something worth studying in the record itself.
That became The Archive, the first phase of the project. From there I developed Mapping the Living Subject, which asks more directly what AI systems can legitimately infer about a person from a long-running digital record.
I've already started comparing how different systems interpret overlapping evidence. One of the things that pushed me toward making this more rigorous is seeing that models can know many of the same facts and still construct noticeably different versions of the same person. They differ not just in conclusions, but in what they treat as evidence and how confidently they fill in gaps.
I also built a related open-source project, T.A.R. — The Akashic Records, around provenance, evidence states, corrections, memory, and auditable AI-assisted research:
https://github.com/THECOMMONWEALTHOFBLACKAMERICA/The-Akashic-record
The Archive is here:
https://dorianmartinsmith.wixsite.com/the-archive-1
I don't have a conventional AI-safety résumé, and I'm not going to dress this up like it came out of a university lab. My advantage is that I have the longitudinal record, I know how it developed, and I've spent years working inside it.
My disadvantage is that I'm too close to the subject.
That's why a significant part of the funding goes toward bringing other people into the evaluation.
Today 12:43 AM
Rapid Grant Application - Airtable.pdf
I read the uploaded form. Fill it like this:
What are you working on? (183/200 characters)
Archive II/III: a longitudinal human-AI evaluation testing whether years of memory and conversation produce an accurate, corrigible model of a person—or merely a convincing narrative.
Link to your work
Primary:
https://dorianmartinsmith.wixsite.com/the-archive-1
If it lets you add another link:
https://github.com/THECOMMONWEALTHOFBLACKAMERICA/The-Akashic-record
If we approve your grant, can we share details publicly?
Select:
Can share publicly with my name
Public URL
https://dorianmartinsmith.wixsite.com/the-archive-1
How are you connected to the BlueDot community?
Don’t claim a connection we don’t have. If the menu has something like Other, New to BlueDot, No prior connection, or Online, choose the closest one. If none of those appear, send me the dropdown choices and I’ll pick the exact one.
How much funding are you requesting?
$10,000
Paste:
I am seeking funding for Archive II/III, the experimental phase of a longitudinal human-AI research project built from several years of real interactions between one person and AI systems.
The core research question is whether an AI system with extensive long-term information about a person actually develops a more accurate model of that person, or whether additional memory can instead produce a highly coherent but incorrect narrative.
Archive II will give multiple AI models access to structured portions of the same longitudinal record and compare their independent conclusions. I will document where models agree, where they disagree, which claims are supported by the underlying record, which claims are inference, and how they respond when the subject challenges or corrects their interpretations.
Archive III extends this into adversarial and prospective evaluation. Independent human reviewers will examine selected model claims and the evidence behind them. Where feasible, predictions or forward-looking interpretations will be recorded before outcomes are known, rather than judged retrospectively. Corrections and model failures will remain part of the published record rather than being removed.
This is particularly relevant to AI safety as systems increasingly use persistent memory, personalization, and long-term context. A system can remember thousands of accurate facts while still developing an inaccurate higher-level model of a person. That creates risks around overconfidence, misclassification, resistant beliefs, confirmation loops, and personalized systems treating inference as established fact.
The project is not intended to prove that AI can fully “know” a person. It is intended to create an unusually deep case study for testing where long-term AI memory improves understanding, where it fails, and whether those failures remain corrigible.
I am requesting $10,000 for a focused evaluation phase:
$3,000 — multi-model API/compute, storage, and experiment infrastructure
$2,500 — independent human evaluators and adversarial review
$1,500 — provenance, comparison, and correction-tracking tooling
$1,250 — research methodology/statistical consultation
$1,000 — privacy, redaction, and data-security work
$750 — documentation and publication
The underlying archive and research program already exist; this funding would be used to test the claims more rigorously rather than simply continue accumulating material.
A useful outcome would include negative results. If years of additional context make models more confidently wrong rather than more accurate, that is exactly the kind of result I want the project to make visible.
This is a better BlueDot pitch than presenting it as funding for The Archive books. The books can remain outputs of the project, but the grant request itself is for AI-memory evaluation, corrigibility testing, adversarial review, and measurement.
Pasted markdown(20260821-061404).md
File
Pasted markdown(20260821-063108).md
File
Archive II/III is a longitudinal human-AI evaluation project testing whether years of accumulated conversations, memories, corrections, and artifacts enable AI systems to form increasingly accurate models of a real person—or simply increasingly coherent narratives that can still be wrong.
The project uses an existing multi-year record of real human-AI interaction rather than a synthetic persona or dataset created specifically for an experiment. Multiple AI systems will analyze equivalent portions of that record, make claims about the subject, and have those claims traced back to evidence, challenged, corrected, and compared across models.
The research focuses on persistent-memory failure modes that may become increasingly important as AI systems become personalized and agentic: unsupported inference becoming treated as fact, corrections being forgotten, multiple models converging on the same false interpretation, confidence increasing without accuracy, and downstream decisions inheriting an incorrect model of the user.
Archive II/III will publish both successes and failures. The goal is not to demonstrate that AI can “know” a person, but to develop practical methods for testing whether long-term AI models of people are accurate, auditable, corrigible, and appropriately uncertain.
The project has five primary goals.
1. Measure whether more memory actually improves person-model accuracy.
I will construct controlled subsets of the existing longitudinal archive and provide equivalent evidence to multiple AI models under different memory/context conditions.2. Separate evidence from inference.
Model claims will be linked back to the underlying record and classified according to whether they are directly supported, inferred, contradicted, uncertain, or unsupported.3. Test corrigibility and error persistence.
Models will receive explicit corrections and contradictory evidence, then be evaluated in later sessions to determine whether those corrections persist or whether the original mistaken interpretation returns.4. Test conclusions independently.
External evaluators will review selected claims against source evidence rather than relying only on my judgment as the subject. Where feasible, prospective predictions will be recorded before outcomes occur.5. Publish a reusable evaluation methodology.
The final work will document model disagreements, correction histories, calibration failures, successful predictions, false narratives, and negative findings. The aim is to produce methods that could later be tested on other longitudinal human-AI datasets.The project succeeds even if the central hypothesis fails. Discovering that additional memory makes models more confidently wrong—or that reliable person-modeling cannot be measured from this dataset—would itself be useful evidence.
Funding goal: $50,000. Minimum viable funding: $10,000.
At the full funding level, approximately:
$15,000 — Project lead stipend: experiment design, corpus preparation, analysis, coordination, and execution.
$8,000 — Multi-model API/compute: repeated controlled runs across model families, context conditions, correction tests, and adversarial experiments.
$7,000 — Independent human evaluators: external scoring of model claims against the underlying record.
$5,000 — Research methodology/statistical consultation: evaluation design, inter-rater reliability, scoring methods, and prospective testing.
$3,500 — Privacy/security/redaction: responsible handling of a sensitive longitudinal corpus.
$3,000 — Evaluation and provenance tooling: linking model claims to sources, corrections, versions, and evidence states.
$2,000 — Storage/experiment infrastructure
$1,500 — Documentation and publication
$5,000 — Contingency
With only $10,000, I would run a smaller evaluation using fewer models and evaluators while prioritizing API/compute, independent review, and publication of an initial study.
Additional funding primarily increases experimental replication and outside evaluation rather than simply expanding the amount of material analyzed.
The project is currently led by Dorian Martin-Smith, who is both the project designer and the primary longitudinal subject.
I have spent more than three years accumulating and preserving the underlying human-AI record through normal use rather than constructing it retrospectively for this study. The archive includes research, corrections, source material, creative projects, technical work, genealogy, decisions, disagreements with AI systems, and generated artifacts across time.
I completed the first Archive phase, which treats that accumulated interaction history as a research object, and developed Mapping the Living Subject, which formalizes questions around AI person-modeling, inference versus evidence, prediction, confidence, memory, and correction.
I have also begun cross-model analysis. Preliminary work has shown that different systems can receive overlapping evidence yet disagree significantly about what counts as fact, what should be inferred, and how confidently those interpretations should be expressed.
A related project, T.A.R. — The Akashic Records, explores source provenance, explicit evidence states, corrections, persistent memory, and auditable AI-assisted research:
https://github.com/THECOMMONWEALTHOFBLACKAMERICA/The-Akashic-record
The Archive:
https://dorianmartinsmith.wixsite.com/the-archive-1
I do not come from a conventional AI-safety laboratory or academic research career, and I do not want to obscure that. The next phase deliberately uses funding to add independent evaluators, methodological advisers, and adversarial review so the findings are not dependent solely on my interpretation.
The most likely failure is methodological rather than operational: the dataset may be unusually deep but still too specific to one subject to support conclusions that generalize beyond this case.
A second risk is subject-researcher bias. Because I know the underlying life history better than an evaluator does, disagreement between me and a model cannot automatically be treated as evidence that the model is wrong. This is why the funded design introduces source-based scoring, outside reviewers, explicit evidence classifications, and prospective tests where possible.
Another possible failure is that model differences prove highly sensitive to prompting, model versions, or context construction, making stable conclusions difficult to reproduce.
There is also a possibility that the central phenomenon is simply weaker than expected: more longitudinal information may straightforwardly improve model accuracy without producing distinctive new failure modes.
None of these outcomes would make the work worthless. If the single-subject methodology fails to generalize, the project can document its limits. If results are prompt-sensitive, that itself becomes an evaluation finding. If long-term context reliably improves accuracy, that is useful evidence too.
The worst outcome would be producing an impressive narrative and mistaking it for evidence. The project is deliberately designed so that a compelling story without measurable support counts as failure rather than success.
$0 has been awarded or received for Archive II/III during the last 12 months. The project has been independently developed and self-funded to date.
I currently have several pending applications submitted in August 2026, including:
BlueDot Impact Rapid Grant: $10,000 requested for a smaller Archive II/III evaluation phase — pending.
Lightcone Commons: Archive II/III funding proposal — pending.
EA Funds / Transformative AI Fund: $50,000 requested — pending.
I also have pending applications for separate projects, including T.A.R., CRFTE, and arts/public-history work. Those funds would not be treated as Archive II/III funding unless explicitly awarded for overlapping research.
If multiple Archive II/III applications are funded, I will disclose the awards and reduce, expand, or reallocate scopes so the same expense is not funded twice.
Funding
Minimum funding (USD)*svg
Max funding goal (USD)*svg
Decision deadline*svg
Logistics
Cause areas
svgScience & technology
svgTechnical AI safety
svgBiomedical
svgAI governance
svgAnimal welfare
svgBiosecurity
svgEA community
svgForecasting
svgGlobal catastrophic risks
svgGlobal health & development
What countries are you (and anyone else working on this) based in?*
This is for Manifund operations and will not be published.
This project will engage in lobbying.svg
Fill it like this:
Minimum funding (USD): 10,000
Max funding goal (USD): 50,000
Decision deadline: October 31, 2026
For Cause areas, select:
Technical AI safety
Science & technology
Global catastrophic risks
I would not select AI governance, forecasting, biosecurity, biomedical, or EA community—the project is fundamentally an empirical AI-memory/evaluation study.
What countries are you (and anyone else working on this) based in?
United States
This project will engage in lobbying:
No / leave unchecked
Archive II/III is research and evaluation, not an effort to influence legislation or political activity.
Pasted markdown(20260821-064624).md
File
IT SAYS WE RECCOMEND WRITING YOUR PROPOSAL IN YOUR OWN WORDS NOT AI THE HOME PAGE FLAG OUT CONTENT WRITTEN AS AI SO SPIVY THIS UP FOR ME:
Yeah — the current version is solid, but it reads a little too much like a grant template. I’d make it more personal, less symmetrical, and more willing to say what you actually noticed and why you care about it. No rewrite can guarantee what an automated detector will label, but this sounds much more like an individual explaining a project rather than a generated proposal.
Archive II/III started with a question I couldn't stop noticing in my own use of AI: if a system talks to the same person for years, does it actually get better at understanding that person, or does it just get better at telling a convincing story about them?
I have more than three years of conversations, corrections, research, arguments, creative work, technical projects, genealogy, decisions, mistakes, and generated material involving the same subject: me. I didn't create that record for an experiment. Most of it existed before I even realized there was an experiment hiding inside it.
That makes it different from giving a model a fictional profile and asking what it thinks. I can go back to the actual record and ask: Where did this belief come from? Was that ever verified? Did I correct it six months ago? Did the correction stick? Did two different models make the same unsupported leap? Is the model more accurate because it remembers more, or simply more confident?
Archive II/III turns that accumulated record into a structured experiment.
Different AI systems will be given comparable parts of the archive and asked to model the same person. I will compare what they get right, what they invent, where they disagree, what they refuse to conclude, and what happens when their interpretations are challenged by new or contradictory evidence.
I'm especially interested in the point where memory stops being simple recall and becomes a model of a human being. That matters more as AI systems become personalized, persistent, and capable of acting on our behalf.
I'm not trying to prove that an AI can "know" somebody. I'm trying to figure out how we can tell when it doesn't.
The first goal is simple: find out whether more memory actually means more accuracy.
I will build controlled versions of the existing archive so different models can be given roughly equivalent information. Then I can compare what changes as more longitudinal context is added.
The second goal is to keep evidence and interpretation from quietly becoming the same thing. If a model says I'm interested in something because I explicitly said so, that's one kind of claim. If it says I'm motivated by something because it inferred a pattern across 40 conversations, that's another. I want those treated differently.
The third goal is to test correction. If a model gets something wrong and I correct it, does that correction survive later conversations? Or does the old interpretation come back because it fits the model's larger story better?
The fourth goal is to get other people involved. One obvious weakness of this project is that I am both the researcher and the person being modeled. I know things about my own history an outside evaluator doesn't. I can also be biased about interpretations I don't like. So selected claims will be reviewed by independent evaluators against the underlying evidence instead of simply asking whether I agree with the model.
I also want to use prospective tests when possible. If a model claims it understands a pattern well enough to predict something, the prediction should be written down before the outcome happens. That is much harder to rationalize after the fact.
The final goal is to publish the failures along with the successes. If five models confidently agree on something and they're all wrong, I want that preserved. If a model makes an interpretation I reject and later evidence supports it, that stays too.
A result doesn't have to flatter the project to be useful.
I'm asking for a maximum of $50,000, with $10,000 as the minimum level where I can run a meaningful version of the experiment.
At $10,000, I would concentrate on the essentials: model/API access, a small group of independent reviewers, basic provenance and comparison tooling, and an initial published evaluation.
At $50,000, I can run the project the way I think it actually needs to be run: more model families, repeated trials instead of one-off prompts, outside evaluators, methodology help, privacy review, better research infrastructure, and enough time to do the analysis carefully instead of fitting it around everything else.
The full budget is approximately:
$15,000 — project lead stipend for experiment design, corpus preparation, testing, analysis, coordination, and documentation
$8,000 — model/API access and compute for repeated multi-model experiments
$7,000 — independent evaluators and adversarial review
$5,000 — methodology/statistical consultation
$3,500 — privacy, security, and redaction work
$3,000 — provenance and evaluation tooling
$2,000 — storage and experiment infrastructure
$1,500 — documentation and publication
$5,000 — contingency for unexpected compute, evaluation, or research costs
More funding doesn't just mean analyzing more conversations. It mainly means more chances for someone other than me to challenge the results.
Right now, this is primarily me: Dorian Martin-Smith.
I'm also the primary longitudinal subject, which is unusual, but that's the reason this particular dataset exists.
I didn't spend three years pretending to be a research subject for an experiment. I was actually using AI: researching my family history, writing, developing projects, changing my mind about things, correcting mistakes, arguing with models, building technical ideas, and living my life while this record accumulated around it.
Eventually I realized there was something worth studying in the record itself.
That became The Archive, the first phase of the project. From there I developed Mapping the Living Subject, which asks more directly what AI systems can legitimately infer about a person from a long-running digital record.
I've already started comparing how different systems interpret overlapping evidence. One of the things that pushed me toward making this more rigorous is seeing that models can know many of the same facts and still construct noticeably different versions of the same person. They differ not just in conclusions, but in what they treat as evidence and how confidently they fill in gaps.
I also built a related open-source project, T.A.R. — The Akashic Records, around provenance, evidence states, corrections, memory, and auditable AI-assisted research:
https://github.com/THECOMMONWEALTHOFBLACKAMERICA/The-Akashic-record
The Archive is here:
https://dorianmartinsmith.wixsite.com/the-archive-1
I don't have a conventional AI-safety résumé, and I'm not going to dress this up like it came out of a university lab. My advantage is that I have the longitudinal record, I know how it developed, and I've spent years working inside it.
My disadvantage is that I'm too close to the subject.
That's why a significant part of the funding goes toward bringing other people into the evaluation.
The biggest risk is probably that this ends up being a fascinating case study of one person but not something that generalizes very far beyond me.
That's possible.
Another problem is bias. If a model says something about me that I think is wrong, my disagreement doesn't automatically prove anything. I'm the subject, but I'm also capable of misunderstanding myself, remembering something incorrectly, or simply disliking an interpretation.
That's why I don't want "Dorian agrees with it" to be the scoring method.
There's also a reproducibility problem. Model behavior can change based on wording, context order, model version, temperature, or what information happened to be included. I may discover that some apparently deep conclusions disappear when the same question is phrased differently.
That would be a result, not something I'd hide.
It's also possible that the dramatic failure mode I'm looking for just isn't very strong. Maybe more memory usually does make models more accurate. Maybe corrections work better than I expect. Maybe there isn't a major difference between long-term person-modeling and ordinary factual memory.
That would be useful to know too.
The outcome I most want to avoid is producing something that sounds profound but can't survive examination.
A model writing a beautiful psychological portrait of someone is not evidence that the portrait is true.
If Archive II/III fails because it exposes how hard that distinction is to measure, I still think that failure would teach us something important.
$0 has been awarded or received for Archive II/III in the last 12 months.
I've built the project independently and funded the work myself so far.
I currently have several applications pending that relate directly to Archive II/III: $10,000 requested from the BlueDot Impact Rapid Grant, an application submitted to Lightcone Commons, and $50,000 requested from the EA Funds Transformative AI Fund.
None of those applications has resulted in an award at the time I'm submitting this proposal.
I also have applications pending for other projects, including T.A.R., CRFTE, and separate arts/public-history work. I don't count those as money raised for Archive II/III.
If more than one funder supports Archive II/III, I will disclose the awards and adjust the scope so nobody is paying twice for the same expense. More funding would let me expand the number of models, reviewers, prospective tests, and replication attempts rather than duplicate an already-funded line item.