Monday, August 03, 2026

Beyond Correlation: How Historians Distinguish Causation, Coincidence, and Contingency in the Age of AI


by Alvin Blackshear | Historian & Researcher  <ablackshear@gmail.com>

There is a particular kind of intellectual vertigo that comes from watching a machine narrate the past with total confidence. Ask an AI system to explain why a revolution happened, why a movement gained momentum, or why a community's history unfolded as it did, and it will almost always answer. The prose will be fluent, the timeline will be clean, and the causal claims will arrive stated as fact rather than as argument. What the answer will rarely do is pause to ask whether the events it has linked together were actually connected at all, or whether they simply happened to occur near one another in time. This is the quiet crisis facing historical writing in an era of generative AI: the technology is exceptionally good at producing narratives that sound causal, and considerably less good at knowing whether those narratives are true.

Historians have spent more than a century developing tools to guard against exactly this failure. Long before anyone worried about large language models hallucinating a tidy story out of scattered facts, historians were worried about themselves doing the same thing. The discipline's entire methodological apparatus, footnotes, source criticism, peer review, historiographical debate, exists in large part to slow down the impulse to declare that because one thing followed another, the first caused the second. What AI-generated history reintroduces, at scale and with new fluency, is an old temptation: the temptation to mistake sequence for cause, and pattern for proof.

The Seduction of Sequence

Correlation is easy to notice and hard to interpret responsibly. Two trends move together, two events happen in the same decade, two figures cross paths, and the mind supplies a connection because connection is a more satisfying story than coincidence. Sarah Woodhouse's recent work on correlation reminds us that not all statistical relationships imply causation, and that there are several distinct varieties of correlation, only one of which reflects a genuine causal link between the variables involved.[1] The others might reflect a shared underlying cause, a coincidence of timing, a feedback loop running in the opposite direction from what intuition suggests, or simply noise that resembles a pattern if you squint at it long enough.

Historians face this same interpretive challenge, except their raw material is not a spreadsheet but the messy residue of human lives: letters, court records, newspapers, oral testimony, material culture. When two developments appear together in this record, a rise in labor unrest alongside a shift in immigration policy, a change in religious practice alongside an economic downturn, a legal victory alongside a change in public rhetoric, the historian's task is to ask what actually connects them, if anything does. Johannes Nagel's comparative analysis of causal reasoning in historical scholarship makes a useful observation here: historians rarely operate with the clean methodological categories that philosophers of science might prefer. Real historical practice is often described as vaguely causal-analytical, blending causal-intentional explanation, attention to unintended consequences, and selective use of theory and comparison, without ever fully resolving into a single rigorous method.[2] This is not a failure of the discipline. It is an honest reflection of how difficult the underlying problem actually is.

AI systems, by contrast, do not experience this difficulty as friction. A model trained to produce plausible continuations of text will happily supply a causal mechanism because causal mechanisms make for better sentences than admissions of uncertainty. The danger is not that AI invents facts out of nothing, though it sometimes does that too. The deeper danger is that it takes real facts, arranges them in a narrative shape, and lets the shape itself imply a causal claim that no historian has actually verified.

Mechanism, Not Just Chronology

If there is a single principle that separates serious historical causal argument from mere narrative sequencing, it is the requirement of an explicit mechanism. It is not enough to say that event A preceded event B and therefore contributed to it. A historian is expected to explain how A produced B: through what institutions, what decisions, what material pressures, what beliefs held by which people.

Daniel Little's work on the microfoundations of social explanation is instructive here. Little argues that sweeping historical forces such as capitalism, the state, or industrialization do not themselves act as causal agents in any direct sense. They are abstractions that describe the aggregate result of countless individual decisions, made by people operating within specific institutions and constrained by specific incentives and beliefs. A causal historical explanation, on this view, has to eventually cash out in terms of those individuals: what they knew, what they wanted, what options were actually available to them, and how their choices interacted with those of others to produce a larger outcome. This is demanding work. It requires the historian to move between scales, from the granular texture of a single decision to the broader pattern that many such decisions eventually create, without losing the thread that connects the two.[3]

Mark Hewitson's argument about the marginalization of causal thinking within the historical profession adds an important wrinkle to this picture. Hewitson suggests that the turn toward language, discourse, and representation that reshaped historical writing over the past several decades, whatever its genuine intellectual gains, also made many historians reluctant to speak plainly about why things happened. Causal explanation came to seem naive, or worse, complicit in older positivist ambitions to discover fixed laws of history. And yet, as Hewitson points out, historians never actually stopped making causal arguments. Every account of a revolution, an election, a migration, or a social movement is saturated with implicit claims about why it occurred. What was lost was not the practice of causal reasoning but the willingness to examine that practice openly and hold it to a clear standard.[4]

This matters enormously for evaluating AI-generated history, because a system that has absorbed decades of scholarship in which causal claims are made implicitly, without explicit justification, will likely reproduce that same implicitness. It will offer causal-sounding narratives without ever surfacing the mechanism that supposedly connects cause to effect, because the training material it draws from often does the same thing. The corrective, for both human historians and the tools they now use, is the same: state the mechanism, or admit that you cannot.

What Coincidence and Contingency Demand of Us

Not everything that happens in history happens for a deep structural reason. Some things happen because a particular person was in a particular place at a particular moment, because a letter arrived a day late, because an illness struck one leader and spared another. Historians call this contingency, and distinguishing it from genuine causation is one of the discipline's more delicate tasks.

The stakes of this distinction are not merely academic. When causal narratives are built where only coincidence existed, the resulting history can flatten the genuine unpredictability and agency that shaped real events into something that looks inevitable, as though it could only have happened the way it did. This has particular consequences for histories that have too often been treated as footnotes to a larger, supposedly more central narrative. African American history, feminist history, and the history of gay and queer communities have all, at various points, been subjected to accounts that either erase contingency entirely, presenting outcomes as foregone conclusions of impersonal social forces, or that erase causation entirely, treating hard-won structural change as a matter of isolated coincidence rather than sustained organizing and argument. Getting the distinction right, separating what was truly contingent from what was the product of deliberate, structural, causal effort, is not a technical nicety. It is a matter of according historical actors the accuracy they deserve, whether that means recognizing the genuine unpredictability they navigated or recognizing the genuine causal power of what they built.

Tay Jeong's work on counterfactual reasoning offers historians one of their sharpest tools for making this distinction. The basic move is to ask what would have happened in the absence of the proposed cause. If removing a particular event or decision from the historical record would plausibly have left the outcome largely unchanged, the causal claim weakens considerably. If removing it would plausibly have produced a genuinely different outcome, the causal claim gains support. Jeong's more technical contribution is to note that counterfactual reasoning is most useful not as a blanket method applied indiscriminately, but in specific situations, particularly where two or more candidate causes seem to be doing similar explanatory work and our untrained intuitions about which one truly mattered are likely to mislead us. This is precise, careful reasoning, and it is exactly the kind of reasoning that a fluent AI narrative can obscure simply by never raising the question in the first place.[5]

The Discipline of Alternative Explanations

Good historical argument does not simply assert a cause and move on. It considers what else might explain the same outcome, and it explains why the preferred causal account is more persuasive than the alternatives. This habit of mind, developed across the humanities, has an unexpected ally in fields that seem far removed from history. Hannah Correia and her coauthors, writing about causal inference in environmental research, describe systematic approaches for distinguishing genuine causal relationships from mere observational association, approaches built around identifying confounding variables and establishing plausible mechanisms before accepting a causal claim.[6] Linbo Wang's comparative overview of causal inference frameworks, spanning the traditions associated with Rubin, Pearl, and structural equation modeling, makes a similar point from a more statistical angle: rigorous causal reasoning requires explicit assumptions, and different frameworks make different assumptions visible in different ways.[7] Though these authors are not writing about history, their underlying discipline, of refusing to accept a causal story until alternatives have been seriously weighed, translates directly into the historian's craft.

This is precisely the discipline that AI-generated narratives tend to skip. A well-trained model can produce a fluent account of why a particular reform succeeded, but it rarely volunteers the alternative explanations a careful historian would consider and reject: that the reform succeeded for reasons largely unrelated to the actors typically credited, that its apparent success was partly an artifact of how success was later measured, or that multiple contributing factors were entangled in ways that resist a single tidy story.

Evidence, Uncertainty, and the Historian's Honesty

Underneath all of this sits the question of evidence. Every causal claim a historian makes rests on some body of evidence, and that evidence varies enormously in quality, completeness, and reliability. Xinyue Chen and colleagues, in their study of how historians actually use visualization in their published work, found that historians face persistent practical and epistemological barriers when trying to represent uncertainty and justify their conclusions to readers and peers.[8] The justification burden, as they describe it, is not a peripheral concern. It is central to what makes a piece of historical writing trustworthy.

Georg Iggers, reflecting late in his career on the historian's role as an engaged intellectual, insisted that this burden cannot be set aside even by historians who write from a place of moral or political conviction. Iggers rejected the idea that historians must choose between detached objectivity and partisan advocacy.[9] Instead, he argued for a middle position: historians can and should be engaged with the pressing concerns of their time, but that engagement must never come at the cost of rigorous evidentiary standards. A historian writing feminist history, or African American history, or the history of gay and queer life, is no less bound by these standards than any other scholar, and indeed the stakes of getting the causal story right, of neither overstating structural inevitability nor understating hard-won agency, are often higher precisely because these histories have so frequently been distorted or dismissed.

Frank Ankersmit pushes this even further by reminding us that the past itself no longer exists in any directly accessible form. Historians do not reproduce it; they represent it, using an inherited semantic apparatus of meaning, truth, and reference that Ankersmit argues must be handled with care rather than assumed to work automatically.[10] If historical writing is, at some level, an act of representation rather than simple description, then AI-generated historical narrative inherits this same burden. It too is representing the past, not retrieving it, and representation always involves choices about emphasis, framing, and inclusion that deserve scrutiny rather than passive acceptance.

A Standard Worth Holding

What emerges from this body of scholarship is not a single formula but a disposition, a way of approaching any historical claim, human-authored or machine-generated, with a consistent set of questions. Is there a stated mechanism connecting cause to effect, or only a sequence dressed up as causation? Have alternative explanations been seriously considered and set aside for identifiable reasons? Has the line between structural causation and genuine contingency been drawn carefully, rather than assumed? Would the outcome have looked meaningfully different if the proposed cause were absent? And does the account acknowledge the limits of its own evidence, rather than presenting a confident narrative where the underlying record is thin or contested?

These questions predate artificial intelligence by decades, and they will outlast any particular model or platform. What AI has changed is the volume and fluency with which causal-sounding narratives can now be produced, and therefore the urgency of applying these questions consistently. The historians whose work informs this discussion were not writing with AI in mind. They were writing because the problem of causation in history has always been genuinely hard, and because getting it wrong has always carried real costs, especially for the histories of communities whose stories have too often been told either as inevitable or as accidental, when the truth usually lies in the harder, more interesting space between. That space, between correlation and causation, between coincidence and contingency, between chronology and mechanism, is where careful historical work has always lived. It is also, increasingly, where the responsible use of AI in historical research will need to live too.

---

Notes:

[1] Sarah Woodhouse, "When Two Things Seem Linked (but Aren't): Understanding Different Types of Correlations," Clearer Thinking, July 2, 2025, https://www.clearerthinking.org/post/when-two-things-seem-linked-but-aren-t-understanding-different-types-of-correlations.

[2] Johannes Nagel, "Causal Explanation in History: An Analysis of Research Strategies," Analyse & Kritik, March 12, 2026, https://doi.org/10.1007/s11577-026-01059-8.

[3] Daniel Little, Microfoundations, Method, and Causation: On the Philosophy of the Social Sciences (New Brunswick, NJ: Transaction Publishers, 1998).

[4] Mark Hewitson, History and Causality (Basingstoke: Palgrave Macmillan, 2014).

[5] Tay Jeong, "When to Use Counterfactuals in Causal Historiography," Sage Journals, February 2, 2025, https://doi.org/10.1177/00491241251314039.

[6] Hannah E. Correia et al., "Best Practices for Moving from Correlation to Causation in Environmental Research," Nature Communications, February 24, 2026, https://rdcu.be/fxl0y.

[7] Linbo Wang, "Causal Inference: A Tale of Three Frameworks," Journal of Data Science, February 11, 2026, https://doi.org/10.6339/25-JDS1211.

[8] Xinyue Chen et al., "How Historians Use Visualization: A Corpus-Backed Taxonomy and Analysis for Cross-Disciplinary Practice," July 10, 2026, https://doi.org/10.48550/arXiv.2605.01456.

[9] Georg Iggers, "The Historian as an Engaged Intellectual: Historical Writing and Social Criticism – A Personal Retrospective," in The Engaged Historian: Perspectives on the Intersections of Politics, Activism and the Historical Profession, ed. Stefan Berger (New York: Berghahn Books, 2019), 277–299.

[10] Frank Ankersmit, Historical Representation (Stanford, CA: Stanford University Press, 2002).

---

Works Cited

Ankersmit, Frank. Historical Representation. Stanford, CA: Stanford University Press, 2002.

Chen, Xinyue, et al. "How Historians Use Visualization: A Corpus-Backed Taxonomy and Analysis for Cross-Disciplinary Practice." July 10, 2026. https://doi.org/10.48550/arXiv.2605.01456.

Correia, Hannah E., et al. "Best Practices for Moving from Correlation to Causation in Environmental Research." Nature Communications, February 24, 2026. https://rdcu.be/fxl0y.

Hewitson, Mark. History and Causality. Basingstoke: Palgrave Macmillan, 2014.

Iggers, Georg. "The Historian as an Engaged Intellectual: Historical Writing and Social Criticism – A Personal Retrospective." In The Engaged Historian: Perspectives on the Intersections of Politics, Activism and the Historical Profession, edited by Stefan Berger, 277–299. New York: Berghahn Books, 2019.

International Commission for the History and Theory of Historiography (ICHTH). 2026 Congress Proceedings. https://www.ichth.net/content/home.html.

Jeong, Tay. "When to Use Counterfactuals in Causal Historiography." Sage Journals, February 2, 2025. https://doi.org/10.1177/00491241251314039.

Little, Daniel. Microfoundations, Method, and Causation: On the Philosophy of the Social Sciences. New Brunswick, NJ: Transaction Publishers, 1998.

Nagel, Johannes. "Causal Explanation in History: An Analysis of Research Strategies." Analyse & Kritik, March 12, 2026. https://doi.org/10.1007/s11577-026-01059-8.

Wang, Linbo. "Causal Inference: A Tale of Three Frameworks." Journal of Data Science, February 11, 2026. https://doi.org/10.6339/25-JDS1211.

Woodhouse, Sarah. "When Two Things Seem Linked (but Aren't): Understanding Different Types of Correlations." Clearer Thinking, July 2, 2025. https://www.clearerthinking.org/post/when-two-things-seem-linked-but-aren-t-understanding-different-types-of-correlations.

Sunday, August 02, 2026

AI for Older Adults


AI for Older Adults

by Alvin Blackshear | Historian & Researcher

Artificial intelligence tools like ChatGPT, Claude, Gemini, Perplexity, and Copilot have moved from novelty to necessity faster than almost any technology in memory. For adults over 60, that speed can feel less like an invitation and more like being left behind at the station. But research on how older adults actually use, and struggle with, these tools tells a more encouraging story: the gap isn't about ability, it's about approach. A 2026 study of older adults' technology support requests found that when queries were rewritten to add missing context and clearer language, the accuracy of AI-generated solutions jumped from 46% to 69%, and related Google search results improved from 35% to 69% accuracy. In other words, the tools already work well enough. The skill worth learning is how to talk to them.

This article walks through four practical steps: choosing the right AI for your needs, writing prompts that get better answers, understanding what AI "agents" are and when to trust them, and avoiding the mistakes that trip up new users of every age, but especially those who didn't grow up parsing tech jargon.

Choosing the Right AI

There is no single "best" AI assistant. Each is tuned for different strengths, and picking the right one for the task at hand saves time and frustration. A simple way to think about it:

·   Best for everyday questions: ChatGPT is the most general-purpose assistant, well suited to answering quick questions, drafting emails, or explaining a topic in conversation.

·   Best for writing and long documents: Claude tends to produce clear, well-organized long-form writing and is a strong choice for letters, essays, or reviewing lengthy documents.

·   Best for web research: Perplexity is built around search, pairing its answers with cited sources so you can click through and verify what it tells you, which is useful when accuracy matters most.

·   Best for Google users: Gemini integrates directly with Gmail, Google Docs, and Google Search, which makes it a natural fit if that's already your everyday ecosystem.

·   Best for Microsoft 365 users: Copilot is built into Word, Outlook, and Excel, so if your family or workplace runs on Microsoft software, it can work inside the programs you already know.

None of these choices is permanent or exclusive. Many people end up using two: one for quick daily questions and another for anything they want double-checked with sources. The point isn't to find the "correct" one, but to match the tool to the task, the same way you'd choose a phone call over a text message depending on what you're trying to accomplish.

Learning to Write Better Prompts

The single biggest lever an older adult has for getting better answers from AI isn't a special app or setting. It's the wording of the question itself. Research on communication between older adults and AI systems identified four recurring patterns that reduce answer quality: verbosity (explaining too much surrounding detail), incompleteness (leaving out a key fact the AI needs), over-specification (adding assumptions that aren't actually true), and under-specification (being too vague about what's wanted). The good news from that same research is that these patterns are fixable, not through retraining how you think, but through a few reliable habits when you type or speak your question.

Ask for plain English. AI models default to whatever register they think fits the question, and that can drift into jargon. Simply adding "explain this in plain English, no technical terms" or "explain it like you would to a smart friend who isn't a computer person" changes the entire tone of the reply.

Request step-by-step explanations. Instead of asking "how do I attach a photo to an email," try "walk me through this one step at a time, and wait for me to say 'next' before moving to the next step." This breaks a wall of instructions into a manageable back-and-forth, and it also gives you a natural point to ask a follow-up question without losing the thread.

Ask for examples. Abstract instructions are harder to follow than concrete ones. Adding "give me a real example" or "show me what that would actually look like on my screen" turns a vague description into something you can match against what you're actually seeing.

Ask the AI to critique its own answer. This is a lesser-known but powerful technique: after receiving an answer, ask "is there anything wrong with this answer, or a simpler way to do it?" AI models are often able to catch their own errors or oversimplifications when prompted to double-check themselves, and this habit builds in a layer of quality control you don't have to supply yourself.

The underlying finding from the research is worth repeating: rephrased, well-specified questions were understood far better by others (93.7% versus 65.8% comprehension in one comparison) and led to solutions people felt confident following. A better-written question isn't a nicety. It's the difference between a correct answer and a confusing one.

Using AI Agents

You may have started to hear the term "AI agent," distinct from a regular chatbot. It's worth understanding the difference, because agents come with real convenience and real risk.

What an agent is. A standard AI chat answers a question and stops. An agent is designed to take a goal and carry out a sequence of actions on its own (searching the web, filling out a form, booking something, or working through a multi-step task) before reporting back. Industry analysis on the state of AI agents in 2026 notes that the defining shift in newer systems is their ability to handle progressively longer and more complex tasks, maintaining context across many steps rather than answering a single quick question.

When to use one. Agents are useful for tasks that are repetitive, well-defined, and low-stakes if something goes slightly wrong, for example, summarizing a folder of documents, comparing prices across several websites, or drafting a set of similar emails. They save the most time exactly where a task has many small, mechanical steps that would otherwise take you an hour of clicking around.

Risks of autonomous agents. The tradeoff for that convenience is reduced oversight. Because an agent acts across many steps without checking in at each one, a small misunderstanding early on can compound into a larger error by the end, the AI equivalent of a wrong turn early in a road trip that takes you far off course before anyone notices. This matters most for anything involving money, personal information, medical decisions, or sending communications on your behalf. Never give an agent standing access to your bank accounts, passwords, or ability to make purchases without a manual confirmation step at the end.

Practical everyday examples. Reasonable, lower-risk uses include: asking an agent to research and summarize options for a specific purchase (with you making the final decision), having it organize digital photos or files, or using it to draft, but not send, a batch of similar messages, such as holiday cards or appointment reminders, for you to review individually before anything goes out.

Avoiding Common Mistakes

Even with the right tool and better prompts, a handful of habits cause most of the frustration people report with AI, regardless of age.

Trusting AI too quickly. AI models can sound confident even when they're wrong. They occasionally generate plausible-sounding but incorrect information, a phenomenon often called "hallucination." Treat a confident tone as no guarantee of accuracy, especially for anything involving health, legal, or financial specifics.

Asking vague questions. As the communication research above showed, vague or incomplete questions are the most common source of poor answers. If a response feels off-target, the fastest fix usually isn't switching tools. It's adding the missing context and asking again.

Accepting the first answer. The first response an AI gives is a starting point, not a verdict. Asking a follow-up like "are you sure?" or "what would make this answer wrong?" often surfaces caveats or corrections the first answer left out.

Failing to verify important facts. For anything that matters, such as a medication interaction, a legal deadline, or a financial figure, cross-check the AI's answer against a trusted human source, such as your doctor, pharmacist, financial advisor, or a reputable website. Tools like Perplexity, which show their sources directly, make this step easier by letting you click through and confirm the claim yourself rather than taking it on faith.

The Bigger Picture

None of this requires becoming a technical expert. The research is fairly consistent on this point: older adults who use AI report high confidence and ease once they receive well-matched, clearly explained answers. The 2026 diary study found 89.8% of older adult participants felt able to answer AI's follow-up questions, and 94.7% felt able to follow the solutions they were given. The barrier isn't capability; it's a mismatch in how questions get asked and answered. Broader research on aging and generative AI echoes the same conclusion, pointing to real promise in health and cognitive support as long as tools are built and used thoughtfully, with attention to privacy and genuine ease of use. Choose a tool that matches the task, ask questions the way outlined above, use agents cautiously and only for low-stakes chores, and verify anything that truly matters. That combination, not fluency in tech jargon, is what determines whether AI becomes a genuinely useful assistant in daily life.

___________________________ 

Sources:

Dhanuka, N. (2026, February 26). State of AI agents 2026: Autonomy is here. Prosus. https://www.prosus.com/news-insights/2026/state-of-ai-agents-2026-autonomy-is-here

Ernst & Young Global Limited. (2026, March). Understanding older generations' adoption of AI [Ripples Research Report]. https://www.ey.com/content/dam/ey-unified-site/ey-com/en-gl/about-us/corporate-responsibility/documents/ey-gl-how-older-generations-are-engaging-with-ai-03-2026.pdf

Liu, T. (2026, May 29). The role of multimodal generative AI in older adults' health management: Systematic scoping review. JMIR AI. https://ai.jmir.org/2026/1/e84695

Shomee, H. H. (2026, January 15). Empowering older adults in digital technology use with foundation models. arXiv. https://doi.org/10.48550/arXiv.2601.10018

 

Saturday, August 01, 2026

Ten Common Hallucinations in Genealogical AI


Ten Common Hallucinations in Genealogical AI

by Alvin Blackshear | Historian & Researcher <ablackshear@gmail.com>

If you've started using AI tools to help build out your family tree, you've probably noticed something exciting: AI can search patterns, suggest connections, and summarize records faster than any human researcher. But you may have also noticed something unsettling. Sometimes the AI is simply wrong, and it's wrong with total confidence. It doesn't hedge. It doesn't say "I'm not sure." It states fabricated facts as if they were pulled straight from a courthouse archive.

This isn't a quirk specific to genealogy tools. It's a well-documented feature of how large language models work. Researchers have found that AI systems are, in a sense, rewarded for guessing rather than admitting uncertainty, much like a student on a hard exam who writes down a plausible-sounding answer rather than leaving the question blank. The same dynamic that produces fabricated legal citations in court filings and nonexistent references in academic papers also produces invented ancestors, false marriages, and imaginary historical explanations in your family tree.

For beginning genealogists who are comfortable with AI tools but new to the discipline of genealogical research, this creates a real risk. Family history has its own standards of evidence (rules developed over generations of professional practice), and AI, for all its fluency, doesn't inherently know or follow them. Below are ten (plus a bonus eleventh) hallucinations you're likely to encounter, along with how to spot each one.

#1. Invented Ancestors

The most obvious and probably most common hallucination is the flat-out invented person. If your records show John Smith born in 1846, married to Sarah Jones, and died in 1912, an AI tool may cheerfully add "William Smith, father of John" to your tree, simply because a father-shaped gap exists and the model has learned that fathers usually appear in these slots.

No birth record names him. No probate record mentions him. No document anywhere establishes his existence. He is, in the most literal sense, a statistical guess dressed up as a fact.

How to catch it: Ask the AI directly, "What specific document supports this person's existence?" If it can't name one, or names one that turns out not to exist (see #5), strike the individual from your tree.

#2. False Parentage

This is a subtler and more dangerous error because it doesn't invent anyone; it just connects the right person to the wrong family. Imagine two men named James Brown, both born in Virginia in 1818, but belonging to two entirely unrelated families. An AI tool, pattern-matching on name, place, and approximate date, may merge them, attaching one man's documented parents to the other man's documented children.

Because online family trees copy from each other constantly, an error like this doesn't stay contained. It can spread across dozens or hundreds of public trees within months, becoming "true" simply through repetition.

How to catch it: Whenever two individuals share a name, county, and rough birth year, ask the AI to list the specific evidence distinguishing them: land records, church parishes, associates named as witnesses, distinguishing middle names. If it can't distinguish them, you probably have two people, not one.

#3. Collapsing Two People into One

Related to false parentage, this hallucination happens when the AI treats "same name, similar era, same county" as sufficient proof of "same person." Common names are especially vulnerable: John Williams, William Johnson, Mary Jones, Elizabeth Brown. In many rural communities of the eighteenth and nineteenth centuries, several unrelated people carried the same name at the same time.

How to catch it: Before accepting that two records refer to one individual, ask what independent evidence (land transactions, church membership, military service, or family testimony) ties them together, not just shared name and geography.

#4. Splitting One Person into Several

The mirror-image error: one real person recorded inconsistently, as "Robert," "Robt.," "Bob," and "R.J.," gets treated by the AI as four different people. Suddenly your family tree has artificially multiplied, with duplicate siblings or duplicate spouses cluttering the record.

Nineteenth-century clerks were not consistent spellers, and names were frequently abbreviated, translated, or anglicized. AI models, trained to treat text strings literally, can miss the obvious inference that a human researcher would make instantly.

How to catch it: When a "new" individual appears with a name variant of someone already in your tree, in the same location and timeframe, check whether the records could plausibly describe one person before treating them as separate.

#5. Invented Sources

Perhaps the easiest hallucination to catch once you know to look for it, and one of the most dangerous when you don't. AI tools will sometimes cite a source that sounds entirely plausible: "County Probate Book 14, page 327" or "Virginia Marriage Register, Volume 8." The citation has the right shape, the right format, the right level of specificity. It just doesn't exist.

This mirrors a broader, well-documented problem. A 2026 analysis of AI-generated academic writing found a sharp rise in nonexistent references following widespread adoption of large language models, with a conservative estimate of nearly 147,000 hallucinated citations appearing in 2025 alone. Similar patterns have shown up in legal filings, where courts have caught lawyers submitting briefs built on cases that were never decided because an AI tool invented them. Genealogy is not immune to the same failure mode; if anything, it may be more vulnerable, since archival citation formats are formulaic and easy for a model to imitate convincingly.

How to catch it: Never accept a citation at face value. Look the source up yourself, in the actual archive, library catalog, or online record collection, before recording it as fact. If you can't locate it, assume it doesn't exist until proven otherwise.

#6. Imaginary Historical Context

AI loves to explain. "The family likely moved west because of the Panic of 1837." It's a tidy, plausible sentence. It may even be historically reasonable. But plausible is not the same as documented, and an AI tool rarely marks the difference clearly. Historical interpretation, offered without evidence specific to your family, quietly becomes historical fiction wearing the authority of fact.

How to catch it: Treat any sentence containing words like "likely," "probably," or "given the context of the time" as a hypothesis to test, not a fact to record. Ask what document (a land sale, a migration record, a letter) actually supports the claimed motive.

#7. Overconfident Translation

If your ancestors left records in German, Polish, Latin, French, or Dutch, AI translation can feel like a superpower, until it isn't. AI models translate old church and civil records with impressive fluency, but that fluency can mask serious errors: invented words to fill illegible gaps, misread names, or entire phrases reconstructed from context rather than actually read. Genealogists working with foreign-language parish records have reported AI tools fabricating names, places, and even whole narrative details that don't exist in the original document.

How to catch it: For any translation affecting a name, date, or relationship, ask for a word-by-word rendering alongside the smooth translation, and compare it against the original image yourself, or have a human speaker of the language check ambiguous passages.

#8. Fabricated Relationships

A record lists a witness, Samuel Jones, at a wedding or a will signing. The AI concludes he must be an uncle, brother, father, or cousin. But witnesses in historical documents were frequently neighbors, business associates, fellow church members, or simply friends, not relatives at all. The AI's eagerness to complete the family picture leads it to assign a relationship where the record specifies none.

How to catch it: Treat every unlabeled name in a record (witness, executor, sponsor, bondsman) as a relationship of unknown type until independent evidence establishes otherwise.

#9. Timeline Repair

Large language models dislike contradictions, and historical records are full of them. If a record shows a birth in 1820 and a marriage in 1832, and something else nearby seems inconsistent, an AI tool may quietly adjust one of the dates so everything appears internally consistent. The apparent contradiction vanishes, but so does the actual historical evidence, replaced by a smoothed-over fiction that fits the model's sense of what "should" be true.

How to catch it: When a date changes between one AI response and another, ask why, and check both figures against the original record. Real historical records are often genuinely inconsistent, and that inconsistency is itself information, not a bug to be silently fixed.

#10. Source Laundering

This is probably the least recognized hallucination, and arguably the most insidious. AI tools trained on the vast ocean of user-submitted online family trees absorb whatever errors already exist there, including all the mistakes described above, made by human researchers over the past two decades. The AI then presents those errors back to you as though they were settled historical fact.

Nothing here is invented from scratch. That's what makes it dangerous. The information already circulates online, attached to seemingly credible trees and family history websites, so it carries an unearned appearance of authority. The AI hasn't lied exactly; it has simply repeated community folklore as though it were consensus.

How to catch it: Be especially skeptical of any claim the AI attributes to "commonly accepted" genealogy or "widely documented" family history. Ask what primary source underlies the claim, not what other online trees say about it.

An Eleventh Hallucination: Certainty Inflation

There's one more pattern worth naming, even though it's rarely discussed: certainty inflation. Professional genealogists are trained to say, when the evidence doesn't fully support a conclusion, "the evidence is insufficient." AI tools are much less comfortable with that kind of humility. Instead, they tend to write things like "based on available evidence, John was almost certainly the son of..."

That phrase, "almost certainly," can represent an enormous leap well beyond what the underlying evidence supports. This isn't accidental. Research into why language models hallucinate has found that the training and evaluation processes used to build these systems reward confident, plausible-sounding answers over honest expressions of uncertainty, the AI equivalent of a student guessing on a test rather than leaving a blank. The result is a model that is statistically inclined to sound more certain than the facts warrant.

Toward a Genealogical AI Evaluation Framework

None of this means AI is useless for family history research; far from it. It means AI output needs to be treated the way a careful historian treats any single source: as a claim to be verified, not a fact to be recorded.

One useful discipline is to treat every AI-generated genealogical statement as a research hypothesis, never as evidence in itself. Accept only what can be traced back to an identifiable primary or reliable secondary source, and preserve uncertainty wherever the historical record itself remains uncertain. In practice, that means running each AI claim through a short mental checklist:

Criterion

Key Question

Source authenticity

Is every factual claim traceable to a real, locatable source?

Evidence chain

Can the conclusion be reconstructed from the cited evidence?

Identity confidence

Has the AI distinguished between similarly named individuals?

Chronological consistency

Are dates plausible without silent correction?

Geographic plausibility

Do locations fit historical boundaries and migration patterns?

Uncertainty disclosure

Does the AI distinguish evidence from inference?

Genealogical Proof Standard alignment

Would the conclusion satisfy professional genealogical standards?


Applying a framework like this won't eliminate hallucinations (no current AI system is free of them), but it will keep you from mistaking a fluent, confident answer for a well-documented one. AI can be a genuinely powerful research assistant for genealogy: it can suggest search strategies, summarize long documents, translate unfamiliar scripts, and spot patterns across large record sets faster than any human could alone. Just remember that its job is to help you find the evidence, not to become the evidence itself.

________________________________

Sources:

Kalai, Adam, et al. "Why Language Models Hallucinate." arXiv, 4 Sept. 2025, https://arxiv.org/abs/2509.04664.

Newsham, Jack. "AI Hallucinations in Court Documents Are a Growing Problem, and Data Shows Lawyers Are Responsible for Many of the Errors." Business Insider, 27 May 2026, https://www.businessinsider.com/increasing-ai-hallucinations-fake-citations-court-records-data-2025-5.

Yin, Yian, et al. "LLM Hallucinations in the Wild: Large-Scale Evidence from Non-Existent Citations." arXiv, 8 May 2026, https://arxiv.org/abs/2605.07723.

 

Friday, July 31, 2026

How Historians Should Evaluate AI


The Historian's AI Evaluation Framework (HAIEF): Twelve Questions Before You Trust a Machine-Written Past

by Alvin Blackshear, retired Historian   <ablackshear@gmail.com>
______________________________________________________________

In the summer of 2026, two of the world's largest consulting firms had to pull back work they had already delivered to clients. PwC withdrew AI-generated reports after discovering they contained fabricated citations and references to sources that did not exist.[1] KPMG ran into the same problem: investigators reviewing one of its AI-assisted reports found that most of its citations were fabricated or misleading.[2] Neither firm works in history. Neither incident involved archives, manuscripts, or a single primary source. That is exactly why they matter here.

It would be easy for historians to read those stories and conclude they describe someone else's problem, a corporate reporting failure, a compliance issue, a cautionary tale about consulting culture. That reading is a mistake. What failed at PwC and KPMG was not domain knowledge. It was sourcing discipline: the basic, mechanical trustworthiness of a citation. If large language models fabricate references in financial and regulatory reporting, where the source material is comparatively narrow and well-indexed, there is no reason to expect better behavior in historical research, where the source base is vast, unevenly digitized, and full of the kind of ambiguity these systems are worst at handling. This is not a story about historians being unusually vulnerable to bad AI output. It is a story about a structural failure mode that will show up wherever citation and provenance matter, and few disciplines depend on citation and provenance more than history does.

There is a second, related problem, and it may be more dangerous than outright fabrication because it is harder to catch. Recent reporting on AI accuracy has noted a troubling pattern: as these systems have become more fluent, they have not necessarily become more honest about what they don't know. Instead, they increasingly deliver incorrect information with the same confident tone as correct information, rather than flagging uncertainty or hedging where hedging is warranted. Any historian will recognize this failure mode immediately, because it is precisely what they are trained to distrust in human sources. A memoir written decades after the fact, a diplomatic cable reporting secondhand rumor as established fact, a witness testifying with total conviction about details that turn out to be wrong — historians have centuries of methodological practice built around the fact that confidence is not evidence. A machine that states falsehoods fluently and without hedging is not a new problem for historical method. It is an old problem wearing a new interface.

Put these two failure modes together, fabricated sourcing and confident error, and you arrive at the central argument of this piece: a language model can perform well on the accuracy benchmarks the AI industry currently uses, and still fail every test a professional historian would apply to a piece of research. General-purpose evaluation metrics measure things like factual accuracy rates, hallucination frequency, and human preference scores. None of those metrics ask whether a source actually exists in the form cited, whether a quotation can be independently located, whether a narrative acknowledges live scholarly disagreement, or whether the chronology holds together under scrutiny. A model can score 95 percent on general factual accuracy and still produce a historical narrative that treats a contested interpretation as settled, leans on secondary literature while ignoring available primary sources, and quietly imports present-day assumptions into a past context that didn't share them. None of that shows up in a hallucination-rate benchmark. All of it would get a paper rejected by a competent peer reviewer.

 That gap, between what general AI evaluation measures and what historical method actually requires, is the subject of this article. This is not a piece about whether historians should use AI, or how to prompt it more effectively; that ground has already been covered well by other writers, including practical workflow guides for AI-assisted historical research. This is a piece about something narrower and, so far, largely unaddressed: how a working historian should grade what the AI hands back. It is meant to function less like an essay and more like a QA protocol, a practical checklist you could actually run against a piece of AI-generated research before you decide whether to trust, cite, or build on it. 

What the research actually shows

The clearest empirical precedent for this kind of evaluation comes from a 2026 comparative study that pitted trained historians against generative AI tools: ChatGPT-4, Claude, Gemini, and Perplexity, on parallel case studies in ancient history. The researchers found a consistent pattern: AI tools were genuinely strong at broad content discovery and thematic synthesis, the kind of work that benefits from fast pattern-matching across large bodies of text. But they were consistently weaker than human historians on source verification, chronological reasoning, provenance, and contextual interpretation, which are the parts of the discipline that depend on judgment rather than retrieval. The study's authors argued for a human-in-the-loop approach that treats AI as a research accelerant rather than a substitute for scholarly adjudication.

A separate line of research has approached the problem from the model side rather than the historian's side, building benchmarks specifically designed to test historical reasoning. That work has found that even the strongest available models struggle with the kind of sophisticated evidentiary reasoning that is routine for a trained historian, the weighing conflicting accounts, recognizing genre and authorial bias, and reconstructing plausible sequences of events from incomplete evidence. This is a useful corrective to the assumption that model capability will simply catch up with historical judgment as systems improve. The gap is not just about scale; it is about a kind of reasoning these systems are not currently built to do well.

Not every recent contribution to this literature is critical of AI's role in historical work. Case studies of experienced historians integrating AI into their actual research workflows, mapping specific tasks to language models, to conceptual work the historian retains, and to conventional computational tools, that describe an approach where AI expands capacity without replacing judgment, provided the historian maintains rigorous cross-checking at every stage. That distinction matters for this piece: the argument here is not that AI is useless for historical work. It's that historians need a way to grade the output that general AI evaluation frameworks were never designed to provide. 

Introducing the framework: HAIEF[3]

What follows is a two-part tool. The first is a practitioner-facing checklist, including twelve diagnostic questions meant to be run quickly against any piece of AI-generated historical research, the way you'd run a gut-check before deciding whether something is worth building on. The second is a more formal, citable rubric that breaks the same underlying concerns into eight scored criteria, useful when you need to document *why* a piece of AI output was accepted or rejected, not just whether it was.

The twelve questions:

1. Are all citations real?

2. Can every quotation be independently located?

3. Is the chronology internally consistent?

4. Does the AI distinguish primary from secondary sources?

5. Are archival collections correctly identified?

6. Are names, dates, and places cross-checked?

7. Does the narrative reflect historiographical debate?

8. Does the AI acknowledge uncertainty?

9. Does it introduce present-day assumptions into historical contexts?

10. Can another historian reproduce the research?

11. Which claims require archival verification?

12. What percentage of the output would I actually trust without checking?

That last question is deliberately blunt, and it's the one I'd argue matters most in practice. Every other question on the list is diagnostic.  It tells you where a piece of AI research is likely to break down. Question twelve forces a number out of you, an honest estimate of how much of the output you'd actually stand behind without independent verification. In my own use, that number is rarely above 40 or 50 percent for anything involving primary-source claims, and treating that as a normal, expected ceiling, rather than a failure of the tool, is the right posture to bring to this work.

The formal rubric maps the same territory onto eight criteria: citation authenticity (do the cited works exist and actually support the claim made), provenance (is every factual assertion traceable to a reliable source), chronological accuracy (do the dates and sequences hold together internally), primary-source fidelity (are quotations and archival references accurate to the original), historiography (does the output acknowledge that scholarly consensus is contested where it is contested), context (are events read within their own historical setting rather than through present-day assumptions), uncertainty (does the output distinguish established fact from inference), and reproducibility (could another historian, working from the same cited evidence, independently reach the same conclusions).

The value of running both versions side by side is that they serve different moments in the workflow. The twelve questions are what you ask while you're reading the output, in real time, deciding whether to keep going or start over. The eight-criterion rubric is what you fill out afterward, when you need a defensible record of why a source was trusted or discarded, which is useful for peer review, for training research assistants, or for documenting due diligence in a professional or editorial context. 

Who this is for

Historians and academic researchers are the most obvious audience, and the stakes for them are the highest: a fabricated citation or an unacknowledged historiographical dispute that makes it into a submitted paper is a professional liability, not just an inconvenience.

Genealogists are arguably the highest-volume, lowest-scrutiny users of AI-assisted historical research, and they are almost entirely unaddressed by the academic literature on this problem. Family history research routinely involves exactly the kind of claims, be it names, dates, places, relationships, that AI systems handle with the least reliability, and genealogists are less likely than academic historians to have institutional habits of source verification already built in.

Journalists covering historical topics under deadline face a particular version of this problem: the twelve-question checklist doubles as a fast fact-checking pass when there isn't time for a full archival dig, and question twelve, what percentage would you trust without checking, is a useful discipline for anyone filing on a clock.

Educators like myself have perhaps the most durable use case. Detecting AI-generated text is a losing, ever-shifting battle. Teaching students to interrogate AI output using a framework like this one, and not asking "did a machine write this" but rather, "would this survive a historian's scrutiny", is a skill that remains useful regardless of how the underlying technology changes. 

Closing

None of this should be mistaken for a claim that a checklist can do a historian's job. A framework like HAIEF can catch a fabricated citation, flag an inconsistent chronology, or surface a claim that needs archival verification before it gets treated as settled. What it cannot do is listen. Jan Burzlaff has argued that AI systems summarize but do not listen, reproduce but do not interpret.  These tools  achieve coherence while faltering at contradiction, and that historical writing is not simply an output to be optimized but a form of presence, a risk taken in the act of trying to make meaning where no prewritten frame will do. That distinction is worth sitting with rather than resolving. A verification framework can tell you whether the sourcing holds up. It cannot tell you whether the history has been understood. Those are different questions, and confusing them, mistaking a clean citation audit for genuine historical interpretation, may turn out to be the more consequential failure mode of all.

Sources

Burzlaff, Jan. "Fragments Not Prompts: Five Principles for Writing History in the Age of AI." Rethinking History (2026). https://doi.org/10.1080/13642529.2025.2546174.

Campbell, Chris. "The Historian in the Age of AI." Transactions of the Royal Historical Society (December 10, 2025). https://www.cambridge.org/core/journals/transactions-of-the-royal-historical-society/article/historian-in-the-age-of-ai/37E3B742A2983DF4DAA2E38D48252F89.

"Generative AI as a Historical Source: Source Criticism, Citation Integrity, and the Jagged Frontier of Digital History." (March 20, 2026). https://acnsci.org/journal/index.php/cte/article/view/1438.

Fox, Yaniv. Using Generative AI in Historical Practice. Cambridge Elements. Cambridge: Cambridge University Press, 2026. https://www.cambridge.org/core/elements/abs/using-generative-ai-in-historical-practice/7C1392A6E9DBD379FAA42E2D16A6D45B.

Henriot, Christian. "The AI-Augmented Research Process: A Historian's Perspective." Preprint, arXiv, August 3, 2025. https://arxiv.org/abs/2508.01779.

"HistoRAG: Embedding Historical Methodology in Retrieval-Augmented Generation." Preprint, arXiv, June 16, 2026. https://arxiv.org/abs/2606.18103.

"Can LLMs Act as Historians?" Preprint, arXiv, April 27, 2026. https://arxiv.org/abs/2604.24690.

Ruchniewicz, Krzysztof. "AI in History Education: The American Historical Association's Guidelines, the Views of Their Authors, and Their Reception within the Historical Community." Studia Historica 62, no. 1 (2026). https://doi.org/10.14746/sh.2026.62.1.003.

Solga, Raymond S., and Mohammed J. Sarwar. "Evaluating Generative AI in Historical Research: A Comparative Study on Identifying Primary Source Evidence in Ancient History." AI & Antiquity 2, no. 1 (February 26, 2026). https://doi.org/10.64946/aiantiquity.v2i1.003.

Financial Times. "PwC Withdraws AI-Generated Reports over Fabricated Citations." July 29, 2026. https://www.ft.com/content/7e149ac8-2ce2-4266-8940-192f9821b33c.

TechRadar Pro. "A Major KPMG Report on AI Was Found to Be Chock-Full of AI Hallucinations." June 12, 2026. https://www.techradar.com/pro/a-major-kpmg-report-on-ai-was-found-to-be-chock-full-of-ai-hallucinations.

Axios. "AI Chatbots Are Getting More Confident — and Sometimes More Wrong." May 30, 2026. https://www.axios.com/2026/05/30/ai-accuracy-chatbots-hallucinations.

Thursday, July 30, 2026

Assemblies of Sorrow: Performances of Black Endangerment in the Jim Crow Era - REVIEW


by Samuel Ng

During the Jim Crow era, Black activists appealed to a diverse population of migrating Black Americans by making them viscerally feel that the threat of anti-Black violence continued to afflict them as a group and to undergird blackness itself. To this end, they organized public gatherings, mostly comprised of Black people, that fostered fears of looming physical harm.

Samuel Galen Ng illuminates this Black consciousness as it emanated from feelings of collective endangerment. The dissemination and intensification of such feelings became a pivotal way of solidifying a national Black consciousness on the eve of the Civil Rights Movement. Ng examines how performances of Black endangerment performed political work that provided Black people with important means of political organizing and insurgency. As Ng shows, the grief and mourning that took place at the performances provided public spaces for individuals and communities to observe specific losses capable of impacting Americans across the country.

Ambitious and interdisciplinary, Assemblies of Sorrow explores an overlooked facet of Black organizing and protest and traces how activists shaped fear and grief into political action.

Samuel Galen Ng is an associate professor of Africana studies at Smith College.

University of Illinois Press
ISBN-13: 978-0252089404

Monday, May 18, 2026

Integration at Second Base: Jackie Robinson and the Quest for Black Citizenship - REVIEW


by Peter Eisenstadt

Jackie Robinson is one of the most enduring icons of the great American pastime―the man who broke baseball’s color line in the twentieth century, opening the door for his fellow professionals and allowing rising generations to dream of fame and glory on the diamond. But for number 42, playing for the Dodgers was just a beginning. As Peter Eisenstadt demonstrates in this compelling new biography, Robinson’s trailblazing journey was more than a role that fate thrust on him―it was politically informed and consciously connected in Robinson’s mind to a vision of integration and full Black citizenship.

When he ventured out of the Negro Leagues and into the majors, as the league’s sole Black player, his triumph could have stopped at mere tokenism. Eisenstadt reveals a more ambitious goal on Robinson’s part, as well as a side to the great sports hero we have never fully appreciated. This book explores the political and spiritual roots of Jackie Robinson’s quest for Black citizenship from his boyhood in Pasadena to his service days―during which he was court-martialed for refusing to change seats on a segregated bus―to a transcendent athletic career that included an MVP award, a World Series victory, and eventually a place in the Hall of Fame. In his life after baseball, Robinson went on to serve as a civil rights leader, columnist, and political advocate.

The determination that spurred his great achievements was always accompanied by an understanding of just how far society still needed to go: despite his success, at the end of his life he was convinced that he “never had it made.” In telling the story of Robinson’s remarkable life, this book sheds invaluable light on the complex meanings of integration.

Peter Eisenstadt is Affiliate Professor of History at Clemson University and the author of Against the Hounds of Hell: A Life of Howard Thurman.

University of Virginia Press
ISBN-13: 978-0813955001