Showing posts with label Historical AI. Show all posts
Showing posts with label Historical AI. Show all posts

Wednesday, August 12, 2026

The Biggest Mistake Historians Make When Using ChatGPT

 

by Alvin Blackshear  |  Historian & Researcher

A historian asks ChatGPT a seemingly straightforward question: “Who was the first African American elected to public office in Delaware after Reconstruction?

Within seconds, the system produces a polished answer. It supplies a name, an election date, a description of the office, and perhaps even several citations. The prose is orderly, confident, and specific. Nothing about it appears careless. The historian copies the response into a research notebook and proceeds to the next question.

Only later does the answer begin to unravel. One citation does not exist. A quotation attributed to a nineteenth-century newspaper is actually a modern paraphrase. The date belongs to another officeholder. An appointment has been confused with an election. Details from two people have been combined into a single biography.

Nothing in the original response sounded implausible. That is precisely the danger.

The greatest mistake historians make when using ChatGPT is not merely trusting an occasional incorrect answer. It is treating AI-generated prose as though it were historical evidence. ChatGPT produces narratives. Historians require evidence. Those activities may sometimes overlap, but they are not the same.

The Abandonment of Historical Method

Historians already possess a sophisticated method for evaluating claims. When examining a letter, newspaper, memoir, photograph, government record, or oral history, they ask familiar questions. Who created the source? When was it created? For what purpose? Under what circumstances? What interests shaped its contents? Can its claims be corroborated? Has another scholar interpreted the evidence differently?

These questions constitute more than an academic ritual. They are the means by which historical knowledge is distinguished from rumor, memory, advocacy, folklore, and invention.

Yet something curious happens when many researchers begin using generative artificial intelligence. The normal habits of source criticism are suspended. A response is accepted because it sounds reasonable. A citation is trusted because it resembles an academic citation. A quotation is repeated because the language seems appropriate to the period.

The historian who would never accept an unsigned reminiscence without examining its provenance may accept an AI-generated paragraph without asking where any of its claims originated.

ChatGPT should therefore be treated as an intelligent research assistant, not as a documentary source. It can suggest search terms, identify possible lines of inquiry, compare concepts, organize notes, formulate questions, and help clarify difficult prose. It can also point a researcher toward archives, books, people, and events that merit investigation. Its usefulness, however, does not convert its output into evidence.

The distinction is fundamental. A research assistant may tell a historian that a document probably exists. The historian must still locate and examine it.

Why Fluency Is So Persuasive

Large language models are especially dangerous when they are almost correct. An obviously absurd answer invites skepticism. A plausible answer, containing real names, accurate background information, and one fabricated detail, may pass unnoticed.

ChatGPT’s authority is largely rhetorical. It writes in complete sentences, arranges events chronologically, supplies transitions, and often presents conclusions without visible hesitation. Readers naturally associate these qualities with knowledge. Confidence, specificity, and coherence become substitutes for provenance.

Recent research suggests that this problem is not simply the result of careless prompting. Kalai and his colleagues argue that common accuracy-based evaluations may reward models for guessing instead of acknowledging uncertainty. If a benchmark gives credit only for correct answers and imposes no meaningful cost for plausible errors, the model is encouraged to answer even when abstention would be more responsible. The researchers propose open scoring rubrics that disclose how errors and abstentions will be evaluated.

This observation has important consequences for historians. A model that always produces an answer may appear more useful than one that frequently says, “I do not have enough evidence.” In historical research, however, an admission of uncertainty may be the more accurate and intellectually responsible response.

The historian must therefore evaluate not only what the system says, but whether the system should have answered at all.

Five Forms of Historical Hallucination

Historical hallucinations frequently assume recognizable forms.

The first is the invented citation. ChatGPT may produce a realistic book title, journal article, archival collection, author, volume number, or page range. Every component may appear academically credible even though the source does not exist.

The second is the misquotation of a primary source. A model may modernize, compress, or reconstruct a statement and then present the resulting language inside quotation marks. The general sentiment may be accurate, but the words are not documentary evidence.

The third is the composite biography. Details belonging to several individuals may be joined into one apparently coherent life. This is especially dangerous when researching people who share names, occupations, institutions, military units, or geographic locations.

The fourth is incorrect chronology. A model may identify real events but arrange them in the wrong order, confuse the date of an appointment with the date of an election, or place a person at an institution before that person arrived there.

The fifth is false causation. ChatGPT may connect two events with phrases such as “therefore,” “as a result,” or “this led to,” even when the evidence establishes only sequence or correlation.

These errors are not equally visible. An invented person may be discovered quickly. A subtle chronological error or unsupported causal inference may survive several rounds of editing because it fits an expected narrative.

Why Better Prompts Are Not Enough

Prompt design can improve an AI response. A historian can request citations, ask the system to distinguish facts from interpretations, require expressions of uncertainty, or instruct it not to invent missing information. Such practices are worthwhile.

They are not verification.

A carefully written prompt may reduce the probability of an error, but it cannot transform generated language into historical evidence. Even a response that includes citations must be checked against the cited material. Even a quotation accompanied by a page number must be located on that page. Even a claim presented as certain must be corroborated.

Research on hallucinations in academic writing identifies fabricated citations, factual distortion, named-entity errors, logical inconsistency, and propagation errors as threats to scholarly integrity. The final responsibility remains with the human author. An historian cannot excuse a false statement by explaining that ChatGPT supplied it.

Better prompting is therefore a research technique, not an evidentiary standard.

Testing AI as Historians Test Evidence

Historians need a repeatable method for testing AI-generated research. The evaluation should move beyond impressions such as “the answer seemed good” or “the model performed well.”

A practical scoring rubric might assess eight dimensions:

1.      Citation existence: Does every cited work or archival collection exist?

2.      Citation accuracy: Does the cited source support the specific claim?

3.      Quotation fidelity: Are quoted words reproduced exactly and in context?

4.      Identity accuracy: Have people with similar names or careers been distinguished?

5.      Chronological accuracy: Are dates and sequences correct?

6.      Geographic and institutional accuracy: Are places, offices, organizations, and jurisdictions correctly identified?

7.      Provenance: Can each important assertion be traced to accessible evidence?

8.      Historiographical consistency: Does the answer acknowledge significant scholarly disagreement?

Each category could be scored from zero to four. A zero would indicate fabrication or complete failure. A score of one would reflect major errors. Two would indicate partial support or substantial ambiguity. Three would represent generally accurate work with minor deficiencies. Four would require full and independently verified support.

A weighted rubric would assign greater penalties to invented citations, false quotations, and merged identities than to minor stylistic or contextual omissions. This matters because not all errors have the same scholarly consequence.

Benchmark testing should also use a fixed set of questions rather than memorable anecdotes. Researchers could create a collection of historical questions at several difficulty levels: well-documented national events, obscure local events, disputed interpretations, biographical identity problems, chronological puzzles, and questions whose correct answer is genuinely unknown.

The same questions could then be submitted to several AI systems under controlled conditions. Evaluators would record accuracy, citation validity, unsupported assertions, appropriate abstentions, and changes across repeated runs. Such testing would reveal not only whether a model produces correct answers, but how it fails.

What RAG Can and Cannot Do

Retrieval-augmented generation, commonly known as RAG, offers one method of improving reliability. A RAG system retrieves documents from a designated collection before generating an answer. For historians, that collection might contain archival finding aids, digitized newspapers, oral-history transcripts, government records, or scholarly articles.

This approach can make AI responses more transparent and more closely connected to identifiable sources. It may reduce fabricated citations and permit researchers to inspect the evidence used in producing an answer.

Yet RAG does not eliminate the need for historical judgment.

Recent research distinguishes factuality from faithfulness. A statement may be faithful to a retrieved document while the document itself is inaccurate, biased, incomplete, or misinterpreted. Conversely, a statement may be factually correct but unsupported by the particular documents retrieved. A serious evaluation must test both dimensions.

For historians, retrieval is only the beginning. The fact that a system retrieved a newspaper article does not establish that the article was truthful. The fact that a claim appears in a memoir does not establish that memory was reliable. The fact that several sources repeat a statement does not prove they were independent.

RAG can retrieve evidence. It cannot perform the full intellectual work of source criticism.

The Historian’s Continuing Responsibility

ChatGPT does not threaten historical scholarship simply because it occasionally invents facts. Historical scholarship has always confronted error, forgery, propaganda, partial memory, and misleading testimony.

The greater danger arises when historians allow the fluency of artificial intelligence to replace the discipline’s oldest habit: skepticism.

For centuries, historians have questioned manuscripts, memoirs, newspapers, photographs, statistics, and archives. Artificial intelligence deserves no exemption from that tradition. Its output should be tested, scored, corroborated, and traced to evidence.

The historian’s task remains unchanged. It is not merely to discover information or compose a persuasive narrative. It is to determine what the surviving evidence permits us to say, what remains uncertain, and why the distinction matters.

References

  1. Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. “Evaluating Large Language Models for Accuracy Incentivizes Hallucinations.” Nature 653 (2026): 1047–1051. https://www.nature.com/articles/s41586-026-10549-w
  2. Rahman, Subhey Sadi, et al. “Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models.” Artificial Intelligence Review 59 (2026), article  70. https://link.springer.com/article/10.1007/s10462-025-11454-w
  3. Fadeeva, Ekaterina, et al. “Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation.” Findings of the Association for Computational Linguistics: ACL 2026 (2026): 6814–6836. https://aclanthology.org/2026.findings-acl.338
  4. Giray, Louie, Md Emdadul Islam, and John Arvin Glo. “AI Hallucinations in Academic Writing: Implications for Research Integrity.” Naunyn-Schmiedeberg’s Archives of Pharmacology (2026). https://pubmed.ncbi.nlm.nih.gov/42217043 

Monday, August 03, 2026

Beyond Correlation: How Historians Distinguish Causation, Coincidence, and Contingency in the Age of AI


by Alvin Blackshear | Historian & Researcher  <ablackshear@gmail.com>

There is a particular kind of intellectual vertigo that comes from watching a machine narrate the past with total confidence. Ask an AI system to explain why a revolution happened, why a movement gained momentum, or why a community's history unfolded as it did, and it will almost always answer. The prose will be fluent, the timeline will be clean, and the causal claims will arrive stated as fact rather than as argument. What the answer will rarely do is pause to ask whether the events it has linked together were actually connected at all, or whether they simply happened to occur near one another in time. This is the quiet crisis facing historical writing in an era of generative AI: the technology is exceptionally good at producing narratives that sound causal, and considerably less good at knowing whether those narratives are true.

Historians have spent more than a century developing tools to guard against exactly this failure. Long before anyone worried about large language models hallucinating a tidy story out of scattered facts, historians were worried about themselves doing the same thing. The discipline's entire methodological apparatus, footnotes, source criticism, peer review, historiographical debate, exists in large part to slow down the impulse to declare that because one thing followed another, the first caused the second. What AI-generated history reintroduces, at scale and with new fluency, is an old temptation: the temptation to mistake sequence for cause, and pattern for proof.

The Seduction of Sequence

Correlation is easy to notice and hard to interpret responsibly. Two trends move together, two events happen in the same decade, two figures cross paths, and the mind supplies a connection because connection is a more satisfying story than coincidence. Sarah Woodhouse's recent work on correlation reminds us that not all statistical relationships imply causation, and that there are several distinct varieties of correlation, only one of which reflects a genuine causal link between the variables involved.[1] The others might reflect a shared underlying cause, a coincidence of timing, a feedback loop running in the opposite direction from what intuition suggests, or simply noise that resembles a pattern if you squint at it long enough.

Historians face this same interpretive challenge, except their raw material is not a spreadsheet but the messy residue of human lives: letters, court records, newspapers, oral testimony, material culture. When two developments appear together in this record, a rise in labor unrest alongside a shift in immigration policy, a change in religious practice alongside an economic downturn, a legal victory alongside a change in public rhetoric, the historian's task is to ask what actually connects them, if anything does. Johannes Nagel's comparative analysis of causal reasoning in historical scholarship makes a useful observation here: historians rarely operate with the clean methodological categories that philosophers of science might prefer. Real historical practice is often described as vaguely causal-analytical, blending causal-intentional explanation, attention to unintended consequences, and selective use of theory and comparison, without ever fully resolving into a single rigorous method.[2] This is not a failure of the discipline. It is an honest reflection of how difficult the underlying problem actually is.

AI systems, by contrast, do not experience this difficulty as friction. A model trained to produce plausible continuations of text will happily supply a causal mechanism because causal mechanisms make for better sentences than admissions of uncertainty. The danger is not that AI invents facts out of nothing, though it sometimes does that too. The deeper danger is that it takes real facts, arranges them in a narrative shape, and lets the shape itself imply a causal claim that no historian has actually verified.

Mechanism, Not Just Chronology

If there is a single principle that separates serious historical causal argument from mere narrative sequencing, it is the requirement of an explicit mechanism. It is not enough to say that event A preceded event B and therefore contributed to it. A historian is expected to explain how A produced B: through what institutions, what decisions, what material pressures, what beliefs held by which people.

Daniel Little's work on the microfoundations of social explanation is instructive here. Little argues that sweeping historical forces such as capitalism, the state, or industrialization do not themselves act as causal agents in any direct sense. They are abstractions that describe the aggregate result of countless individual decisions, made by people operating within specific institutions and constrained by specific incentives and beliefs. A causal historical explanation, on this view, has to eventually cash out in terms of those individuals: what they knew, what they wanted, what options were actually available to them, and how their choices interacted with those of others to produce a larger outcome. This is demanding work. It requires the historian to move between scales, from the granular texture of a single decision to the broader pattern that many such decisions eventually create, without losing the thread that connects the two.[3]

Mark Hewitson's argument about the marginalization of causal thinking within the historical profession adds an important wrinkle to this picture. Hewitson suggests that the turn toward language, discourse, and representation that reshaped historical writing over the past several decades, whatever its genuine intellectual gains, also made many historians reluctant to speak plainly about why things happened. Causal explanation came to seem naive, or worse, complicit in older positivist ambitions to discover fixed laws of history. And yet, as Hewitson points out, historians never actually stopped making causal arguments. Every account of a revolution, an election, a migration, or a social movement is saturated with implicit claims about why it occurred. What was lost was not the practice of causal reasoning but the willingness to examine that practice openly and hold it to a clear standard.[4]

This matters enormously for evaluating AI-generated history, because a system that has absorbed decades of scholarship in which causal claims are made implicitly, without explicit justification, will likely reproduce that same implicitness. It will offer causal-sounding narratives without ever surfacing the mechanism that supposedly connects cause to effect, because the training material it draws from often does the same thing. The corrective, for both human historians and the tools they now use, is the same: state the mechanism, or admit that you cannot.

What Coincidence and Contingency Demand of Us

Not everything that happens in history happens for a deep structural reason. Some things happen because a particular person was in a particular place at a particular moment, because a letter arrived a day late, because an illness struck one leader and spared another. Historians call this contingency, and distinguishing it from genuine causation is one of the discipline's more delicate tasks.

The stakes of this distinction are not merely academic. When causal narratives are built where only coincidence existed, the resulting history can flatten the genuine unpredictability and agency that shaped real events into something that looks inevitable, as though it could only have happened the way it did. This has particular consequences for histories that have too often been treated as footnotes to a larger, supposedly more central narrative. African American history, feminist history, and the history of gay and queer communities have all, at various points, been subjected to accounts that either erase contingency entirely, presenting outcomes as foregone conclusions of impersonal social forces, or that erase causation entirely, treating hard-won structural change as a matter of isolated coincidence rather than sustained organizing and argument. Getting the distinction right, separating what was truly contingent from what was the product of deliberate, structural, causal effort, is not a technical nicety. It is a matter of according historical actors the accuracy they deserve, whether that means recognizing the genuine unpredictability they navigated or recognizing the genuine causal power of what they built.

Tay Jeong's work on counterfactual reasoning offers historians one of their sharpest tools for making this distinction. The basic move is to ask what would have happened in the absence of the proposed cause. If removing a particular event or decision from the historical record would plausibly have left the outcome largely unchanged, the causal claim weakens considerably. If removing it would plausibly have produced a genuinely different outcome, the causal claim gains support. Jeong's more technical contribution is to note that counterfactual reasoning is most useful not as a blanket method applied indiscriminately, but in specific situations, particularly where two or more candidate causes seem to be doing similar explanatory work and our untrained intuitions about which one truly mattered are likely to mislead us. This is precise, careful reasoning, and it is exactly the kind of reasoning that a fluent AI narrative can obscure simply by never raising the question in the first place.[5]

The Discipline of Alternative Explanations

Good historical argument does not simply assert a cause and move on. It considers what else might explain the same outcome, and it explains why the preferred causal account is more persuasive than the alternatives. This habit of mind, developed across the humanities, has an unexpected ally in fields that seem far removed from history. Hannah Correia and her coauthors, writing about causal inference in environmental research, describe systematic approaches for distinguishing genuine causal relationships from mere observational association, approaches built around identifying confounding variables and establishing plausible mechanisms before accepting a causal claim.[6] Linbo Wang's comparative overview of causal inference frameworks, spanning the traditions associated with Rubin, Pearl, and structural equation modeling, makes a similar point from a more statistical angle: rigorous causal reasoning requires explicit assumptions, and different frameworks make different assumptions visible in different ways.[7] Though these authors are not writing about history, their underlying discipline, of refusing to accept a causal story until alternatives have been seriously weighed, translates directly into the historian's craft.

This is precisely the discipline that AI-generated narratives tend to skip. A well-trained model can produce a fluent account of why a particular reform succeeded, but it rarely volunteers the alternative explanations a careful historian would consider and reject: that the reform succeeded for reasons largely unrelated to the actors typically credited, that its apparent success was partly an artifact of how success was later measured, or that multiple contributing factors were entangled in ways that resist a single tidy story.

Evidence, Uncertainty, and the Historian's Honesty

Underneath all of this sits the question of evidence. Every causal claim a historian makes rests on some body of evidence, and that evidence varies enormously in quality, completeness, and reliability. Xinyue Chen and colleagues, in their study of how historians actually use visualization in their published work, found that historians face persistent practical and epistemological barriers when trying to represent uncertainty and justify their conclusions to readers and peers.[8] The justification burden, as they describe it, is not a peripheral concern. It is central to what makes a piece of historical writing trustworthy.

Georg Iggers, reflecting late in his career on the historian's role as an engaged intellectual, insisted that this burden cannot be set aside even by historians who write from a place of moral or political conviction. Iggers rejected the idea that historians must choose between detached objectivity and partisan advocacy.[9] Instead, he argued for a middle position: historians can and should be engaged with the pressing concerns of their time, but that engagement must never come at the cost of rigorous evidentiary standards. A historian writing feminist history, or African American history, or the history of gay and queer life, is no less bound by these standards than any other scholar, and indeed the stakes of getting the causal story right, of neither overstating structural inevitability nor understating hard-won agency, are often higher precisely because these histories have so frequently been distorted or dismissed.

Frank Ankersmit pushes this even further by reminding us that the past itself no longer exists in any directly accessible form. Historians do not reproduce it; they represent it, using an inherited semantic apparatus of meaning, truth, and reference that Ankersmit argues must be handled with care rather than assumed to work automatically.[10] If historical writing is, at some level, an act of representation rather than simple description, then AI-generated historical narrative inherits this same burden. It too is representing the past, not retrieving it, and representation always involves choices about emphasis, framing, and inclusion that deserve scrutiny rather than passive acceptance.

A Standard Worth Holding

What emerges from this body of scholarship is not a single formula but a disposition, a way of approaching any historical claim, human-authored or machine-generated, with a consistent set of questions. Is there a stated mechanism connecting cause to effect, or only a sequence dressed up as causation? Have alternative explanations been seriously considered and set aside for identifiable reasons? Has the line between structural causation and genuine contingency been drawn carefully, rather than assumed? Would the outcome have looked meaningfully different if the proposed cause were absent? And does the account acknowledge the limits of its own evidence, rather than presenting a confident narrative where the underlying record is thin or contested?

These questions predate artificial intelligence by decades, and they will outlast any particular model or platform. What AI has changed is the volume and fluency with which causal-sounding narratives can now be produced, and therefore the urgency of applying these questions consistently. The historians whose work informs this discussion were not writing with AI in mind. They were writing because the problem of causation in history has always been genuinely hard, and because getting it wrong has always carried real costs, especially for the histories of communities whose stories have too often been told either as inevitable or as accidental, when the truth usually lies in the harder, more interesting space between. That space, between correlation and causation, between coincidence and contingency, between chronology and mechanism, is where careful historical work has always lived. It is also, increasingly, where the responsible use of AI in historical research will need to live too.

---

Notes:

[1] Sarah Woodhouse, "When Two Things Seem Linked (but Aren't): Understanding Different Types of Correlations," Clearer Thinking, July 2, 2025, https://www.clearerthinking.org/post/when-two-things-seem-linked-but-aren-t-understanding-different-types-of-correlations.

[2] Johannes Nagel, "Causal Explanation in History: An Analysis of Research Strategies," Analyse & Kritik, March 12, 2026, https://doi.org/10.1007/s11577-026-01059-8.

[3] Daniel Little, Microfoundations, Method, and Causation: On the Philosophy of the Social Sciences (New Brunswick, NJ: Transaction Publishers, 1998).

[4] Mark Hewitson, History and Causality (Basingstoke: Palgrave Macmillan, 2014).

[5] Tay Jeong, "When to Use Counterfactuals in Causal Historiography," Sage Journals, February 2, 2025, https://doi.org/10.1177/00491241251314039.

[6] Hannah E. Correia et al., "Best Practices for Moving from Correlation to Causation in Environmental Research," Nature Communications, February 24, 2026, https://rdcu.be/fxl0y.

[7] Linbo Wang, "Causal Inference: A Tale of Three Frameworks," Journal of Data Science, February 11, 2026, https://doi.org/10.6339/25-JDS1211.

[8] Xinyue Chen et al., "How Historians Use Visualization: A Corpus-Backed Taxonomy and Analysis for Cross-Disciplinary Practice," July 10, 2026, https://doi.org/10.48550/arXiv.2605.01456.

[9] Georg Iggers, "The Historian as an Engaged Intellectual: Historical Writing and Social Criticism – A Personal Retrospective," in The Engaged Historian: Perspectives on the Intersections of Politics, Activism and the Historical Profession, ed. Stefan Berger (New York: Berghahn Books, 2019), 277–299.

[10] Frank Ankersmit, Historical Representation (Stanford, CA: Stanford University Press, 2002).

---

Works Cited

Ankersmit, Frank. Historical Representation. Stanford, CA: Stanford University Press, 2002.

Chen, Xinyue, et al. "How Historians Use Visualization: A Corpus-Backed Taxonomy and Analysis for Cross-Disciplinary Practice." July 10, 2026. https://doi.org/10.48550/arXiv.2605.01456.

Correia, Hannah E., et al. "Best Practices for Moving from Correlation to Causation in Environmental Research." Nature Communications, February 24, 2026. https://rdcu.be/fxl0y.

Hewitson, Mark. History and Causality. Basingstoke: Palgrave Macmillan, 2014.

Iggers, Georg. "The Historian as an Engaged Intellectual: Historical Writing and Social Criticism – A Personal Retrospective." In The Engaged Historian: Perspectives on the Intersections of Politics, Activism and the Historical Profession, edited by Stefan Berger, 277–299. New York: Berghahn Books, 2019.

International Commission for the History and Theory of Historiography (ICHTH). 2026 Congress Proceedings. https://www.ichth.net/content/home.html.

Jeong, Tay. "When to Use Counterfactuals in Causal Historiography." Sage Journals, February 2, 2025. https://doi.org/10.1177/00491241251314039.

Little, Daniel. Microfoundations, Method, and Causation: On the Philosophy of the Social Sciences. New Brunswick, NJ: Transaction Publishers, 1998.

Nagel, Johannes. "Causal Explanation in History: An Analysis of Research Strategies." Analyse & Kritik, March 12, 2026. https://doi.org/10.1007/s11577-026-01059-8.

Wang, Linbo. "Causal Inference: A Tale of Three Frameworks." Journal of Data Science, February 11, 2026. https://doi.org/10.6339/25-JDS1211.

Woodhouse, Sarah. "When Two Things Seem Linked (but Aren't): Understanding Different Types of Correlations." Clearer Thinking, July 2, 2025. https://www.clearerthinking.org/post/when-two-things-seem-linked-but-aren-t-understanding-different-types-of-correlations.