The Historian's AI Evaluation Framework (HAIEF): Twelve Questions Before You Trust a Machine-Written Past
by Alvin Blackshear, retired Historian <ablackshear@gmail.com>
__________________________________________________________________________
In the summer of 2026, two of the world's largest consulting
firms had to pull back work they had already delivered to clients. PwC withdrew
AI-generated reports after discovering they contained fabricated citations and
references to sources that did not exist.[1]
KPMG ran into the same problem: investigators reviewing one of its AI-assisted
reports found that most of its citations were fabricated or misleading.[2]
Neither firm works in history. Neither incident involved archives, manuscripts,
or a single primary source. That is exactly why they matter here.
It would be easy for historians to read those stories and
conclude they describe someone else's problem, a corporate reporting failure, a
compliance issue, a cautionary tale about consulting culture. That reading is a
mistake. What failed at PwC and KPMG was not domain knowledge. It was sourcing
discipline: the basic, mechanical trustworthiness of a citation. If large
language models fabricate references in financial and regulatory reporting,
where the source material is comparatively narrow and well-indexed, there is no
reason to expect better behavior in historical research, where the source base
is vast, unevenly digitized, and full of the kind of ambiguity these systems
are worst at handling. This is not a story about historians being unusually
vulnerable to bad AI output. It is a story about a structural failure mode that
will show up wherever citation and provenance matter, and few disciplines
depend on citation and provenance more than history does.
There is a second, related problem, and it may be more
dangerous than outright fabrication because it is harder to catch. Recent
reporting on AI accuracy has noted a troubling pattern: as these systems have
become more fluent, they have not necessarily become more honest about what
they don't know. Instead, they increasingly deliver incorrect information with
the same confident tone as correct information, rather than flagging
uncertainty or hedging where hedging is warranted. Any historian will recognize
this failure mode immediately, because it is precisely what they are trained to
distrust in human sources. A memoir written decades after the fact, a
diplomatic cable reporting secondhand rumor as established fact, a witness
testifying with total conviction about details that turn out to be wrong —
historians have centuries of methodological practice built around the fact that
confidence is not evidence. A machine that states falsehoods fluently and
without hedging is not a new problem for historical method. It is an old
problem wearing a new interface.
Put these two failure modes together, fabricated sourcing
and confident error, and you arrive at the central argument of this piece: a
language model can perform well on the accuracy benchmarks the AI industry
currently uses, and still fail every test a professional historian would apply
to a piece of research. General-purpose evaluation metrics measure things like
factual accuracy rates, hallucination frequency, and human preference scores.
None of those metrics ask whether a source actually exists in the form cited,
whether a quotation can be independently located, whether a narrative
acknowledges live scholarly disagreement, or whether the chronology holds
together under scrutiny. A model can score 95 percent on general factual
accuracy and still produce a historical narrative that treats a contested
interpretation as settled, leans on secondary literature while ignoring
available primary sources, and quietly imports present-day assumptions into a
past context that didn't share them. None of that shows up in a
hallucination-rate benchmark. All of it would get a paper rejected by a
competent peer reviewer.
What the research actually shows
The clearest empirical precedent for this kind of evaluation comes from a 2026 comparative study that pitted trained historians against generative AI tools: ChatGPT-4, Claude, Gemini, and Perplexity, on parallel case studies in ancient history. The researchers found a consistent pattern: AI tools were genuinely strong at broad content discovery and thematic synthesis, the kind of work that benefits from fast pattern-matching across large bodies of text. But they were consistently weaker than human historians on source verification, chronological reasoning, provenance, and contextual interpretation, which are the parts of the discipline that depend on judgment rather than retrieval. The study's authors argued for a human-in-the-loop approach that treats AI as a research accelerant rather than a substitute for scholarly adjudication.
A separate line of research has approached the problem from
the model side rather than the historian's side, building benchmarks
specifically designed to test historical reasoning. That work has found that
even the strongest available models struggle with the kind of sophisticated
evidentiary reasoning that is routine for a trained historian, the weighing
conflicting accounts, recognizing genre and authorial bias, and reconstructing
plausible sequences of events from incomplete evidence. This is a useful
corrective to the assumption that model capability will simply catch up with
historical judgment as systems improve. The gap is not just about scale; it is
about a kind of reasoning these systems are not currently built to do well.
Not every recent contribution to this literature is critical of AI's role in historical work. Case studies of experienced historians integrating AI into their actual research workflows, mapping specific tasks to language models, to conceptual work the historian retains, and to conventional computational tools, that describe an approach where AI expands capacity without replacing judgment, provided the historian maintains rigorous cross-checking at every stage. That distinction matters for this piece: the argument here is not that AI is useless for historical work. It's that historians need a way to grade the output that general AI evaluation frameworks were never designed to provide.
Introducing the framework: HAIEF[3]
What follows is a two-part tool. The first is a practitioner-facing checklist, including twelve diagnostic questions meant to be run quickly against any piece of AI-generated historical research, the way you'd run a gut-check before deciding whether something is worth building on. The second is a more formal, citable rubric that breaks the same underlying concerns into eight scored criteria, useful when you need to document *why* a piece of AI output was accepted or rejected, not just whether it was.
The twelve questions:
1. Are all citations real?
2. Can every quotation be independently located?
3. Is the chronology internally consistent?
4. Does the AI distinguish primary from secondary sources?
5. Are archival collections correctly identified?
6. Are names, dates, and places cross-checked?
7. Does the narrative reflect historiographical debate?
8. Does the AI acknowledge uncertainty?
9. Does it introduce present-day assumptions into historical
contexts?
10. Can another historian reproduce the research?
11. Which claims require archival verification?
12. What percentage of the output would I actually trust
without checking?
The formal rubric maps the same territory onto eight
criteria: citation authenticity (do the cited works exist and actually support
the claim made), provenance (is every factual assertion traceable to a reliable
source), chronological accuracy (do the dates and sequences hold together
internally), primary-source fidelity (are quotations and archival references
accurate to the original), historiography (does the output acknowledge that
scholarly consensus is contested where it is contested), context (are events
read within their own historical setting rather than through present-day
assumptions), uncertainty (does the output distinguish established fact from
inference), and reproducibility (could another historian, working from the same
cited evidence, independently reach the same conclusions).
The value of running both versions side by side is that they serve different moments in the workflow. The twelve questions are what you ask while you're reading the output, in real time, deciding whether to keep going or start over. The eight-criterion rubric is what you fill out afterward, when you need a defensible record of why a source was trusted or discarded, which is useful for peer review, for training research assistants, or for documenting due diligence in a professional or editorial context.
Who this is for
Historians and academic researchers are the most obvious audience, and the stakes for them are the highest: a fabricated citation or an unacknowledged historiographical dispute that makes it into a submitted paper is a professional liability, not just an inconvenience.
Genealogists are arguably the highest-volume,
lowest-scrutiny users of AI-assisted historical research, and they are almost
entirely unaddressed by the academic literature on this problem. Family history
research routinely involves exactly the kind of claims, be it names, dates,
places, relationships, that AI systems handle with the least reliability, and
genealogists are less likely than academic historians to have institutional
habits of source verification already built in.
Journalists covering historical topics under deadline face a particular version of this problem: the twelve-question checklist doubles as a fast fact-checking pass when there isn't time for a full archival dig, and question twelve, what percentage would you trust without checking, is a useful discipline for anyone filing on a clock.
Educators like myself have perhaps the most durable use case. Detecting AI-generated text is a losing, ever-shifting battle. Teaching students to interrogate AI output using a framework like this one, and not asking "did a machine write this" but rather, "would this survive a historian's scrutiny", is a skill that remains useful regardless of how the underlying technology changes.
Closing
None of this should be mistaken for a claim that a checklist can do a historian's job. A framework like HAIEF can catch a fabricated citation, flag an inconsistent chronology, or surface a claim that needs archival verification before it gets treated as settled. What it cannot do is listen. Jan Burzlaff has argued that AI systems summarize but do not listen, reproduce but do not interpret. These tools achieve coherence while faltering at contradiction, and that historical writing is not simply an output to be optimized but a form of presence, a risk taken in the act of trying to make meaning where no prewritten frame will do. That distinction is worth sitting with rather than resolving. A verification framework can tell you whether the sourcing holds up. It cannot tell you whether the history has been understood. Those are different questions, and confusing them, mistaking a clean citation audit for genuine historical interpretation, may turn out to be the more consequential failure mode of all.
Sources
Burzlaff, Jan. "Fragments Not Prompts: Five Principles for Writing History in the Age of AI." Rethinking History (2026). https://doi.org/10.1080/13642529.2025.2546174.
Campbell,
Chris. "The Historian in the Age of AI." Transactions of
the Royal Historical Society (December 10, 2025).
https://www.cambridge.org/core/journals/transactions-of-the-royal-historical-society/article/historian-in-the-age-of-ai/37E3B742A2983DF4DAA2E38D48252F89.
"Generative
AI as a Historical Source: Source Criticism, Citation Integrity, and the Jagged
Frontier of Digital History." (March 20, 2026).
https://acnsci.org/journal/index.php/cte/article/view/1438.
Fox, Yaniv.
Using Generative AI in Historical Practice. Cambridge Elements.
Cambridge: Cambridge University Press, 2026.
https://www.cambridge.org/core/elements/abs/using-generative-ai-in-historical-practice/7C1392A6E9DBD379FAA42E2D16A6D45B.
Henriot,
Christian. "The AI-Augmented Research Process: A Historian's
Perspective." Preprint, arXiv, August 3, 2025.
https://arxiv.org/abs/2508.01779.
"HistoRAG:
Embedding Historical Methodology in Retrieval-Augmented Generation."
Preprint, arXiv, June 16, 2026. https://arxiv.org/abs/2606.18103.
"Can
LLMs Act as Historians?" Preprint, arXiv, April 27, 2026.
https://arxiv.org/abs/2604.24690.
Ruchniewicz,
Krzysztof. "AI in History Education: The American Historical Association's
Guidelines, the Views of Their Authors, and Their Reception within the
Historical Community." Studia Historica 62, no. 1 (2026).
https://doi.org/10.14746/sh.2026.62.1.003.
Solga,
Raymond S., and Mohammed J. Sarwar. "Evaluating Generative AI in
Historical Research: A Comparative Study on Identifying Primary Source Evidence
in Ancient History." AI & Antiquity 2, no. 1 (February 26,
2026). https://doi.org/10.64946/aiantiquity.v2i1.003.
Financial
Times. "PwC Withdraws AI-Generated Reports over Fabricated
Citations." July 29, 2026.
https://www.ft.com/content/7e149ac8-2ce2-4266-8940-192f9821b33c.
TechRadar
Pro. "A Major KPMG Report on AI Was Found to Be Chock-Full of AI
Hallucinations." June 12, 2026.
https://www.techradar.com/pro/a-major-kpmg-report-on-ai-was-found-to-be-chock-full-of-ai-hallucinations.
Axios.
"AI Chatbots Are Getting More Confident — and Sometimes More Wrong."
May 30, 2026.
https://www.axios.com/2026/05/30/ai-accuracy-chatbots-hallucinations.
No comments:
Post a Comment