by Alvin Blackshear | Historian & Researcher <ablackshear@gmail.com>
Historians have always worked at the boundary between evidence and
interpretation, testing claims against sources, weighing provenance, and living
comfortably with uncertainty. Generative artificial intelligence has entered
this terrain quickly and unevenly, offering historians an extraordinary tool
for discovery while introducing new categories of risk that the discipline has
not yet fully reckoned with. The question facing the field is not whether AI
belongs in historical research. It already does. The real question is how
historians verify what AI produces, and whether existing habits of source
criticism are sufficient for a technology that can fabricate footnotes, invent
quotations, and drift toward revisionist narratives when prompted to do so.
This article argues that AI can meaningfully accelerate historical research,
but it cannot replace professional source criticism, and it proposes the outline
of a repeatable verification protocol suited to the particular demands of
archival research, primary source authentication, citation validation,
chronology checks, and historiographical consistency.
The stakes of getting this right extend beyond any single article or
dissertation. History as a discipline depends on a shared confidence that its
claims can be traced back to something real, a document, an object, a testimony
that survived. When that chain of evidence is compromised, even quietly and even
with good intentions, the damage does not stay contained to one footnote. It
erodes the credibility of the discipline as a whole, at a moment when public
trust in expertise is already under strain. Historians who adopt AI tools
without adopting equally rigorous verification habits are not simply risking
their own reputations. They are testing the durability of a field that has
spent centuries building methods for distinguishing memory from evidence and
evidence from invention.
The Promise and the Ceiling
Large language models can summarize secondary literature, suggest starting
bibliographies, identify patterns across large text corpora, and help
historians locate materials they might otherwise have missed. Used well, these
tools compress hours of preliminary searching into minutes. The American
Historical Association's Ad Hoc Committee on Artificial Intelligence in History
Education, in its 2025 guiding principles, acknowledges this potential while
insisting on a firm boundary: generative AI can mimic some of the work
historians do, but this should never be mistaken for the work itself. The
committee's language is worth sitting with. AI produces texts, images, audio,
and video, not truths. It selects words and patterns from training data rather
than comprehending the past in its full complexity and contradiction. When a
model encounters a gap in its training, it does not pause to acknowledge
uncertainty. It fills the gap with something plausible sounding, a process the
field now calls hallucination.
For historians, hallucination is not an abstract technical curiosity. It is a
direct threat to the evidentiary basis of scholarship. A fabricated citation
that looks correctly formatted, attached to a real-sounding archive or journal,
can pass a cursory glance and enter a footnote before anyone notices the source
does not exist. The AHA's guiding principles are unambiguous on this point: any
reference generated by AI must be checked against the original before it is
used, and a reference included in a footnote without that check is never
acceptable practice, regardless of how the historian frames their working
process.
Why Automated Fact-Checking Falls Short
It is tempting to assume that the solution to AI hallucination is simply more
AI, in the form of automated fact-checking systems layered on top of generative
tools. Recent research complicates that assumption considerably. A 2026 study
led by Kelly Amaddio examined how people respond to automated fact-checkers
correcting misinformation, using a sample of more than thirteen hundred
participants exposed to a piece of gun control misinformation and a subsequent
AI-generated correction. The findings revealed a persistent and troubling
pattern: prior beliefs shaped both continued endorsement of the misinformation
after correction and perceptions of the fact-checking system's reliability.
When the fact-checker was described as only moderately accurate, close to the
accuracy levels people already associate with existing AI systems, prior
beliefs predicted whether the correction actually changed minds. Even when
participants were told the system was highly accurate, the effect of prior
belief did not disappear, only lessened.
This matters for historians because it demonstrates that automated verification
does not operate in a neutral space. It is filtered through the same human
biases that shape how any evidence is received, and it can be defeated by
claims that already align with what someone wants to believe. Historical
scholarship, particularly on contested and politically charged topics, is
especially vulnerable to this dynamic. A fact-checking layer bolted onto a
language model cannot substitute for the disciplinary training that teaches
historians to interrogate their own assumptions as rigorously as they interrogate
a source.
Revisionism as a Distinct Problem
Contested historical claims present a challenge that generic fact-checking
frameworks were never designed to solve, because the disagreement is not always
about a discrete, verifiable fact but about interpretation, framing, and the
political stakes attached to memory. A 2026 study by Francesco Ortu and
collaborators addresses this directly. The researchers built a benchmark called
Historical Misinfo, drawing on five hundred contested historical events from
forty five countries, each paired with a documented factual narrative and a
documented revisionist counter-narrative. They then tested how large language
models responded under neutral prompting compared with prompting that
explicitly requested the revisionist version of events.
The results are instructive. Under neutral conditions, most models leaned
closer to the factual reference. But when users explicitly requested a
revisionist framing, every model tested showed a sharp increase in revisionist
output, with little resistance or self-correction. This finding should give
historians real pause. It suggests that a model's apparent reliability on a
given topic says little about how that same model will behave when a user,
deliberately or not, nudges it toward a distorted account. For historians
working on genocide, colonialism, contested territorial claims, or any subject
where competing national narratives exist, this is not a marginal concern. It
is central to responsible use.
What Verification-First Research Looks Like Elsewhere
Historians are not alone in confronting this problem. Related work emerging
from adjacent fields offers useful models for how disciplines might restructure
their relationship with AI around verification rather than convenience. Lei
You, Lele Cao, and Iryna Gurevych have argued that AI-assisted peer review in
the sciences should be verification-first rather than review-mimicking,
proposing that AI tools function as adversarial auditors generating checkable
evidence rather than as systems that simply predict what a human reviewer might
conclude. Belinda Mo extends this argument to the broader landscape of
autonomous AI research agents, warning of a widening verification gap between
what AI systems can produce and what any human can realistically check, and
calling for verification infrastructure built around observable workflows and
clear attribution.
A study by Raymond Solga and Mohammed Sarwar offers perhaps the closest
parallel to the historian's own situation. Comparing traditional historical
methods against several generative AI tools in identifying and validating
primary sources in ancient history, the researchers found that AI performed
well at broad content discovery and thematic synthesis but struggled
significantly with genre boundaries, provenance transparency, and the kind of
contextual interpretation that trained historians bring almost instinctively to
a source. Their proposed solution, a human-in-the-loop framework built around
model pluralism and provenance-first protocols, maps closely onto what
historians already know from decades of archival practice, simply extended to
include AI as one input among many that must be checked rather than trusted.
There is also a temporal dimension worth naming. Carlo Iacono's recent audit of
AI-assisted research output found that the models cited in published work are
frequently outdated by the time readers encounter them, sometimes by close to a
year, with model versions superseded and behaviors shifting in ways that are rarely
documented. For historians, this is a reminder that any claim about what a
given AI tool can or cannot do should be dated and treated as provisional
rather than as a permanent property of the technology.
Toward a Repeatable Protocol
Drawing these threads together, a workable verification protocol for historians
should rest on several commitments. Every AI-generated citation must be checked
against the primary source or an authoritative secondary source before it
appears in any written work, with no exceptions for citations that merely look
plausible. Chronology and provenance claims generated by AI require independent
confirmation, since these are precisely the areas where current tools show the
weakest performance. Contested or politically charged claims deserve heightened
scrutiny, given the demonstrated tendency of models to drift toward whatever
framing a prompt implies. Verification work should be documented as it happens,
noting which tool was used, when, and what was checked, both for the historian's
own accountability and because the tools themselves change quickly enough that
undocumented claims about their reliability age poorly. And throughout, the
historian's own expertise must precede rather than follow AI use, since
evaluating a model's output well requires the same depth of subject knowledge
that historians have always needed to evaluate any source.
None of this diminishes the real value AI offers historians willing to use it
carefully. It does insist that the discipline's oldest habit, treating every
claim as something to be verified rather than assumed, remains the right
response to a genuinely new technology. Generative AI can accelerate the search
for evidence. It cannot substitute for the judgment that turns evidence into
history.
There is a version of this moment that history has lived through before, in
smaller ways, each time a new technology promised to change how the past is
studied. Microfilm, digitization, keyword-searchable databases, all reshaped
the pace and reach of archival work without ever touching the core discipline
of asking what a source can and cannot tell us. Generative AI is different in
degree, not necessarily in kind. It compresses years of searching into minutes,
but it also compresses the distance between plausible and true, and that
compression is precisely where a historian's training earns its keep. The AHA's
committee put it plainly: there are no shortcuts to expertise, and the skills
that let a scholar recognize a hallucinated citation or a subtly revisionist
framing can only be built through the same sustained engagement with sources
that has always defined the craft. A historian who verifies rigorously will
find in AI a genuine research partner. A historian who does not will find,
sooner or later, that the past they have written no longer holds.
______________________________
Sources
American Historical Association, Ad Hoc Committee on Artificial Intelligence in
History Education. "Guiding Principles for Artificial Intelligence in
History Education." Approved by AHA Council, July 29, 2025.
https://www.historians.org/resource/guiding-principles-for-artificial-intelligence-in-history-education/
Amaddio, Kelly M., et al. "Prior Beliefs and Automated Fact Checking:
Limits on the Accuracy of AI Verification." February 5, 2026. https://doi.org/10.1371/journal.pone.0342332
Iacono, Carlo. "The Half-Lives of Generative-AI Evidence." July 27,
2026. https://doi.org/10.48550/arXiv.2607.24032
Mo, Belinda. "The Age of AI Agents Demands a New Scientific Paradigm to
Sustain Trustworthy Science." June 18, 2026.
https://doi.org/10.48550/arXiv.2607.26064
Ortu, Francesco, et al. "Preserving Historical Truth: Detecting Historical
Revisionism in Large Language Models." February 22, 2026.
https://doi.org/10.48550/arXiv.2602.17433
Solga, Raymond S., and Mohammed J. Sarwar. "Evaluating Generative AI in
Historical Research." AI & Antiquity (2026).
https://doi.org/10.64946/aiantiquity.v2i1.003
You, Lei, Lele Cao, and Iryna Gurevych. "Preventing the Collapse of Peer
Review Requires Verification-First AI." February 12, 2026.
https://doi.org/10.48550/arXiv.2601.16909




