Sunday, August 09, 2026

The Fastest Man in the World Has a Message About Slowing Down


What Usain Bolt's Jardiance Commercial Can Teach Us About CKD

Usain Bolt built his entire identity on speed. In 2009, he covered 100 meters in 9.58 seconds, a world record that still stands nearly two decades later. He is, by any measure, the fastest human being who has ever lived. So it's a little disorienting to see him on American television these days, not sprinting, but talking about the value of slowing down.

The commercial is for Jardiance, called "Slow Can Be Good Too," and it deliberately turns Bolt's public identity upside down. Speed made him famous. But the ad isn't selling speed. It's making the case that slowing something down, specifically the progression of a chronic disease, can be its own kind of victory. That contrast is what makes the advertising work. Our brains are wired to notice contradictions, and pairing the world's fastest man with a message about deceleration creates exactly the kind of cognitive friction that makes people pay attention.

Buried inside the ad is a simple line that matters far more than the celebrity casting: "When was the last time you had your kidneys checked?" That question is worth pausing on, because most people genuinely don't know the answer.

The Commercial Contains a More Important Message

It's tempting to read the ad as just another celebrity pharmaceutical spot. But the word that matters here isn't Jardiance. It's progression. Chronic kidney disease, or CKD, is a condition that typically develops slowly, often over years, and it rarely announces itself with dramatic symptoms in its earlier stages. A person can lose a meaningful share of kidney function and feel completely normal.

That quiet progression is precisely why CKD is such an underrecognized problem. According to the Centers for Disease Control and Prevention's most recent national estimates, roughly 14 percent of U.S. adults, or about 37 million people, have chronic kidney disease [1]. Far more striking is how many of them don't know it. The CDC has found that about 9 in 10 adults with CKD are unaware they have it [1]. That is not a small awareness gap. It is nearly the entire affected population moving through the disease without a diagnosis, often until it has advanced considerably.

What Exactly Is CKD?

Think of the kidneys as filters. Roughly the size of fists, sitting toward the back of the abdomen, they filter waste and excess fluid out of the blood, help regulate blood pressure, and keep the body's chemistry in balance. Chronic kidney disease means that filtering ability has been damaged and is diminished, and that the damage has persisted for at least three months.

CKD is typically staged using two pieces of information rather than one, which is part of why it's harder to explain than, say, a cholesterol number. As kidney function declines, the body loses its ability to clear waste products and regulate fluid balance, which over time raises the risk of cardiovascular disease, anemia, bone problems, and eventually kidney failure requiring dialysis or transplantation. The disease doesn't move in a straight line for everyone, but for many people it is progressive, meaning it tends to move in one direction without intervention.

The Diabetes Connection

Type 2 diabetes is the single biggest driver of CKD in the United States, and the relationship between the two is well established. Persistently elevated blood glucose damages the small blood vessels inside the kidneys' filtering units over time, gradually reducing their ability to do their job. The CDC's newest national data illustrate just how tightly linked these conditions are: an estimated 38 percent of adults with diabetes and 21 percent of adults with high blood pressure were estimated to have CKD [1]. Even among people with prediabetes, more than 1 in 10 were estimated to have CKD already [1].

This is where A1C, the blood test that reflects average blood sugar over roughly the previous three months, enters the picture. A1C doesn't measure kidney function directly, but it is one of the best available signals of the metabolic stress that, over years, can damage the kidneys. Managing blood sugar isn't just about avoiding diabetes complications in the abstract. It's directly tied to kidney protection.

Do You Know Your Numbers?

This is the practical heart of the matter, and it's worth being specific about three separate numbers, because they measure different things and people often conflate them.

A1C reflects longer-term blood glucose control and is used to screen for and monitor diabetes and prediabetes.

eGFR, or estimated glomerular filtration rate, is calculated from a blood creatinine measurement and estimates how efficiently the kidneys are filtering blood. A lower eGFR generally indicates reduced kidney function.

uACR, the urine albumin-to-creatinine ratio, looks for albumin, a protein, leaking into the urine. Healthy kidneys generally keep albumin in the blood rather than letting it pass into urine, so its presence can be an early sign of kidney damage, sometimes appearing before eGFR changes at all.

Current clinical guidance reflects how important it is to check both eGFR and uACR together rather than relying on either alone. The American Diabetes Association's 2026 Standards of Care recommend assessing kidney function with both a random urine albumin-to-creatinine ratio and estimated glomerular filtration rate at least annually in people with type 1 diabetes of five or more years' duration, and in all people with type 2 diabetes regardless of treatment [2]. The rationale is straightforward: albuminuria and eGFR each carry independent information about risk, and at any given eGFR, the degree of albuminuria is associated with the risk of cardiovascular disease, CKD progression, and death [2].

So three questions worth bringing to your next checkup: Do you know your A1C? Do you know your eGFR? Have you had a urine albumin test? For most adults without diabetes or hypertension, routine kidney screening isn't automatically part of an annual physical, so this is a conversation to have directly with a clinician, especially for anyone with diabetes, high blood pressure, cardiovascular disease, or a family history of kidney disease.

What Jardiance Actually Has to Do With It

Jardiance, known generically as empagliflozin, belongs to a drug class called SGLT2 inhibitors. It was originally approved for type 2 diabetes and later for certain forms of heart failure. In September 2023, the FDA expanded its approval again, this time specifically to reduce the risk of sustained decline in eGFR, end-stage kidney disease, cardiovascular death, and hospitalization in adults with chronic kidney disease at risk of progression [3]. That approval was based on the EMPA-KIDNEY trial, a large, dedicated CKD study in which empagliflozin reduced the composite risk of kidney disease progression or cardiovascular death by 28 percent relative to placebo [4].

It's worth being precise about what that means. Jardiance is not a treatment for everyone with CKD, and it isn't a substitute for blood pressure control, glucose management, or other standard therapies. The manufacturer notes it is not recommended for CKD in patients with polycystic kidney disease or those requiring significant immunosuppressive therapy for kidney disease [5], since it isn't expected to help those populations. Like any prescription medicine, it carries its own risks and side effects, including dehydration and genital or urinary tract infections, and whether it's appropriate is a decision for a patient and their physician, not a conclusion to draw from a 30-second ad.

Back to Bolt

For most of his career, Usain Bolt spent every fraction of a second trying to go faster. His latest television role asks viewers to consider the opposite proposition: that sometimes slowing something down is itself a form of winning. Perhaps the most useful thing about the commercial isn't the drug being advertised at all. It's the question it leaves behind, the one easy to overlook amid the celebrity wattage: Do you know how well your kidneys are working? For nearly 9 in 10 people currently living with CKD, the honest answer is still no. That's the number worth changing.


Notes

1.     Centers for Disease Control and Prevention. Chronic Kidney Disease in the United States (updated March 2026). CS 363495-A.

2.     American Diabetes Association. 11. Chronic Kidney Disease and Risk Management: Standards of Care in Diabetes—2026. Diabetes Care, 2026.

3.     U.S. Food and Drug Administration approval announcement, via Eli Lilly and Company / Boehringer Ingelheim, "US FDA approves Jardiance for the treatment of adults with chronic kidney disease," September 22, 2023.

4.     Pharmacy Times. "FDA Approves Empagliflozin for Adults With Chronic Kidney Disease at Risk of Progression."

5.     Eli Lilly and Company. "US FDA approves Jardiance for the treatment of adults with chronic kidney disease" (full release, contraindication details), September 22, 2023.

6.     American Diabetes Association Professional Practice Committee. Screening for Chronic Kidney Disease in People with Diabetes (visual guide), April 2026.

7.     National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK). Kidney Disease Statistics for the United States.

8.     KDIGO. 2026 Clinical Practice Guideline for Diabetes and CKD (public review draft, March 2026).

9.     HCPLive. "Empagliflozin (Jardiance) Receives CKD Approval from US FDA."

10.Boehringer Ingelheim / Eli Lilly patient prescribing information, Jardiance Chronic Kidney Disease page.

Lessons from eDiscovery for RAG: Building Trustworthy Retrieval Systems


A follow-up to "What is RAG? Retrieval-Augmented Generation Systems with Archival Sources"

by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

Retrieval Is Not a New Problem

In my last article, I tested Retrieval-Augmented Generation (RAG) systems against archival sources and found that even well-engineered pipelines struggle with the basic question of whether they found the right material before they ever tried to answer a question. That struggle is not new. It has a name, a case history, and a professional discipline built around it: eDiscovery.

For three decades, litigators and their technical partners have grappled with a problem that RAG developers are only now rediscovering: how do you find the right documents inside an enormous, messy, heterogeneous collection, and how do you prove, after the fact, that you found them? In litigation, the collection is made of emails, contracts, scanned memos, spreadsheets, and chat logs, scattered across custodians and file formats. In RAG, the collection is made of chunked documents, embeddings, and vector indices. The materials differ. The underlying task, sorting a haystack for a small number of relevant needles under conditions of uncertainty, is the same.

This matters because the AI field tends to treat retrieval as a purely technical challenge: pick a better embedding model, tune the reranker, adjust top-k. Wu and colleagues' recent survey of RAG architectures catalogs exactly this kind of technical churn across retrieval pipelines, hybrid search strategies, reranking layers, and evaluation methods (Wu et al., 2026). What that survey does not resolve, and what most RAG engineering discussions skip entirely, is the governance question that eDiscovery has spent decades answering: how do you know your retrieval process is good enough to trust, and how do you demonstrate that to someone who was not in the room when you built it?

Precision versus Recall, Revisited

Every eDiscovery practitioner learns the same lesson early: precision and recall pull against each other, and the cost of getting the balance wrong is not abstract. Set the net too wide and reviewers drown in irrelevant documents, burning hours and budget on material that never should have surfaced. Set the net too narrow and you risk missing the one email that determines the outcome of a case. Cormack and Grossman's foundational work on Continuous Active Learning showed that recall-oriented review, iteratively refined through relevance feedback, could achieve high recall with a fraction of the manual review effort that exhaustive linear review required (Cormack and Grossman, 2015). That paper reframed retrieval not as a one-time query but as a process, one that improves through repeated rounds of human judgment and system adjustment.

RAG systems face a structurally identical tradeoff, just with different vocabulary. Top-k retrieval determines how many chunks get pulled before generation begins. Reranking determines which of those chunks survive to reach the model's context window. Filtering determines which sources are excluded before retrieval even starts. Elkiran and Rasheed's comparative study of retriever-reranker pairings makes the tradeoff explicit: different combinations of retrieval and reranking strategies produce measurably different quality and efficiency profiles, and no single configuration dominates across all use cases (Elkiran and Rasheed, 2026). That is precision versus recall wearing a different uniform. A RAG system tuned for maximum precision will answer confidently and sometimes wrongly, having missed the one chunk that would have changed the answer. A system tuned for recall will surface enough context to be safe, but risks diluting the model's attention with irrelevant material, an effect not unlike a review team wading through thousands of non-responsive documents in search of the few that matter.

The TREC Total Recall track exists precisely because "how much recall is enough" cannot be answered by intuition. It has to be measured, with defined test collections, defined relevance judgments, and reproducible scoring. RAG evaluation is only beginning to build equivalent infrastructure.

Chain of Custody versus Provenance

A litigator asks where a document came from. A historian asks about its provenance. A RAG developer asks which chunk produced a given answer. These are the same question asked by three different professions, and only one of those professions has built a rigorous, court-tested framework for answering it.

Chain of custody in eDiscovery is not paperwork for its own sake. It exists because a document's evidentiary value depends on being able to trace it back to its source without gaps: who collected it, when, from which custodian, and whether it was altered along the way. Knollmeyer and colleagues' review of RAG evaluation dimensions identifies citation quality and context relevance as core axes on which modern systems are judged (Knollmeyer et al., 2026), which is encouraging, but citation quality in most RAG systems today means little more than "a source was attached." It rarely means the source was verified, the retrieval path was logged, or the chunk boundaries were preserved in a way that lets a reviewer reconstruct exactly what the model saw.

This gap matters more as RAG systems move into higher-stakes domains. Zhang and colleagues' comparative study of retrieval strategies in clinical note extraction found that traceability between retrieved evidence and generated output is essential in a domain where an ungrounded claim carries real consequences (Zhang et al., 2026). The parallel to legal discovery is not a stretch. A clinician relying on a RAG system's summary of a patient's history needs the same assurance a litigator needs from a document review platform: an unbroken, inspectable link between the source material and the conclusion drawn from it.

Toward Defensible AI Retrieval

Legal discovery imposes four requirements on retrieval decisions: they must be explainable, reproducible, documented, and defensible under adversarial scrutiny. Those four words are worth sitting with, because RAG evaluation is converging on exactly the same standard, even if the field has not yet named it that way.

Tan and colleagues' work on provably robust aggregation addresses a version of the defensibility problem directly, introducing mathematically grounded methods for making retrieval resistant to corrupted or adversarial evidence (Tan et al., 2026). That is the RAG equivalent of a chain-of-custody challenge in court: what happens when the retrieved evidence itself cannot be trusted, whether through poisoning, duplication, or simple corpus noise? Cho and Lee's redundancy-aware evaluation framework tackles an adjacent problem that will feel familiar to any eDiscovery practitioner: high-similarity corpora, full of near-duplicate documents, distort both retrieval metrics and reviewer judgment if not handled explicitly (Cho and Lee, 2026). Legal collections are famously redundant, forwarded emails, cc'd threads, and near-identical drafts. RAG corpora, especially those built from scraped or aggregated sources, are no different, and the evaluation methods built for one domain transfer cleanly to the other.

I want to propose a working term for what this convergence is pointing toward: Defensible AI Retrieval. A defensible retrieval system is not simply one that performs well on a benchmark. It is one whose outputs can be explained after the fact, whose retrieval logic can be reproduced by a third party, whose decisions are documented well enough to survive scrutiny, and whose failure modes are known rather than discovered by accident. The Sedona Conference's TAR Case Law Primer spends hundreds of pages working through exactly this standard for legal technology, and its core insight, that defensibility is established through process documentation and validation, not through claims of accuracy alone, applies almost without modification to AI retrieval systems (Sedona Conference, 2023).

Quality Assurance: The Discipline RAG Has Not Yet Built

If there is one area where eDiscovery is unambiguously ahead of RAG, it is quality assurance. Legal review workflows are built around statistical validation from the outset: random sampling of the review population, second-level review of a subset of coded documents, dedicated privilege review passes, and formal error correction protocols. None of this is optional or aspirational. It is baked into the workflow because courts require it and because the cost of an undetected error, a privileged document produced by mistake, a responsive document never found, is high enough to justify the overhead.

RAG evaluation has started to build analogous practices, but the culture around them is still immature. Hallucination detection, groundedness scoring, and faithfulness metrics are now common terms in RAG papers. Knollmeyer and colleagues' framework treats faithfulness and context relevance as first-class evaluation dimensions alongside more traditional retrieval accuracy measures (Knollmeyer et al., 2026), and Sotic and Kamps' study of information-seeking behavior in RAG systems adds something the retrieval metrics literature often misses: how users actually interact with retrieved evidence, what they trust, what they miss, and where their mental models diverge from what the system actually did (Sotic and Kamps, 2026). That user-facing lens is closer to what eDiscovery calls second-level review, a human check on whether the system's output actually holds up.

What is still missing from most RAG evaluation pipelines is the statistical rigor that eDiscovery treats as baseline. Cormack and Grossman's 2016 paper on engineering quality and reliability in technology-assisted review argued that retrieval systems should be engineered for measurable quality, not tuned by intuition and then trusted (Cormack and Grossman, 2016). Random sampling with confidence intervals, documented error rates, and formal validation protocols are standard in eDiscovery and still rare in RAG deployment. A scoring rubric borrowed from legal QA might look like this for a RAG system under evaluation: sample a statistically significant subset of generated answers, score each on groundedness (is every claim traceable to a retrieved chunk), faithfulness (does the answer avoid contradicting its sources), completeness (did retrieval surface the material a human reviewer would consider necessary), and false-negative risk (what was missed, and how costly would that omission be in context). Reporting these as point estimates with confidence intervals, rather than as a single aggregate accuracy score, would bring RAG evaluation much closer to the standard eDiscovery already treats as routine.

Ten Lessons RAG Developers Should Borrow

1. Retrieval must be measurable. Intuition is not a validation strategy.

2. Every answer needs provenance. A citation that cannot be traced to its source chunk is not a citation.

3. Retrieval decisions should be explainable, not just accurate.

4. Sampling is mandatory. No system should be trusted on the basis of spot checks alone.

5. Human validation never disappears. Automation reduces review burden; it does not eliminate the need for oversight.

6. Metadata matters. Knowing where a chunk came from, and how it was processed, is as important as its content.

7. False negatives are often more dangerous than false positives. A missed document, or a missed fact, can be more damaging than an irrelevant one.

8. Evaluation should be continuous, not a one-time benchmark run before launch.

9. Retrieval quality should be statistically measured, with error rates and confidence intervals, not asserted.

10. Trust is earned through transparency, not claimed through marketing.

Conclusion

AI did not invent the challenge of trustworthy retrieval. Long before vector databases and large language models existed, courts demanded retrieval systems capable of finding, documenting, defending, and reproducing evidence, and an entire professional discipline grew up around meeting that demand. Retrieval-Augmented Generation (RAG) represents a genuine technological advance. It is not, however, an entirely new discipline, and treating it as one means reinventing quality assurance practices, defensibility standards, and precision-recall tradeoffs that the legal profession has already spent decades refining under adversarial, high-stakes conditions.

The future of trustworthy AI retrieval may depend less on inventing new evaluation methods from scratch than on rediscovering, and adapting, the hard-earned lessons of eDiscovery. The tools will keep changing. The underlying discipline, measurable, provenance-driven, statistically validated, and honest about its failure modes, does not need to be reinvented. It needs to be borrowed.

---

Notes

Cho, H., & Lee, J.-Y. (2026). RARE: Redundancy-aware retrieval evaluation framework for high-similarity corpora. ACL 2026. https://aclanthology.org/2026.acl-long.923

Cormack, G. V., & Grossman, M. R. (2015). Autonomy and reliability of continuous active learning for technology-assisted review. https://arxiv.org/abs/1504.06868

Cormack, G. V., & Grossman, M. R. (2016). Engineering quality and reliability in technology-assisted review. Proceedings of SIGIR 2016. https://plg.uwaterloo.ca/~gvcormac/PAPERSx.html

Elkiran, H., & Rasheed, J. (2026). Evaluating retriever reranker pairings in RAG based on quality and efficiency trade-offs. Discover Computing. https://doi.org/10.1007/s10791-026-10156-3

Knollmeyer, S., et al. (2026). Evaluating retrieval augmented generation: A comprehensive review of evaluation dimensions, question types, and application. SN Computer Science. https://link.springer.com/article/10.1007/s42979-026-05134-x

Redgrave LLP (2026). "Putting Gen AI to the Test: A Document Review Accuracy Study. Relativity aiR for Review vs. Active Learning". https://resources.relativity.com/document-review-accuracy-analysis-study-lp.html

The Sedona Conference. (2023). The Sedona Conference TAR case law primer (2nd ed.). https://www.thesedonaconference.org/node/10365

Sotic, B. N., & Kamps, J. (2026). Information seeking behavior in LLM-based RAG: Mental models and missing information. Proceedings of ACM SIGIR 2026. https://doi.org/10.1145/3805712.3808537

Tan, X., et al. (2026). PRA-RAG: Provably robust aggregation in retrieval-augmented generation against retrieval corruption. Findings of ACL 2026. https://aclanthology.org/2026.findings-acl.1794

Wu, S., et al. (2026). Retrieval-augmented generation for natural language processing: A survey. Artificial Intelligence Review. https://link.springer.com/article/10.1007/s10462-026-11605-7

Zhang, H., et al. (2026). Optimising clinical information extraction: A comparative study of retrieval-augmented generation techniques in clinical notes. Journal of Biomedical Informatics. https://www.sciencedirect.com/science/article/pii/S1532046426000778

Saturday, August 08, 2026

Shrink Solves Murder - REVIEW


by Philippa Perry

When a body is found near Beachy Head, the police chalk it up to suicide — a tragic but not uncommon end in these parts. But local psychotherapist Patricia Phillips isn’t convinced. The victim? Her three o’clock patient, Henry Clayton. The cause of death is supposedly self-inflicted. Yet Pat can’t shake the belief that someone wanted Henry Clayton dead. She spends her working life listening to histories and secrets, and she has a nose for when a story doesn’t quite ring true. Drawn from the therapy room to the crime scene, Pat begins to notice what others appear to overlook. At her side is her best friend Prichard — a home-brewer of fearsome, stomach-turning concoctions, an excellent cook, and a man who seems to get along with everyone. Which makes him useful for infiltrating village life. As Pat and Prichard look beneath the village’s thin veneer of normality — one that barely conceals its appetites — they discover a killer hiding in plain sight.

Shrink Solves Murder is a warm, witty, and perceptive crime caper from the nation’s favorite therapist.

Penguin Random House
ISBN-13: 978-1529155327

#PenguinRandomHouse

#ShrinkSolvesMurder


What is RAG ?


What Is Retrieval-Augmented Generation (RAG), and Is It a Groundbreaking Way of Interacting with Historical Documents?

by Alvin Blackshear | Historian & Researcher

For as long as archives have existed, the historian's craft has depended on a single, stubborn skill: finding the right document among thousands of wrong ones. Card catalogs gave way to keyword search engines, which gave way to digitized finding aids, yet the underlying task never changed. A researcher formulates a question, searches a collection using the vocabulary available at the time, and hopes the terms they chose overlap with the terms an archivist or a scanner happened to preserve. Retrieval-Augmented Generation, or RAG, is the first technology in this long lineage that promises to close the gap between what a historian actually wants to know and what a search box is capable of returning. Whether it fully delivers on that promise, and under what conditions, is the question this article sets out to examine.

Why Historians Need RAG Rather Than a Generic Chatbot

It is tempting to assume that any large language model, given enough context, can answer historical questions competently. This assumption is mistaken, and understanding why is essential to understanding what RAG actually contributes.

A generic LLM answers from its training data, a fixed and often outdated snapshot of text scraped from the internet. It has no reliable way to point to a specific archival folder, newspaper issue, or census page as the source of its claim, and it has every incentive, structurally speaking, to produce a fluent and confident answer even when no such source exists in its memory. For a historian, whose entire discipline rests on the traceability of a claim back to a primary document, this is disqualifying. An answer without a citation is not history. It is a plausible-sounding guess.

RAG changes the architecture in a specific way. Instead of relying solely on what a model memorized during training, a RAG system retrieves relevant passages from an external, defined collection, be it a digitized newspaper archive, a set of scanned letters, or a database of military records, and then asks the language model to generate an answer grounded in those retrieved passages. The research group behind the paper "Improving Access to Historical Archives with Real-time RAG-based Systems," led by Stergios Konstantinidis and colleagues, demonstrates this architecture at scale, applying it to a Swiss newspaper archive of roughly half a million segments spanning nearly two and a half centuries. Their system pairs a semantic retrieval and reranking pipeline with grounded answer generation, which means the model's output is tethered to specific archival text rather than floating free of any documentary anchor. This is the essential difference that makes RAG, rather than an ordinary chatbot, the appropriate tool for archival research.

How RAG Differs from Conventional Archival Search

Traditional archival search tools, including many of the platforms libraries and museums have relied on for decades, are built on lexical matching. A researcher types a term, and the system returns documents containing that exact term or a close variant. This works reasonably well when the researcher already knows the vocabulary used by the source material, but historical vocabulary shifts constantly. Place names change. Institutional terminology evolves. Spelling was inconsistent long before anyone standardized it. A search for a term as it is understood today may simply miss a document that describes the same event using nineteenth-century phrasing.

RAG systems address this through semantic retrieval, which represents both the query and the archival text as vectors in a shared conceptual space, allowing the system to surface documents that are meaningfully related even when the wording differs substantially. A master's thesis from the University of Twente, authored by Gauri Bhatnagar and focused on the CLARIAH Media Suite, an audiovisual heritage collection, tested exactly this proposition with real humanities scholars and archival professionals. The study found that RAG genuinely helps with cross-lingual access and exploratory, interpretive queries that lexical search struggles with, but it also surfaced a critical limitation: the system performed well on factual retrieval and named-entity identification, while struggling with interpretive reasoning and the kind of complex synthesis that sits at the heart of humanities scholarship. In other words, RAG can help a historian find the needle, but it cannot yet be trusted to explain what the needle means within a broader historiographical argument.

A separate, more industry-facing piece published by Veridian Software, a company that builds digital newspaper and archive platforms, frames RAG in more practical terms for institutions considering adoption, describing it as a natural extension of the search tools archives already maintain rather than a wholesale replacement of them.

Common Failure Modes in Archival RAG Systems

No technology arrives without its own characteristic failures, and archival RAG is no exception. Several recurring problems deserve explicit attention from anyone applying these tools to historical research.

OCR errors.  Most archival text exists as scanned images that have been converted to machine-readable text through optical character recognition, a process that is notoriously imperfect on older, damaged, or low-quality print. The Konstantinidis paper directly measures this problem and reports that an LLM-based refinement step reduced character error rates by up to roughly forty-five percent and word error rates by close to sixty-one percent, compared to the raw OCR output. That is a substantial improvement, but it is not perfection, and any residual error can silently distort a retrieved passage or an answer built on top of it.

Incomplete retrieval.  Even the best semantic search will not surface every relevant document, particularly when a collection is heterogeneous or when relevant material uses unusual phrasing, an obscure genre, or a minority language. A RAG system that returns five confident-sounding sources may be quietly ignoring a sixth that contradicts them.

Provenance loss.  When a language model synthesizes multiple retrieved passages into a single fluent answer, it is easy for the boundary between one source and another to blur. A historian needs to know precisely which claim came from which document, on which page, from which collection. A system that merges sources into an undifferentiated paragraph has failed a basic archival standard even if every individual fact happens to be accurate.

Temporal confusion.  Archives frequently contain material spanning centuries, and a retrieval system optimized purely for topical similarity has no inherent understanding of chronological sequence. It is entirely possible for a RAG system to present an 1850s account and a 1950s retrospective as though they were contemporaneous, a mistake that would mislead any historian relying on the system's framing rather than checking dates independently.

Fabricated citations.  Perhaps the most consequential failure mode is the language model's tendency to generate a citation that looks correct in form, a plausible archive box number or a newspaper date, but does not correspond to any actual retrieved document. Even well-grounded RAG systems are not fully immune to this, especially when a query has no good answer in the underlying collection and the model is inclined to produce something anyway rather than admit the gap.

The Missing Piece: A Historian's Evaluation Standard

Here is where the existing research reveals a striking gap. Every technical paper reviewed for this article evaluates RAG systems using measures native to information retrieval and natural language processing: retrieval precision, recall, embedding quality, OCR accuracy, system latency, and hallucination rate. A recent survey of RAG evaluation methods by Aoran Gan catalogs this landscape comprehensively, and a companion technical guide from Toloka lays out practical metrics for groundedness and answer faithfulness. These are legitimate and necessary measures. But none of them ask the questions a professional historian actually asks when assessing whether a piece of evidence is trustworthy.

Did the system retrieve the most authoritative version of a source, or merely the most textually similar one? Did it note when retrieved evidence contradicted itself, or did it quietly favor the passage that supported a cleaner narrative? Did it preserve the chain of custody and provenance for each claim? Did it distinguish a primary account from a later secondary interpretation of that account? Did it respect the order of events rather than collapsing decades into an undifferentiated blend? Did it accurately synthesize material drawn from separate collections without conflating them? Did it acknowledge uncertainty where the evidence was genuinely ambiguous, rather than resolving that ambiguity for the sake of a tidy answer? And did it surface evidence that was difficult to find but still relevant, rather than defaulting to whatever ranked highest by similarity score?

This gap between technical evaluation and historiographical evaluation is the foundation for what can be called a Historian's AI Evaluation Framework, or HAIEF. Rather than asking whether a RAG system is efficient, HAIEF asks whether it is trustworthy by the standards the historical profession has used for generations: provenance, contextualization, chronology, corroboration, and transparency about uncertainty. A system might score extremely well on NDCG or answer-correctness metrics, the kind reported by Konstantinidis and colleagues, whose reranking pipeline improved NDCG@10 from roughly sixty-six percent to eighty-seven percent, and still fail a historian's basic sniff test if it cannot show its provenance trail or if it silently smooths over contradictory testimony.

Toward a Historian's Testing Framework

A workable HAIEF would need to translate these five values into repeatable, scorable tests. Provenance could be scored by checking whether every claim in a generated answer links back to a specific, verifiable source with collection, folder, and page-level detail. Contextualization could be scored by asking whether the system situates a document within its original archival and historical setting rather than presenting it as free-floating text. Chronology could be tested with queries that require ordering events correctly across a span of decades or centuries. Corroboration could be tested by deliberately including contradictory sources in a test collection and observing whether the system flags the disagreement or resolves it artificially. Transparency about uncertainty could be measured by presenting queries with genuinely unresolved historical questions and checking whether the system hedges appropriately rather than manufacturing false confidence.

Benchmark Collections and a Comparative Rubric

Testing such a framework requires benchmark collections that reflect the actual diversity of archival material historians work with: personal letters and correspondence, census and vital records, digitized newspapers, military service and pension records, oral history transcripts, and photograph collections with accompanying metadata. Each of these genres carries its own provenance conventions and its own failure risks, and a system that performs well on newspapers may perform poorly on oral histories, where nuance, tone, and interviewer influence complicate straightforward retrieval.

A repeatable scoring rubric applying HAIEF criteria could then be used to compare general-purpose assistants such as ChatGPT, Claude, and Gemini against dedicated archival RAG systems purpose-built for a single institution's holdings. Such a comparison would likely show that general-purpose assistants excel at fluent synthesis but lag on provenance transparency, while dedicated archival systems, closer in spirit to the Konstantinidis pipeline, perform better on traceability but may still struggle with the interpretive reasoning that Bhatnagar's thesis identified as a persistent weakness across the board.

Recommendations for Archives Building AI-Assisted Discovery Tools

Institutions developing these tools should treat historiographical evaluation as a first-class requirement rather than an afterthought bolted on after a system already works technically. This means involving historians and archivists directly in system design, not merely as end-user testers after launch. It means building citation and provenance display into the interface itself, so that a user can see, at a glance, exactly which archival item supports each claim. It means deliberately including contradictory or ambiguous material in test collections so that a system's handling of disagreement can be assessed before deployment, rather than discovered by an unlucky researcher months later. And it means resisting the temptation to optimize purely for fluent, confident-sounding answers, since fluency and historical reliability are not the same thing and can, in the worst cases, actively work against one another.

Conclusion

RAG is not a gimmick, and the evidence gathered here, from a large-scale newspaper archive study to a humanities-focused thesis to a growing body of evaluation literature, makes clear that it represents a genuine advance over keyword search for anyone working with digitized historical material. But calling it groundbreaking requires a qualification. It is groundbreaking as a retrieval technology. It is not yet groundbreaking as a historiographical instrument, because none of the existing research has built evaluation standards around the values historians actually rely on to judge evidence. Closing that gap, through a framework like HAIEF, is not a minor refinement. It is the difference between a tool that finds documents quickly and a tool that a historian can actually trust.

____________________________ 

Sources:

Bhatnagar, Gauri. "Enhancing Multimodal Archival Search and Discovery in the Sound and Vision Archive Using Large Language Models and Retrieval-Augmented Generation." Master's thesis, University of Twente, August 2025. https://purl.utwente.nl/essays/108957.

Gan, Aoran. "Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey." arXiv, April 21, 2025. https://doi.org/10.48550/arXiv.2504.14891.

Jancovic, Marek. "RAG for Historians: A Radically New Way of Interacting with Historical Documents." LinkedIn, March 29, 2026. https://www.linkedin.com/pulse/rag-historians-radically-new-way-interacting-marek-jancovic-ew3xe/.

Konstantinidis, Stergios, Hayman Lotfy, Alexis Erne, Faruk Zahiragic, Min-Yen Kan, and Michalis Vlachos. "Improving Access to Historical Archives with Real-time RAG-based Systems." arXiv, July 3, 2026. https://arxiv.org/html/2607.03440v1.

Toloka. "RAG Evaluation: A Technical Guide to Measuring Retrieval-Augmented Generation." Toloka Blog, August 15, 2025. https://toloka.ai/blog/rag-evaluation-a-technical-guide-to-measuring-retrieval-augmented-generation.

Veridian Software. "What Is RAG and How Could It Support Digital Collection Search?" Veridian Software Knowledge Base, June 29, 2025. https://veridiansoftware.com/knowledge-base/a-new-way-to-search-digital-collections-introducing-rag.

Thursday, August 06, 2026

Why Metadata Matters in AI Systems


by Alvin Blackshear  | Historian & Researcher  <ablackshear@gmail.com>

There is a quiet assumption embedded in most conversations about artificial intelligence: that the intelligence lives in the model. Bigger parameter counts, cleverer architectures, more refined training regimes. These are the variables that dominate public discussion and marketing copy alike. Yet anyone who has built or evaluated a retrieval-augmented generation system knows that this framing leaves out something essential. A language model, however capable, is only as trustworthy as the information it is given to reason over. And the quality of that information is determined less by the words on the page than by the structure surrounding them: the metadata.

Metadata is often treated as an administrative afterthought, the digital equivalent of a filing label. This framing badly undersells what metadata actually does inside a modern AI system. It is the connective tissue that turns a pile of documents into something resembling knowledge. It carries context, provenance, chronology, and relationships. It tells a system not just what a piece of information says, but where it came from, when it was true, how much to trust it, and how it relates to everything else the system knows. Strip metadata away, and even the most sophisticated retrieval pipeline is reasoning in the dark.

Metadata Is the Context Behind the Data

Every document a system retrieves carries an implicit history. It was written by someone, at some point, for some purpose, and it exists in some relationship to other documents. Without metadata capturing that history, an AI system has no way to distinguish a peer-reviewed clinical guideline from an anonymous forum post, or a policy that was superseded last year from one still in force. Both simply become "text" to a retrieval engine, indistinguishable in weight or authority.

This is not a hypothetical failure mode. Research on metadata generation within relational data warehouse processes has shown that when systems are built to actively generate and manage metadata as part of the retrieval pipeline, rather than treating it as a static byproduct of storage, the downstream reasoning improves measurably. Framing metadata as a first-class citizen in system design, rather than an afterthought bolted onto search indexes, appears to be one of the more consistent findings across recent work in this area.

Metadata Drives Retrieval in RAG Systems

Retrieval-augmented generation (RAG) depends entirely on finding the right passage at the right moment. This sounds simple until one considers how much ambiguity lives inside ordinary language, especially around time. A query about "the current guidelines" or "recent developments" contains a temporal reference that a keyword search cannot resolve on its own. Recent work on handling fuzzy time expressions in RAG systems tackled exactly this problem, introducing temporal metadata filtering that allows a system to reason about vague date references rather than ignoring them. The reported result was notable: a 15.7 percent improvement in retrieval hit rate when temporal metadata was incorporated into the filtering process. That is not a marginal gain. It represents the difference between a system that reliably surfaces the correct document and one that gets there by chance roughly one in six fewer times.

This finding deserves attention because it reframes what "better retrieval" actually means. Much of the public conversation about improving RAG systems centers on embedding quality or chunking strategy. Those matter, but the temporal robustness research suggests that metadata-aware filtering can produce gains of a similar order of magnitude, often at a fraction of the computational cost of retraining or fine-tuning an embedding model.

Metadata Improves AI Accuracy and Reduces Hallucination

Hallucination remains the most persistent trust problem in generative AI, and much of the discourse around it treats the phenomenon as a model-level flaw to be solved through better training. But a meaningful portion of hallucination in RAG systems is a retrieval problem wearing a generation costume. If a system retrieves irrelevant, outdated, or low-authority material and passes it to the language model as context, the model will do exactly what it is designed to do: synthesize a fluent answer from what it was given. The result looks like a hallucination, but its root cause is upstream.

Work on reliable retrieval-augmented feature generation has examined precisely this relationship, showing that the reliability of downstream reasoning is closely tied to the quality of what gets retrieved in the first place. This is a useful corrective to any evaluation strategy that only scores final outputs. If a benchmark measures answer accuracy without also scoring retrieval precision and metadata fidelity, it risks rewarding systems that are lucky rather than sound, and it makes diagnosing failure far harder than it needs to be.

Provenance and Source Trustworthiness

Few domains illustrate the stakes of provenance more clearly than clinical decision support. A framework for auditable, source-verified clinical AI, integrating retrieval-augmented generation with explicit data provenance, was designed around a simple but demanding principle: every recommendation the system makes should be traceable back to a verifiable source, with that source's reliability made visible rather than assumed. This is provenance metadata doing real work, not as a compliance checkbox but as a functional requirement for a system operating in a high-stakes domain.

The broader lesson generalizes well beyond medicine. Any AI system that produces recommendations, summaries, or answers benefits from being able to answer the question "where did this come from" with something more specific than "the training data" or "a document somewhere in the index." Provenance metadata is what makes that specificity possible.

Temporal Reasoning: Knowing When Something Was True

Facts have expiration dates, even when the sentences describing them do not visibly change. A statement that was accurate in 2022 may be false today, and a document that has been formally superseded may still sit in a retrieval index looking exactly as authoritative as its replacement. The fuzzy time expression research referenced earlier addresses this directly, and it is worth returning to because temporal reasoning touches nearly every other capability discussed here. Explainability depends on knowing when a claim was made. Trust depends on knowing whether a claim still holds. Citation depends on being able to tell a user not just what a source says but whether it remains current. Metadata that encodes validity windows, supersession relationships, and publication dates gives a system the raw material to reason about all of this rather than presenting every retrieved passage as equally timeless.

Metadata Enables Explainable AI

Explainability is frequently discussed as though it were purely a matter of model interpretability: attention visualizations, feature attributions, and the like. But for retrieval-augmented systems, a more practical and arguably more useful form of explainability comes from metadata itself. If a system can show which document it drew from, who authored it, when it was published, and how authoritative that source is judged to be, it has already given a user most of what they need to evaluate the answer's credibility. Research introducing governance-driven agentic retrieval chains has pushed this idea further, proposing explicit evidence ledgers that log authority scores, temporal validity, and conflict resolution decisions across a multi-step retrieval process. This kind of ledger transforms explainability from a static disclosure into an auditable trail, which is a meaningfully different and more rigorous standard.

Metadata Supports Citation Generation

Citation is where metadata's value becomes most visible to an end user. A system that can point to the exact source of a claim, along with its date, authorship, and context, is offering something categorically different from a system that simply asserts a fact. Work applying retrieval-augmented generation to art provenance research within a major cultural heritage index demonstrates this well. In a domain like art history, where the chain of ownership and attribution is often the entire substance of the inquiry, metadata is not a supporting feature. It is the subject matter itself. The success of that research in surfacing accurate, explainable provenance chains for historical artworks offers a useful proof of concept for citation generation in far less specialized domains.

Metadata Helps AI Understand Document Relationships

Individual documents rarely stand alone. They respond to one another, build on one another, and sometimes contradict one another. Research on adaptive information management for retrieval-augmented generation has explored how systems can maintain coherent reasoning across multiple retrieval steps by tracking these relationships in working memory rather than treating each retrieval as an isolated event. This has particular relevance for historical research, where understanding how one document relates to, revises, or is revised by another is often the actual analytical task. Metadata that captures these relational links allows an AI system to support that kind of layered inquiry rather than flattening it into a series of disconnected facts.

Metadata Enables Better AI Evaluation

Evaluation of AI systems has tended to focus heavily on final-answer accuracy, but a growing body of work argues for scoring the retrieval and reasoning process itself. Research on self-evaluation driven strategy optimization in agentic retrieval introduced a self-assessment step in which the system evaluates the quality of retrieved evidence before generating a final answer. This kind of built-in scoring rubric, applied to metadata quality and source reliability rather than only to output fluency, represents a more rigorous approach to benchmarking. It suggests that future evaluation standards for RAG systems should include explicit metrics for provenance completeness, temporal accuracy, and citation traceability, not only for the correctness of the final generated text.

Metadata Improves Security and Governance

Metadata also does quiet but essential work in access control. Permission-aware retrieval, in which a system respects who is allowed to see which documents based on classification, ownership, or sensitivity metadata, is what makes AI systems viable in enterprise and regulated environments. The governance-driven framework mentioned earlier extends this idea by building authority scoring and audit trails directly into the retrieval chain, treating governance not as a separate layer bolted on top of the system but as something woven through its metadata architecture from the start.

The Future: Metadata-Aware AI Agents

Looking ahead, the most interesting frontier is not a smarter language model but a more metadata-literate one. Future agentic systems will likely reason over metadata with the same deliberateness they currently apply to text: weighing provenance, tracking temporal validity, assigning confidence scores, generating citations automatically, integrating with knowledge graphs, handling multimodal metadata across images, audio, and video, and preserving chain of custody across long, multi-step workflows. This shift would move AI systems away from merely producing plausible-sounding text and toward producing answers that are transparent, traceable, and genuinely defensible under scrutiny.

That is a meaningful distinction, and it is worth sitting with. Plausibility and defensibility are not the same standard. A fluent answer can be plausible and still wrong. A defensible answer, grounded in well-structured metadata, gives a user the means to check it. As AI systems take on more consequential tasks, from clinical support to historical research to enterprise decision-making, that difference is likely to matter more, not less. The path toward more trustworthy AI runs directly through the unglamorous, essential work of getting metadata right.

---

Sources

Chen, L.-C., Chen, H.-W., & Chen, M.-S. (2026). DeFuzzRAG: Handling Fuzzy Time Expressions for Temporal Robustness in Retrieval-Augmented Generation. Proceedings of AAAI-26. https://ojs.aaai.org/index.php/AAAI/article/view/40276

An Auditable and Source-Verified Framework for Clinical AI Decision Support: Integrating Retrieval-Augmented Generation with Data Provenance. Frontiers in Artificial Intelligence (2026). https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2026.1737532/full

Reliable Retrieval-Augmented Feature Generation with Large Language Model Reasoning. Knowledge and Information Systems (Springer), 2026. https://link.springer.com/article/10.1007/s10115-026-02792-4

A ReAct- and RAG-Based Framework for Metadata Generation and Access in Relational Data Warehouse Processes. Big Data and Cognitive Computing (2026). https://doi.org/10.3390/bdcc10060172

LedgerRAG: Governance-Driven Agentic Chain of Retrieval for Dynamic Knowledge Scenarios. Electronics (MDPI), 2026. https://www.mdpi.com/2079-9292/15/7/1376

Reasoning with Memory: Adaptive Information Management for Retrieval-Augmented Generation. Findings of ACL 2026. https://aclanthology.org/2026.findings-acl.1834

Reflective RAG: Self-Evaluation Driven Strategy Optimization in Agentic Retrieval-Augmented Generation. Findings of ACL 2026. https://aclanthology.org/2026.findings-acl.648

Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index (2026). https://eprints.whiterose.ac.uk/id/eprint/238115

Wednesday, August 05, 2026

I Asked ChatGPT to Find My Professor Profile. It Failed. Here's Why It Matters.


by Alvin Blackshear  |  Historian & Researcher   <ablackshear@gmail.com>

Why Historians Should Never Trust "I Couldn't Find It."

I asked ChatGPT to find my profile on Rate My Professor. It told me, with no hedging at all, that it could not find a professor profile under that name. That answer was wrong. When I added the name of the college where I taught, the model found my profile immediately, in the same database it had just searched a moment before.

When I asked why its first answer had been incorrect, ChatGPT explained that the initial search had been too narrow and that it had overstated the result. It went further, telling me that it had not been appropriate to state definitively that no profile existed. That admission is more interesting than the original mistake. It names, almost exactly, the failure that researchers of every kind run into constantly: absence of evidence quietly becomes evidence of absence, and a system built to retrieve information ends up asserting a negative it has no real basis for.

This is a small experiment. But small experiments are often where large problems become visible. What follows is an attempt to take this one seriously, not as a complaint about a chatbot, but as a case study in how retrieval systems fail, why those failures are harder to catch than ordinary hallucinations, and what a better evaluation standard would look like.

The Failure Was Not a Hallucination

It is worth being precise about what did not happen here. The model did not invent a fake Rate My Professor page. It did not fabricate reviews, ratings, or a biography. Hallucination, in the sense the AI research community usually means it, involves generating content that has no basis in the retrieved evidence or in reality. That is not what happened in this case.

Instead, the model committed a quieter and, in some ways, more dangerous error. It ran an incomplete retrieval, treated the incomplete result as sufficient, and then converted that insufficient evidence into a categorical claim. Four distinct failures stacked on top of each other:

  1. The retrieval step was incomplete, since the search used only a name and not the additional identifying detail (the institution) that would have disambiguated the result.
  2. The evidence gathered was insufficient to support any firm conclusion, positive or negative.
  3. The model's confidence should have been low, given how thin the search actually was.
  4. The final answer failed to reflect that low confidence, presenting a tentative non-result as though it were a settled fact.

Hallucinations tend to draw attention because they are often vivid and easy to catch. A fabricated citation or an invented quote can be checked against a source and shown to be false. This kind of error is harder to catch precisely because it looks like careful, responsible behavior. The model appeared to be reporting a fact rather than making one up. It sounded modest. It just was not modest enough about the right thing.

Why This Matters

Strip the incident down to its underlying logic and it reads like this: I searched. I did not find it. Therefore it does not exist. Stated plainly, that reasoning is invalid, and most people would recognize the problem immediately if it were presented as a syllogism in a logic class. Yet this is exactly the shortcut that retrieval systems take when they are not explicitly designed to resist it.

Anyone who has done serious research recognizes this pattern because it is the default condition of research itself. Archives are incomplete. Metadata is inconsistent. Names appear under variant spellings, married names, initials, or transliterations. Sources are catalogued under headings that make sense to an archivist but not to a later researcher. Databases frequently require an additional piece of context, an institution, a date range, a middle name, before they will surface a record that has been sitting there the entire time.

None of this is exotic. It is the ordinary condition of information work. The question is not whether an AI system will encounter it. The question is whether the system is built to recognize when it has hit this wall, or whether it will simply report the wall as an empty room.

Historians Have Solved This Problem for Centuries

There is a useful comparison to be made with historical methodology, which has spent a very long time developing habits of mind for exactly this situation. A historian who cannot find a record does not usually conclude that the record never existed. Instead, the working assumption is procedural: did I search enough archives, are there alternate spellings of this name, did the person's name change through marriage or naturalization, is this source catalogued under a different heading, is there another repository that might hold a more complete record.

The distinction that matters here is between two very different sentences. "It doesn't exist" is a conclusion about the world. "I haven't found evidence yet" is a conclusion about the search. Good historians are trained, often through painful experience, to default to the second sentence and to treat the first as a claim that requires an extraordinary amount of additional work before it can be responsibly made.

This is not a minor stylistic habit. It is the entire discipline's defense against a kind of error that predates computers by centuries: mistaking the limits of one's own search for the limits of what actually happened.

What Actually Failed

Returning to the four evaluation points named earlier, each deserves a closer look from an evaluator's perspective rather than a user's perspective.

Retrieval was incomplete because the search relied on a single identifier, a name, in a database where names alone are frequently ambiguous or thin. A single search pass is rarely sufficient in any system where entities can share names, use variant spellings, or be indexed under partial records.

Evidence was insufficient because a single unsuccessful query does not constitute a thorough investigation. One null result from one search string tells a system almost nothing about whether a record exists elsewhere in the same database under a slightly different query.

Confidence should have been low because the retrieval process itself provided no strong signal either way. Low retrieval coverage should map to low confidence in any well calibrated system, and a wide gap between actual coverage and stated confidence is itself a measurable defect.

The final answer should have expressed that uncertainty rather than erasing it. The correct output was never "no profile exists." The correct output was closer to "I could not locate one with the search I ran, and here is what might help me look further."

How a Better AI Should Respond

It is useful to imagine the same interaction handled well. Instead of stating that there is no professor profile, a more careful system might have said something like this: I could not locate one using your name alone. If you can tell me the institution where you taught, I can search more precisely.

Notice what that sentence does. It preserves uncertainty instead of erasing it. It invites the user to supply the missing retrieval key rather than closing the inquiry. It treats the null result as information about the search process, not information about the world. This is a small rewording, but it represents a genuinely different epistemic posture, one that treats retrieval as provisional until proven otherwise.

Why This Matters for RAG Systems

This incident connects directly to a broader question about how Retrieval-Augmented Generation (RAG) systems are built and evaluated. RAG systems are only as good as the retrieval pipeline underneath them, and that pipeline depends on several interacting components: retrieval quality, metadata structure, search strategy, query expansion, ranking, and the underlying coverage of the source material itself.

Recent research on RAG evaluation has tried to formalize exactly this dependency. A comprehensive review of RAG evaluation dimensions by Knollmeyer and colleagues works through how retrieval quality, question types, and evaluation metrics interact, and it makes a case that evaluation frameworks need to look separately at retrieval performance and generation performance rather than treating a RAG system's output as a single undifferentiated block. That separation matters here, because the language model itself did not fail at composition or fluency. The retrieval layer beneath it simply stopped searching too early, and the generation layer then dressed that shortfall up as a confident answer.

Work on uncertainty-aware retrieval reinforces the same point from a different angle. Research on uncertainty-aware dynamic retrieval argues that retrieval decisions should be driven by the model's own uncertainty rather than by fixed, one-shot retrieval rules, so that a system recognizes when its confidence is too low to stop searching. Adaptive multi-source retrieval frameworks make a related argument, showing that incorporating query complexity and confidence-aware fusion allows a system to keep pulling in additional sources precisely in the situations where a single retrieval key, like a name without an institution, is not enough to settle the question. Hypothesis-and-verification approaches to retrieval push this further still, treating an initial null result not as a final answer but as a hypothesis to be tested against additional evidence before any conclusion is reported. Self-evaluation approaches to agentic retrieval add another layer, having the system review its own retrieval strategy before committing to a final answer, which is exactly the internal check that would have caught this failure before it reached the user.

None of this is exotic academic theorizing removed from practical stakes. It describes, almost mechanically, what would have needed to happen for ChatGPT to say something more honest than "no profile exists."

An AI Evaluation Rubric

It is possible to turn this single incident into a reusable evaluation checklist, one that applies whenever an AI system reports that something cannot be found. Work on faithfulness-aware uncertainty quantification for fact-checking RAG outputs provides useful scaffolding here, since it separates the question of whether an answer is faithful to its retrieved evidence from the separate question of whether that evidence is factually sufficient in the first place. Applied to a "not found" claim, the following questions become a practical audit:

Were multiple search strategies attempted, or did the system stop after a single query. Were alternate identifiers used, such as an institution, a date, or a variant spelling. Was metadata expanded to account for inconsistent cataloguing. Was uncertainty reported explicitly, rather than folded silently into a flat statement. Did the model distinguish between "not found" and "does not exist." Could one additional retrieval key plausibly change the result.

That last question turned out to be the entire story in this case. One additional identifier, the name of the institution, was all it took to convert a false negative into a correct answer. A system that had asked itself that question before answering would have caught its own mistake.

Implications Beyond Historians

This failure mode is not confined to historians looking up their own teaching profiles. It shows up wherever retrieval sits underneath a serious decision. Legal researchers checking whether a precedent exists face the same risk of concluding "no such case" from an incomplete search of case law. Genealogists working with inconsistent historical spellings run into it constantly. Medical literature searches can miss relevant studies indexed under different terminology, with consequences far more serious than a mistaken professor lookup. Journalists checking whether a claim has prior reporting, researchers conducting literature reviews, and enterprise knowledge bases used inside companies all depend on retrieval systems that can quietly convert "I didn't find it in my index" into "it isn't true." A recent Nature article on synthesizing scientific literature with retrieval-augmented language models makes a related point in the scholarly context specifically, arguing that trustworthy synthesis of academic evidence depends on retrieval systems that are transparent about their own coverage limits rather than presenting synthesized answers as though they rested on complete evidence.

In every one of these domains, the underlying mechanism is identical to what happened with a single Rate My Professor search. One missing retrieval key can completely change the answer, and a system that does not know how to say "I haven't looked hard enough yet" will eventually say something false with total confidence.

Conclusion

Return to where this started. ChatGPT eventually found my profile once I supplied the one piece of context it needed, the institution where I taught. The information had been there the entire time. The system's retrieval capability was never really in question. What failed was something quieter and, in the end, more consequential: the model treated an incomplete search as a complete answer and reported a negative conclusion before it had exhausted reasonable ways of looking.

The most important lesson here was not that ChatGPT failed to find my professor profile. It was that, for a moment, it mistook an incomplete search for a finished one. For historians, for researchers in any field, and for anyone building AI systems meant to be trusted with real questions, that distinction is not a technicality. It is the whole matter.

---

References

Asai, Akari, et al. "Synthesizing Scientific Literature with Retrieval-Augmented Language Models." Nature (2026). https://www.nature.com/articles/s41586-025-10072-4

Fadeeva, Ekaterina, et al. "Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation." Findings of ACL 2026. https://aclanthology.org/2026.findings-acl.338

Knollmeyer, Simon, et al. "Evaluating Retrieval Augmented Generation: A Comprehensive Review of Evaluation Dimensions, Question Types, and Application." SN Computer Science 7 (2026). https://link.springer.com/article/10.1007/s42979-026-05134-x

Wu, Haiyan, et al. "Reflective RAG: Self-Evaluation Driven Strategy Optimization in Agentic Retrieval-Augmented Generation." Findings of ACL 2026. https://aclanthology.org/2026.findings-acl.648

"AdaKG-RAG: Adaptive KG-guided Retrieval-Augmented Generation with Hypothesis-and-Verification Retrieval for Multi-hop Question Answering." Journal of King Saud University Computer and Information Sciences (2026). https://link.springer.com/article/10.1007/s44443-026-01028-3

"UA-RAG: Uncertainty-aware Dynamic Retrieval-Augmented Generation." Neurocomputing (2026). https://www.sciencedirect.com/science/article/pii/S0925231226017558

Tuesday, August 04, 2026

AI Fact Verification for Historians: Building a Discipline-Specific Protocol for a New Research Environment


by Alvin Blackshear | Historian & Researcher  <ablackshear@gmail.com>

Historians have always worked at the boundary between evidence and interpretation, testing claims against sources, weighing provenance, and living comfortably with uncertainty. Generative artificial intelligence has entered this terrain quickly and unevenly, offering historians an extraordinary tool for discovery while introducing new categories of risk that the discipline has not yet fully reckoned with. The question facing the field is not whether AI belongs in historical research. It already does. The real question is how historians verify what AI produces, and whether existing habits of source criticism are sufficient for a technology that can fabricate footnotes, invent quotations, and drift toward revisionist narratives when prompted to do so.

This article argues that AI can meaningfully accelerate historical research, but it cannot replace professional source criticism, and it proposes the outline of a repeatable verification protocol suited to the particular demands of archival research, primary source authentication, citation validation, chronology checks, and historiographical consistency.

The stakes of getting this right extend beyond any single article or dissertation. History as a discipline depends on a shared confidence that its claims can be traced back to something real, a document, an object, a testimony that survived. When that chain of evidence is compromised, even quietly and even with good intentions, the damage does not stay contained to one footnote. It erodes the credibility of the discipline as a whole, at a moment when public trust in expertise is already under strain. Historians who adopt AI tools without adopting equally rigorous verification habits are not simply risking their own reputations. They are testing the durability of a field that has spent centuries building methods for distinguishing memory from evidence and evidence from invention.

The Promise and the Ceiling

Large language models can summarize secondary literature, suggest starting bibliographies, identify patterns across large text corpora, and help historians locate materials they might otherwise have missed. Used well, these tools compress hours of preliminary searching into minutes. The American Historical Association's Ad Hoc Committee on Artificial Intelligence in History Education, in its 2025 guiding principles, acknowledges this potential while insisting on a firm boundary: generative AI can mimic some of the work historians do, but this should never be mistaken for the work itself. The committee's language is worth sitting with. AI produces texts, images, audio, and video, not truths. It selects words and patterns from training data rather than comprehending the past in its full complexity and contradiction. When a model encounters a gap in its training, it does not pause to acknowledge uncertainty. It fills the gap with something plausible sounding, a process the field now calls hallucination.

For historians, hallucination is not an abstract technical curiosity. It is a direct threat to the evidentiary basis of scholarship. A fabricated citation that looks correctly formatted, attached to a real-sounding archive or journal, can pass a cursory glance and enter a footnote before anyone notices the source does not exist. The AHA's guiding principles are unambiguous on this point: any reference generated by AI must be checked against the original before it is used, and a reference included in a footnote without that check is never acceptable practice, regardless of how the historian frames their working process.

Why Automated Fact-Checking Falls Short

It is tempting to assume that the solution to AI hallucination is simply more AI, in the form of automated fact-checking systems layered on top of generative tools. Recent research complicates that assumption considerably. A 2026 study led by Kelly Amaddio examined how people respond to automated fact-checkers correcting misinformation, using a sample of more than thirteen hundred participants exposed to a piece of gun control misinformation and a subsequent AI-generated correction. The findings revealed a persistent and troubling pattern: prior beliefs shaped both continued endorsement of the misinformation after correction and perceptions of the fact-checking system's reliability. When the fact-checker was described as only moderately accurate, close to the accuracy levels people already associate with existing AI systems, prior beliefs predicted whether the correction actually changed minds. Even when participants were told the system was highly accurate, the effect of prior belief did not disappear, only lessened.

This matters for historians because it demonstrates that automated verification does not operate in a neutral space. It is filtered through the same human biases that shape how any evidence is received, and it can be defeated by claims that already align with what someone wants to believe. Historical scholarship, particularly on contested and politically charged topics, is especially vulnerable to this dynamic. A fact-checking layer bolted onto a language model cannot substitute for the disciplinary training that teaches historians to interrogate their own assumptions as rigorously as they interrogate a source.

Revisionism as a Distinct Problem

Contested historical claims present a challenge that generic fact-checking frameworks were never designed to solve, because the disagreement is not always about a discrete, verifiable fact but about interpretation, framing, and the political stakes attached to memory. A 2026 study by Francesco Ortu and collaborators addresses this directly. The researchers built a benchmark called Historical Misinfo, drawing on five hundred contested historical events from forty five countries, each paired with a documented factual narrative and a documented revisionist counter-narrative. They then tested how large language models responded under neutral prompting compared with prompting that explicitly requested the revisionist version of events.

The results are instructive. Under neutral conditions, most models leaned closer to the factual reference. But when users explicitly requested a revisionist framing, every model tested showed a sharp increase in revisionist output, with little resistance or self-correction. This finding should give historians real pause. It suggests that a model's apparent reliability on a given topic says little about how that same model will behave when a user, deliberately or not, nudges it toward a distorted account. For historians working on genocide, colonialism, contested territorial claims, or any subject where competing national narratives exist, this is not a marginal concern. It is central to responsible use.

What Verification-First Research Looks Like Elsewhere

Historians are not alone in confronting this problem. Related work emerging from adjacent fields offers useful models for how disciplines might restructure their relationship with AI around verification rather than convenience. Lei You, Lele Cao, and Iryna Gurevych have argued that AI-assisted peer review in the sciences should be verification-first rather than review-mimicking, proposing that AI tools function as adversarial auditors generating checkable evidence rather than as systems that simply predict what a human reviewer might conclude. Belinda Mo extends this argument to the broader landscape of autonomous AI research agents, warning of a widening verification gap between what AI systems can produce and what any human can realistically check, and calling for verification infrastructure built around observable workflows and clear attribution.

A study by Raymond Solga and Mohammed Sarwar offers perhaps the closest parallel to the historian's own situation. Comparing traditional historical methods against several generative AI tools in identifying and validating primary sources in ancient history, the researchers found that AI performed well at broad content discovery and thematic synthesis but struggled significantly with genre boundaries, provenance transparency, and the kind of contextual interpretation that trained historians bring almost instinctively to a source. Their proposed solution, a human-in-the-loop framework built around model pluralism and provenance-first protocols, maps closely onto what historians already know from decades of archival practice, simply extended to include AI as one input among many that must be checked rather than trusted.

There is also a temporal dimension worth naming. Carlo Iacono's recent audit of AI-assisted research output found that the models cited in published work are frequently outdated by the time readers encounter them, sometimes by close to a year, with model versions superseded and behaviors shifting in ways that are rarely documented. For historians, this is a reminder that any claim about what a given AI tool can or cannot do should be dated and treated as provisional rather than as a permanent property of the technology.

Toward a Repeatable Protocol

Drawing these threads together, a workable verification protocol for historians should rest on several commitments. Every AI-generated citation must be checked against the primary source or an authoritative secondary source before it appears in any written work, with no exceptions for citations that merely look plausible. Chronology and provenance claims generated by AI require independent confirmation, since these are precisely the areas where current tools show the weakest performance. Contested or politically charged claims deserve heightened scrutiny, given the demonstrated tendency of models to drift toward whatever framing a prompt implies. Verification work should be documented as it happens, noting which tool was used, when, and what was checked, both for the historian's own accountability and because the tools themselves change quickly enough that undocumented claims about their reliability age poorly. And throughout, the historian's own expertise must precede rather than follow AI use, since evaluating a model's output well requires the same depth of subject knowledge that historians have always needed to evaluate any source.

None of this diminishes the real value AI offers historians willing to use it carefully. It does insist that the discipline's oldest habit, treating every claim as something to be verified rather than assumed, remains the right response to a genuinely new technology. Generative AI can accelerate the search for evidence. It cannot substitute for the judgment that turns evidence into history.

There is a version of this moment that history has lived through before, in smaller ways, each time a new technology promised to change how the past is studied. Microfilm, digitization, keyword-searchable databases, all reshaped the pace and reach of archival work without ever touching the core discipline of asking what a source can and cannot tell us. Generative AI is different in degree, not necessarily in kind. It compresses years of searching into minutes, but it also compresses the distance between plausible and true, and that compression is precisely where a historian's training earns its keep. The AHA's committee put it plainly: there are no shortcuts to expertise, and the skills that let a scholar recognize a hallucinated citation or a subtly revisionist framing can only be built through the same sustained engagement with sources that has always defined the craft. A historian who verifies rigorously will find in AI a genuine research partner. A historian who does not will find, sooner or later, that the past they have written no longer holds.

______________________________

Sources

American Historical Association, Ad Hoc Committee on Artificial Intelligence in History Education. "Guiding Principles for Artificial Intelligence in History Education." Approved by AHA Council, July 29, 2025. https://www.historians.org/resource/guiding-principles-for-artificial-intelligence-in-history-education/

Amaddio, Kelly M., et al. "Prior Beliefs and Automated Fact Checking: Limits on the Accuracy of AI Verification." February 5, 2026. https://doi.org/10.1371/journal.pone.0342332

Iacono, Carlo. "The Half-Lives of Generative-AI Evidence." July 27, 2026. https://doi.org/10.48550/arXiv.2607.24032

Mo, Belinda. "The Age of AI Agents Demands a New Scientific Paradigm to Sustain Trustworthy Science." June 18, 2026. https://doi.org/10.48550/arXiv.2607.26064

Ortu, Francesco, et al. "Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models." February 22, 2026. https://doi.org/10.48550/arXiv.2602.17433

Solga, Raymond S., and Mohammed J. Sarwar. "Evaluating Generative AI in Historical Research." AI & Antiquity (2026). https://doi.org/10.64946/aiantiquity.v2i1.003

You, Lei, Lele Cao, and Iryna Gurevych. "Preventing the Collapse of Peer Review Requires Verification-First AI." February 12, 2026. https://doi.org/10.48550/arXiv.2601.16909