Wednesday, August 12, 2026

The Biggest Mistake Historians Make When Using ChatGPT

 

by Alvin Blackshear  |  Historian & Researcher

A historian asks ChatGPT a seemingly straightforward question: “Who was the first African American elected to public office in Delaware after Reconstruction?

Within seconds, the system produces a polished answer. It supplies a name, an election date, a description of the office, and perhaps even several citations. The prose is orderly, confident, and specific. Nothing about it appears careless. The historian copies the response into a research notebook and proceeds to the next question.

Only later does the answer begin to unravel. One citation does not exist. A quotation attributed to a nineteenth-century newspaper is actually a modern paraphrase. The date belongs to another officeholder. An appointment has been confused with an election. Details from two people have been combined into a single biography.

Nothing in the original response sounded implausible. That is precisely the danger.

The greatest mistake historians make when using ChatGPT is not merely trusting an occasional incorrect answer. It is treating AI-generated prose as though it were historical evidence. ChatGPT produces narratives. Historians require evidence. Those activities may sometimes overlap, but they are not the same.

The Abandonment of Historical Method

Historians already possess a sophisticated method for evaluating claims. When examining a letter, newspaper, memoir, photograph, government record, or oral history, they ask familiar questions. Who created the source? When was it created? For what purpose? Under what circumstances? What interests shaped its contents? Can its claims be corroborated? Has another scholar interpreted the evidence differently?

These questions constitute more than an academic ritual. They are the means by which historical knowledge is distinguished from rumor, memory, advocacy, folklore, and invention.

Yet something curious happens when many researchers begin using generative artificial intelligence. The normal habits of source criticism are suspended. A response is accepted because it sounds reasonable. A citation is trusted because it resembles an academic citation. A quotation is repeated because the language seems appropriate to the period.

The historian who would never accept an unsigned reminiscence without examining its provenance may accept an AI-generated paragraph without asking where any of its claims originated.

ChatGPT should therefore be treated as an intelligent research assistant, not as a documentary source. It can suggest search terms, identify possible lines of inquiry, compare concepts, organize notes, formulate questions, and help clarify difficult prose. It can also point a researcher toward archives, books, people, and events that merit investigation. Its usefulness, however, does not convert its output into evidence.

The distinction is fundamental. A research assistant may tell a historian that a document probably exists. The historian must still locate and examine it.

Why Fluency Is So Persuasive

Large language models are especially dangerous when they are almost correct. An obviously absurd answer invites skepticism. A plausible answer, containing real names, accurate background information, and one fabricated detail, may pass unnoticed.

ChatGPT’s authority is largely rhetorical. It writes in complete sentences, arranges events chronologically, supplies transitions, and often presents conclusions without visible hesitation. Readers naturally associate these qualities with knowledge. Confidence, specificity, and coherence become substitutes for provenance.

Recent research suggests that this problem is not simply the result of careless prompting. Kalai and his colleagues argue that common accuracy-based evaluations may reward models for guessing instead of acknowledging uncertainty. If a benchmark gives credit only for correct answers and imposes no meaningful cost for plausible errors, the model is encouraged to answer even when abstention would be more responsible. The researchers propose open scoring rubrics that disclose how errors and abstentions will be evaluated.

This observation has important consequences for historians. A model that always produces an answer may appear more useful than one that frequently says, “I do not have enough evidence.” In historical research, however, an admission of uncertainty may be the more accurate and intellectually responsible response.

The historian must therefore evaluate not only what the system says, but whether the system should have answered at all.

Five Forms of Historical Hallucination

Historical hallucinations frequently assume recognizable forms.

The first is the invented citation. ChatGPT may produce a realistic book title, journal article, archival collection, author, volume number, or page range. Every component may appear academically credible even though the source does not exist.

The second is the misquotation of a primary source. A model may modernize, compress, or reconstruct a statement and then present the resulting language inside quotation marks. The general sentiment may be accurate, but the words are not documentary evidence.

The third is the composite biography. Details belonging to several individuals may be joined into one apparently coherent life. This is especially dangerous when researching people who share names, occupations, institutions, military units, or geographic locations.

The fourth is incorrect chronology. A model may identify real events but arrange them in the wrong order, confuse the date of an appointment with the date of an election, or place a person at an institution before that person arrived there.

The fifth is false causation. ChatGPT may connect two events with phrases such as “therefore,” “as a result,” or “this led to,” even when the evidence establishes only sequence or correlation.

These errors are not equally visible. An invented person may be discovered quickly. A subtle chronological error or unsupported causal inference may survive several rounds of editing because it fits an expected narrative.

Why Better Prompts Are Not Enough

Prompt design can improve an AI response. A historian can request citations, ask the system to distinguish facts from interpretations, require expressions of uncertainty, or instruct it not to invent missing information. Such practices are worthwhile.

They are not verification.

A carefully written prompt may reduce the probability of an error, but it cannot transform generated language into historical evidence. Even a response that includes citations must be checked against the cited material. Even a quotation accompanied by a page number must be located on that page. Even a claim presented as certain must be corroborated.

Research on hallucinations in academic writing identifies fabricated citations, factual distortion, named-entity errors, logical inconsistency, and propagation errors as threats to scholarly integrity. The final responsibility remains with the human author. An historian cannot excuse a false statement by explaining that ChatGPT supplied it.

Better prompting is therefore a research technique, not an evidentiary standard.

Testing AI as Historians Test Evidence

Historians need a repeatable method for testing AI-generated research. The evaluation should move beyond impressions such as “the answer seemed good” or “the model performed well.”

A practical scoring rubric might assess eight dimensions:

1.      Citation existence: Does every cited work or archival collection exist?

2.      Citation accuracy: Does the cited source support the specific claim?

3.      Quotation fidelity: Are quoted words reproduced exactly and in context?

4.      Identity accuracy: Have people with similar names or careers been distinguished?

5.      Chronological accuracy: Are dates and sequences correct?

6.      Geographic and institutional accuracy: Are places, offices, organizations, and jurisdictions correctly identified?

7.      Provenance: Can each important assertion be traced to accessible evidence?

8.      Historiographical consistency: Does the answer acknowledge significant scholarly disagreement?

Each category could be scored from zero to four. A zero would indicate fabrication or complete failure. A score of one would reflect major errors. Two would indicate partial support or substantial ambiguity. Three would represent generally accurate work with minor deficiencies. Four would require full and independently verified support.

A weighted rubric would assign greater penalties to invented citations, false quotations, and merged identities than to minor stylistic or contextual omissions. This matters because not all errors have the same scholarly consequence.

Benchmark testing should also use a fixed set of questions rather than memorable anecdotes. Researchers could create a collection of historical questions at several difficulty levels: well-documented national events, obscure local events, disputed interpretations, biographical identity problems, chronological puzzles, and questions whose correct answer is genuinely unknown.

The same questions could then be submitted to several AI systems under controlled conditions. Evaluators would record accuracy, citation validity, unsupported assertions, appropriate abstentions, and changes across repeated runs. Such testing would reveal not only whether a model produces correct answers, but how it fails.

What RAG Can and Cannot Do

Retrieval-augmented generation, commonly known as RAG, offers one method of improving reliability. A RAG system retrieves documents from a designated collection before generating an answer. For historians, that collection might contain archival finding aids, digitized newspapers, oral-history transcripts, government records, or scholarly articles.

This approach can make AI responses more transparent and more closely connected to identifiable sources. It may reduce fabricated citations and permit researchers to inspect the evidence used in producing an answer.

Yet RAG does not eliminate the need for historical judgment.

Recent research distinguishes factuality from faithfulness. A statement may be faithful to a retrieved document while the document itself is inaccurate, biased, incomplete, or misinterpreted. Conversely, a statement may be factually correct but unsupported by the particular documents retrieved. A serious evaluation must test both dimensions.

For historians, retrieval is only the beginning. The fact that a system retrieved a newspaper article does not establish that the article was truthful. The fact that a claim appears in a memoir does not establish that memory was reliable. The fact that several sources repeat a statement does not prove they were independent.

RAG can retrieve evidence. It cannot perform the full intellectual work of source criticism.

The Historian’s Continuing Responsibility

ChatGPT does not threaten historical scholarship simply because it occasionally invents facts. Historical scholarship has always confronted error, forgery, propaganda, partial memory, and misleading testimony.

The greater danger arises when historians allow the fluency of artificial intelligence to replace the discipline’s oldest habit: skepticism.

For centuries, historians have questioned manuscripts, memoirs, newspapers, photographs, statistics, and archives. Artificial intelligence deserves no exemption from that tradition. Its output should be tested, scored, corroborated, and traced to evidence.

The historian’s task remains unchanged. It is not merely to discover information or compose a persuasive narrative. It is to determine what the surviving evidence permits us to say, what remains uncertain, and why the distinction matters.

References

  1. Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. “Evaluating Large Language Models for Accuracy Incentivizes Hallucinations.” Nature 653 (2026): 1047–1051. https://www.nature.com/articles/s41586-026-10549-w
  2. Rahman, Subhey Sadi, et al. “Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models.” Artificial Intelligence Review 59 (2026), article  70. https://link.springer.com/article/10.1007/s10462-025-11454-w
  3. Fadeeva, Ekaterina, et al. “Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation.” Findings of the Association for Computational Linguistics: ACL 2026 (2026): 6814–6836. https://aclanthology.org/2026.findings-acl.338
  4. Giray, Louie, Md Emdadul Islam, and John Arvin Glo. “AI Hallucinations in Academic Writing: Implications for Research Integrity.” Naunyn-Schmiedeberg’s Archives of Pharmacology (2026). https://pubmed.ncbi.nlm.nih.gov/42217043 

Monday, August 10, 2026

A Rubric for Assessing AI-Generated FOIA Summaries

by Alvin Blackshear | Historian & Researcher


Intro: A Gap Hiding in Plain Sight

Two literatures are growing quickly in 2026 and, so far, growing apart. The first is the literature on large language model (LLM) evaluation: rubric-driven frameworks for scoring factuality, faithfulness, hallucination, and task completion across generative outputs. The second is the literature on FOIA (Freedom of Information Act) automation: case studies and vendor analyses describing how agencies are using AI to triage requests, search collections, and draft redactions faster than manual review allows. Both bodies of work are maturing on their own terms. Neither has, in any formal sense, met the other.

That absence is notable given the setting. FOIA is not an ordinary text-summarization domain. A FOIA summary sits downstream of a legal disclosure decision. It distills a record that has already been searched, reviewed, exempted, and redacted, and it typically becomes part of what a requester, a court, an oversight body, or a journalist relies on to understand what an agency did and why. If an AI-generated summary drops a redacted date, blurs the line between what was withheld and what was released, or introduces even a plausible-sounding inference not supported by the underlying record, the error does not stay contained to a chatbot transcript. It can shape a public record, an appeal, or a legal filing.

Despite that stakes profile, a search of the current literature turns up remarkably little that combines the two threads. Papers on LLM evaluation rubrics rarely mention FOIA. Papers on FOIA automation rarely propose a scoring framework specific to summary quality. Almost none attempt a rubric purpose-built for the FOIA summarization task, one that treats redaction fidelity, exemption logic, and traceability back to the released record as first-class evaluation criteria, on par with the factuality and hallucination metrics borrowed from general-purpose LLM evaluation. That gap is the occasion for this article: a proposed rubric for assessing AI-generated FOIA summaries, and a look at how the leading purpose-built FOIA technology on the market today, Relativity FOIA and Relativity aiR, fits into, and exposes, that same gap.

Why FOIA Summarization Needs Its Own Rubric

General-purpose summarization rubrics, the ones built for news articles, meeting transcripts, or customer support tickets, optimize for coherence, conciseness, and topical coverage. Those criteria matter for FOIA summaries too, but they are not sufficient. A FOIA summary carries legal weight that a meeting-notes summary does not. Three features of the FOIA context justify a dedicated evaluation instrument:

·         Redaction is structural, not incidental. A FOIA record is rarely handed over whole. Exemptions under the statute (and analogous state public-records laws) remove specific categories of information, such as personal privacy, law enforcement techniques, or deliberative process, while leaving the rest disclosable. A summary that fails to distinguish "this is what was released" from "this is what was withheld and why" misrepresents the disclosure itself, not just the document.

·         Legal meaning is precise and often reader-facing. Terms like "responsive," "exempt," "segregable," and the specific statutory exemption cited (e.g., Exemption 5, Exemption 7(C)) carry defined legal consequences. A summary that paraphrases those terms loosely can change what a reader believes the agency determined.

·         The audience includes adversarial and oversight readers. Unlike an internal meeting summary, a FOIA summary may be read by litigators, journalists, or Inspectors General specifically looking for inconsistency between the summary, the redactions, and the underlying record. The evaluation bar has to anticipate that scrutiny.

These features are why importing an off-the-shelf LLM summarization rubric is inadequate on its own, and why FOIA-specific criteria, including redaction and exemption handling, legal-meaning preservation, and traceability to the released record, need to sit alongside more familiar accuracy and hallucination metrics.

A Proposed Rubric for AI-Generated FOIA Summaries

The rubric below is offered as a starting framework rather than a finished instrument. It borrows its general shape, weighted criteria that are each independently scorable, from the structured evaluation approaches now common in LLM benchmarking, while grounding the specific criteria in FOIA practice.

 Criterion

Weight

 Factual accuracy

25%

 Completeness of disclosed information

20%

 Correct treatment of redactions and exemptions

15%

 Preservation of legal meaning

15%

 Hallucination rate

10%

 Neutrality and absence of editorial bias

5%

 Citation/traceability back to released records

5%

 Readability and organization

5%


Factual accuracy (25%)
carries the largest weight because it is the load-bearing criterion for every downstream use of the summary. Evaluators should check whether every factual claim in the summary is verifiable against the record as released, with particular attention to dates, names of programs or offices (as distinct from names withheld under privacy exemptions), dollar figures, and sequences of events.

Completeness of disclosed information (20%) asks whether the summary captures the material substance of what was released, not merely a plausible-sounding excerpt of it. A summary that is accurate but omits a disclosed finding, a key admission, or a significant data point can be just as misleading as one that fabricates something, because it gives a false impression of a "clean" record.

Correct treatment of redactions and exemptions (15%) is the most FOIA-specific criterion on the list. It asks whether the summary clearly signals where information was withheld, whether it correctly reflects the cited exemption basis where that basis is itself disclosed, and, critically, whether it avoids inferring or reconstructing withheld content from context. That kind of inference is a known failure mode of generative summarization when a document has gaps.

Preservation of legal meaning (15%) evaluates whether statutory and procedural terms are used correctly and whether the summary's characterization of the agency's determination (responsive, non-responsive, partially exempt, referred to another agency, etc.) matches the record's actual disposition.

Hallucination rate (10%) is scored separately from general factual accuracy because it targets a distinct failure mode: content that has no basis anywhere in the source record, rather than content that is merely imprecise. In FOIA summarization this includes fabricated exemption rationales, invented dates, or synthesized "likely" content behind a redaction.

Neutrality and absence of editorial bias (5%) checks that the summary does not editorialize about the agency's conduct, insert framing not present in the record, or adopt a tone that implies a conclusion (such as wrongdoing, cover-up, or exoneration) the record itself does not support.

Citation/traceability back to released records (5%) asks whether a reviewer can trace each summary statement back to a specific page, Bates number, or section of the released record. This matters enormously for defensibility if the summary is later challenged.

Readability and organization (5%) is weighted lowest deliberately. It matters for usability, but a highly readable summary that fails the criteria above is a worse outcome than an awkward one that passes them.

Applied consistently, a rubric of this kind would let researchers and agencies compare AI-generated FOIA summaries across vendors, models, and prompt designs using a shared, FOIA-native standard, something the current literature does not yet offer.

Where the Industry Stands: Relativity FOIA and Relativity aiR

Relativity is a useful lens for testing this gap against real-world deployment, because it is one of the few vendors building AI specifically into the FOIA review workflow rather than adapting a general eDiscovery tool after the fact.

Relativity FOIA, announced as generally available inside RelativityOne Government, unifies records request intake, case management, AI-supported review, disclosure, and reporting into a single connected workflow. The product's own framing makes clear that AI is meant to assist rather than replace the reviewer: every prioritized record arrives with a confidence score, a plain-language explanation, a counterpoint, and supporting citations, and it is the human reviewer, not the model, who makes the final call on each document. The accompanying launch announcement strikes the same note, describing the platform as pairing defensible AI and automation with Relativity's established review technology so that agencies can standardize exemption rationale and accelerate review. That is a framing built around assistance and consistency, not autonomous disclosure decisions.

Relativity aiR, the generative AI layer underlying that workflow, does not market a standalone "summarization" feature by that name, but its analysis outputs function as summaries in practice. RelativityOne's government release notes describe a mid-2026 update to aiR for Review that lets teams build custom analyses across both text and images, producing what the release notes themselves describe as "insights, extractions, and summaries." In other words, the document-level explanations aiR for Review generates, the kind of output a FOIA reviewer would read to understand a record before deciding what to release, are a summarization capability in every functional sense, even without that specific product label.

The philosophy behind these design choices is spelled out most directly in Brian Thompson's commentary on government AI infrastructure. Thompson argues that courts and oversight bodies require agency decisions to be explainable, traceable, and defensible, and that AI can only meet that bar when it is purpose-built for legal and public sector work in a way that preserves audit trails and keeps a verifiable link between inputs and outputs, with human validation built in rather than bolted on. That is a strong statement of principle: transparency, defensibility, explainability, and reproducibility as design requirements. But it remains a statement of principle rather than a published measurement framework.

The clearest evidence of both the industry's evaluation ambition and its current limits comes from an independent study Relativity has publicized: Redgrave LLP's head-to-head comparison of aiR for Review against a traditional active-learning managed review. The results were striking on the metrics the study did measure. aiR for Review reached 88 percent recall against 64 percent for the active-learning workflow, and its elusion rate (responsive documents wrongly left in the discard pile) came in at 1 percent versus 3 percent for the manual process, while consuming roughly 18 attorney-hours against an estimated 1,123 hours for the 24-person manual review team. The study is equally candid about tradeoffs: aiR for Review's precision, at 29 percent, trailed the active-learning workflow's 39 percent, a gap the authors attribute in part to the low "richness," meaning the scarcity of responsive documents, in the test population.

That study, however, measured a responsiveness-classification task, finding documents relevant to a legal standard, not a summarization task. It offers precision and recall figures for document identification, not for the fidelity of AI-generated summaries. Reviewed against the rubric above and against the broader FOIA-specific commentary, six gaps stand out in Relativity's current public-facing material:

·         There is no formal, published FOIA summarization evaluation rubric comparable to the one proposed here.

·         There are no benchmark scores specific to AI-generated FOIA summaries, as distinct from the responsiveness-classification benchmarks that do exist.

·         There is no hallucination testing specific to FOIA summaries. The closest published work addresses document-level responsiveness accuracy, not summary-level fidelity.

·         There is no disclosed human evaluation methodology specific to FOIA summarization, meaning no description of how human reviewers score AI summary outputs against a defined standard.

·         There are no precision/recall studies of AI-generated summaries themselves, again as distinct from precision/recall for document classification.

·         There is no academic validation paper describing how Relativity measures FOIA summary quality specifically, as opposed to review-workflow efficiency more broadly.

None of this is a criticism of the product design, which is explicitly built around human validation, citation-backed outputs, and audit trails, features that a rigorous FOIA summarization rubric would actually reward under the "traceability" and "correct treatment of redactions and exemptions" criteria. It is, instead, a description of an open research space. The underlying architecture for defensible AI summarization exists, but the formal instrument for measuring whether it produces good FOIA summaries, as opposed to good responsiveness calls, does not yet exist in the public record.

Implications for Research and Practice

The rubric proposed here is deliberately modest in scope: eight criteria, weighted to reflect FOIA's legal stakes, designed to be usable by researchers, agency FOIA officers, and vendors alike. Its value lies less in the specific weights, which should be stress-tested and adjusted through empirical study, than in establishing that FOIA summarization deserves a dedicated evaluation standard rather than an inherited one.

For researchers, the immediate opportunity is to pair this kind of rubric with a labeled dataset of AI-generated FOIA summaries scored by trained reviewers, producing the kind of benchmark and hallucination-rate figures currently missing from the literature. For agencies and vendors, the opportunity is to publish exactly the kind of validation study that Relativity has already modeled for document responsiveness (blind expert review, ground-truth comparison, transparent methodology) but aimed squarely at summary quality rather than classification accuracy.

As FOIA request volumes climb and agencies lean further into AI-assisted review, the gap between AI that helps reviewers understand records and AI whose summaries have been rigorously measured against a FOIA-specific standard is the gap this research agenda should close.


Sources

1.     Relativity, "Relativity FOIA: Purpose-Built FOIA Software for Federal and State Agencies," June 2026. https://relativity.com/blog/relativity-foia-purpose-built-foia-software-for-federal-and-state-agencies

2.     Relativity, "Relativity Launches Relativity FOIA to Streamline Public Disclosure Operations for Government Agencies," June 1, 2026. https://relativity.com/news-events/relativity-launches-relativity-foia-to-streamline-public-disclosure-operations-for-government-agencies

3.     Relativity, "What's New in RelativityOne Government (FOIA)." https://help.relativity.com/RelativityOne/Content/What_s_New/What_s_new_in_RelativityOne_Government.htm

4.     Brian Thompson, "AI, Operational Infrastructure, and the Future of Government Legal Work," Relativity Blog, April 8, 2026. https://www.relativity.com/blog/ai-operational-infrastructure-and-the-future-of-government-legal-work

5.     Robert Keeling and Ray Mangum (Redgrave LLP), "Results Are In: 5 Lessons from an Independent Study of aiR for Review," Relativity Blog, June 9, 2026. https://www.relativity.com/blog/results-are-in-5-lessons-from-an-independent-study-of-air-for-review



Sunday, August 09, 2026

The Fastest Man in the World Has a Message About Slowing Down


What Usain Bolt's Jardiance Commercial Can Teach Us About CKD

Usain Bolt built his entire identity on speed. In 2009, he covered 100 meters in 9.58 seconds, a world record that still stands nearly two decades later. He is, by any measure, the fastest human being who has ever lived. So it's a little disorienting to see him on American television these days, not sprinting, but talking about the value of slowing down.

The commercial is for Jardiance, called "Slow Can Be Good Too," and it deliberately turns Bolt's public identity upside down. Speed made him famous. But the ad isn't selling speed. It's making the case that slowing something down, specifically the progression of a chronic disease, can be its own kind of victory. That contrast is what makes the advertising work. Our brains are wired to notice contradictions, and pairing the world's fastest man with a message about deceleration creates exactly the kind of cognitive friction that makes people pay attention.

Buried inside the ad is a simple line that matters far more than the celebrity casting: "When was the last time you had your kidneys checked?" That question is worth pausing on, because most people genuinely don't know the answer.

The Commercial Contains a More Important Message

It's tempting to read the ad as just another celebrity pharmaceutical spot. But the word that matters here isn't Jardiance. It's progression. Chronic kidney disease, or CKD, is a condition that typically develops slowly, often over years, and it rarely announces itself with dramatic symptoms in its earlier stages. A person can lose a meaningful share of kidney function and feel completely normal.

That quiet progression is precisely why CKD is such an underrecognized problem. According to the Centers for Disease Control and Prevention's most recent national estimates, roughly 14 percent of U.S. adults, or about 37 million people, have chronic kidney disease [1]. Far more striking is how many of them don't know it. The CDC has found that about 9 in 10 adults with CKD are unaware they have it [1]. That is not a small awareness gap. It is nearly the entire affected population moving through the disease without a diagnosis, often until it has advanced considerably.

What Exactly Is CKD?

Think of the kidneys as filters. Roughly the size of fists, sitting toward the back of the abdomen, they filter waste and excess fluid out of the blood, help regulate blood pressure, and keep the body's chemistry in balance. Chronic kidney disease means that filtering ability has been damaged and is diminished, and that the damage has persisted for at least three months.

CKD is typically staged using two pieces of information rather than one, which is part of why it's harder to explain than, say, a cholesterol number. As kidney function declines, the body loses its ability to clear waste products and regulate fluid balance, which over time raises the risk of cardiovascular disease, anemia, bone problems, and eventually kidney failure requiring dialysis or transplantation. The disease doesn't move in a straight line for everyone, but for many people it is progressive, meaning it tends to move in one direction without intervention.

The Diabetes Connection

Type 2 diabetes is the single biggest driver of CKD in the United States, and the relationship between the two is well established. Persistently elevated blood glucose damages the small blood vessels inside the kidneys' filtering units over time, gradually reducing their ability to do their job. The CDC's newest national data illustrate just how tightly linked these conditions are: an estimated 38 percent of adults with diabetes and 21 percent of adults with high blood pressure were estimated to have CKD [1]. Even among people with prediabetes, more than 1 in 10 were estimated to have CKD already [1].

This is where A1C, the blood test that reflects average blood sugar over roughly the previous three months, enters the picture. A1C doesn't measure kidney function directly, but it is one of the best available signals of the metabolic stress that, over years, can damage the kidneys. Managing blood sugar isn't just about avoiding diabetes complications in the abstract. It's directly tied to kidney protection.

Do You Know Your Numbers?

This is the practical heart of the matter, and it's worth being specific about three separate numbers, because they measure different things and people often conflate them.

A1C reflects longer-term blood glucose control and is used to screen for and monitor diabetes and prediabetes.

eGFR, or estimated glomerular filtration rate, is calculated from a blood creatinine measurement and estimates how efficiently the kidneys are filtering blood. A lower eGFR generally indicates reduced kidney function.

uACR, the urine albumin-to-creatinine ratio, looks for albumin, a protein, leaking into the urine. Healthy kidneys generally keep albumin in the blood rather than letting it pass into urine, so its presence can be an early sign of kidney damage, sometimes appearing before eGFR changes at all.

Current clinical guidance reflects how important it is to check both eGFR and uACR together rather than relying on either alone. The American Diabetes Association's 2026 Standards of Care recommend assessing kidney function with both a random urine albumin-to-creatinine ratio and estimated glomerular filtration rate at least annually in people with type 1 diabetes of five or more years' duration, and in all people with type 2 diabetes regardless of treatment [2]. The rationale is straightforward: albuminuria and eGFR each carry independent information about risk, and at any given eGFR, the degree of albuminuria is associated with the risk of cardiovascular disease, CKD progression, and death [2].

So three questions worth bringing to your next checkup: Do you know your A1C? Do you know your eGFR? Have you had a urine albumin test? For most adults without diabetes or hypertension, routine kidney screening isn't automatically part of an annual physical, so this is a conversation to have directly with a clinician, especially for anyone with diabetes, high blood pressure, cardiovascular disease, or a family history of kidney disease.

What Jardiance Actually Has to Do With It

Jardiance, known generically as empagliflozin, belongs to a drug class called SGLT2 inhibitors. It was originally approved for type 2 diabetes and later for certain forms of heart failure. In September 2023, the FDA expanded its approval again, this time specifically to reduce the risk of sustained decline in eGFR, end-stage kidney disease, cardiovascular death, and hospitalization in adults with chronic kidney disease at risk of progression [3]. That approval was based on the EMPA-KIDNEY trial, a large, dedicated CKD study in which empagliflozin reduced the composite risk of kidney disease progression or cardiovascular death by 28 percent relative to placebo [4].

It's worth being precise about what that means. Jardiance is not a treatment for everyone with CKD, and it isn't a substitute for blood pressure control, glucose management, or other standard therapies. The manufacturer notes it is not recommended for CKD in patients with polycystic kidney disease or those requiring significant immunosuppressive therapy for kidney disease [5], since it isn't expected to help those populations. Like any prescription medicine, it carries its own risks and side effects, including dehydration and genital or urinary tract infections, and whether it's appropriate is a decision for a patient and their physician, not a conclusion to draw from a 30-second ad.

Back to Bolt

For most of his career, Usain Bolt spent every fraction of a second trying to go faster. His latest television role asks viewers to consider the opposite proposition: that sometimes slowing something down is itself a form of winning. Perhaps the most useful thing about the commercial isn't the drug being advertised at all. It's the question it leaves behind, the one easy to overlook amid the celebrity wattage: Do you know how well your kidneys are working? For nearly 9 in 10 people currently living with CKD, the honest answer is still no. That's the number worth changing.


Notes

1.     Centers for Disease Control and Prevention. Chronic Kidney Disease in the United States (updated March 2026). CS 363495-A.

2.     American Diabetes Association. 11. Chronic Kidney Disease and Risk Management: Standards of Care in Diabetes—2026. Diabetes Care, 2026.

3.     U.S. Food and Drug Administration approval announcement, via Eli Lilly and Company / Boehringer Ingelheim, "US FDA approves Jardiance for the treatment of adults with chronic kidney disease," September 22, 2023.

4.     Pharmacy Times. "FDA Approves Empagliflozin for Adults With Chronic Kidney Disease at Risk of Progression."

5.     Eli Lilly and Company. "US FDA approves Jardiance for the treatment of adults with chronic kidney disease" (full release, contraindication details), September 22, 2023.

6.     American Diabetes Association Professional Practice Committee. Screening for Chronic Kidney Disease in People with Diabetes (visual guide), April 2026.

7.     National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK). Kidney Disease Statistics for the United States.

8.     KDIGO. 2026 Clinical Practice Guideline for Diabetes and CKD (public review draft, March 2026).

9.     HCPLive. "Empagliflozin (Jardiance) Receives CKD Approval from US FDA."

10.Boehringer Ingelheim / Eli Lilly patient prescribing information, Jardiance Chronic Kidney Disease page.

Lessons from eDiscovery for RAG: Building Trustworthy Retrieval Systems


A follow-up to "What is RAG? Retrieval-Augmented Generation Systems with Archival Sources"

by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

Retrieval Is Not a New Problem

In my last article, I tested Retrieval-Augmented Generation (RAG) systems against archival sources and found that even well-engineered pipelines struggle with the basic question of whether they found the right material before they ever tried to answer a question. That struggle is not new. It has a name, a case history, and a professional discipline built around it: eDiscovery.

For three decades, litigators and their technical partners have grappled with a problem that RAG developers are only now rediscovering: how do you find the right documents inside an enormous, messy, heterogeneous collection, and how do you prove, after the fact, that you found them? In litigation, the collection is made of emails, contracts, scanned memos, spreadsheets, and chat logs, scattered across custodians and file formats. In RAG, the collection is made of chunked documents, embeddings, and vector indices. The materials differ. The underlying task, sorting a haystack for a small number of relevant needles under conditions of uncertainty, is the same.

This matters because the AI field tends to treat retrieval as a purely technical challenge: pick a better embedding model, tune the reranker, adjust top-k. Wu and colleagues' recent survey of RAG architectures catalogs exactly this kind of technical churn across retrieval pipelines, hybrid search strategies, reranking layers, and evaluation methods (Wu et al., 2026). What that survey does not resolve, and what most RAG engineering discussions skip entirely, is the governance question that eDiscovery has spent decades answering: how do you know your retrieval process is good enough to trust, and how do you demonstrate that to someone who was not in the room when you built it?

Precision versus Recall, Revisited

Every eDiscovery practitioner learns the same lesson early: precision and recall pull against each other, and the cost of getting the balance wrong is not abstract. Set the net too wide and reviewers drown in irrelevant documents, burning hours and budget on material that never should have surfaced. Set the net too narrow and you risk missing the one email that determines the outcome of a case. Cormack and Grossman's foundational work on Continuous Active Learning showed that recall-oriented review, iteratively refined through relevance feedback, could achieve high recall with a fraction of the manual review effort that exhaustive linear review required (Cormack and Grossman, 2015). That paper reframed retrieval not as a one-time query but as a process, one that improves through repeated rounds of human judgment and system adjustment.

RAG systems face a structurally identical tradeoff, just with different vocabulary. Top-k retrieval determines how many chunks get pulled before generation begins. Reranking determines which of those chunks survive to reach the model's context window. Filtering determines which sources are excluded before retrieval even starts. Elkiran and Rasheed's comparative study of retriever-reranker pairings makes the tradeoff explicit: different combinations of retrieval and reranking strategies produce measurably different quality and efficiency profiles, and no single configuration dominates across all use cases (Elkiran and Rasheed, 2026). That is precision versus recall wearing a different uniform. A RAG system tuned for maximum precision will answer confidently and sometimes wrongly, having missed the one chunk that would have changed the answer. A system tuned for recall will surface enough context to be safe, but risks diluting the model's attention with irrelevant material, an effect not unlike a review team wading through thousands of non-responsive documents in search of the few that matter.

The TREC Total Recall track exists precisely because "how much recall is enough" cannot be answered by intuition. It has to be measured, with defined test collections, defined relevance judgments, and reproducible scoring. RAG evaluation is only beginning to build equivalent infrastructure.

Chain of Custody versus Provenance

A litigator asks where a document came from. A historian asks about its provenance. A RAG developer asks which chunk produced a given answer. These are the same question asked by three different professions, and only one of those professions has built a rigorous, court-tested framework for answering it.

Chain of custody in eDiscovery is not paperwork for its own sake. It exists because a document's evidentiary value depends on being able to trace it back to its source without gaps: who collected it, when, from which custodian, and whether it was altered along the way. Knollmeyer and colleagues' review of RAG evaluation dimensions identifies citation quality and context relevance as core axes on which modern systems are judged (Knollmeyer et al., 2026), which is encouraging, but citation quality in most RAG systems today means little more than "a source was attached." It rarely means the source was verified, the retrieval path was logged, or the chunk boundaries were preserved in a way that lets a reviewer reconstruct exactly what the model saw.

This gap matters more as RAG systems move into higher-stakes domains. Zhang and colleagues' comparative study of retrieval strategies in clinical note extraction found that traceability between retrieved evidence and generated output is essential in a domain where an ungrounded claim carries real consequences (Zhang et al., 2026). The parallel to legal discovery is not a stretch. A clinician relying on a RAG system's summary of a patient's history needs the same assurance a litigator needs from a document review platform: an unbroken, inspectable link between the source material and the conclusion drawn from it.

Toward Defensible AI Retrieval

Legal discovery imposes four requirements on retrieval decisions: they must be explainable, reproducible, documented, and defensible under adversarial scrutiny. Those four words are worth sitting with, because RAG evaluation is converging on exactly the same standard, even if the field has not yet named it that way.

Tan and colleagues' work on provably robust aggregation addresses a version of the defensibility problem directly, introducing mathematically grounded methods for making retrieval resistant to corrupted or adversarial evidence (Tan et al., 2026). That is the RAG equivalent of a chain-of-custody challenge in court: what happens when the retrieved evidence itself cannot be trusted, whether through poisoning, duplication, or simple corpus noise? Cho and Lee's redundancy-aware evaluation framework tackles an adjacent problem that will feel familiar to any eDiscovery practitioner: high-similarity corpora, full of near-duplicate documents, distort both retrieval metrics and reviewer judgment if not handled explicitly (Cho and Lee, 2026). Legal collections are famously redundant, forwarded emails, cc'd threads, and near-identical drafts. RAG corpora, especially those built from scraped or aggregated sources, are no different, and the evaluation methods built for one domain transfer cleanly to the other.

I want to propose a working term for what this convergence is pointing toward: Defensible AI Retrieval. A defensible retrieval system is not simply one that performs well on a benchmark. It is one whose outputs can be explained after the fact, whose retrieval logic can be reproduced by a third party, whose decisions are documented well enough to survive scrutiny, and whose failure modes are known rather than discovered by accident. The Sedona Conference's TAR Case Law Primer spends hundreds of pages working through exactly this standard for legal technology, and its core insight, that defensibility is established through process documentation and validation, not through claims of accuracy alone, applies almost without modification to AI retrieval systems (Sedona Conference, 2023).

Quality Assurance: The Discipline RAG Has Not Yet Built

If there is one area where eDiscovery is unambiguously ahead of RAG, it is quality assurance. Legal review workflows are built around statistical validation from the outset: random sampling of the review population, second-level review of a subset of coded documents, dedicated privilege review passes, and formal error correction protocols. None of this is optional or aspirational. It is baked into the workflow because courts require it and because the cost of an undetected error, a privileged document produced by mistake, a responsive document never found, is high enough to justify the overhead.

RAG evaluation has started to build analogous practices, but the culture around them is still immature. Hallucination detection, groundedness scoring, and faithfulness metrics are now common terms in RAG papers. Knollmeyer and colleagues' framework treats faithfulness and context relevance as first-class evaluation dimensions alongside more traditional retrieval accuracy measures (Knollmeyer et al., 2026), and Sotic and Kamps' study of information-seeking behavior in RAG systems adds something the retrieval metrics literature often misses: how users actually interact with retrieved evidence, what they trust, what they miss, and where their mental models diverge from what the system actually did (Sotic and Kamps, 2026). That user-facing lens is closer to what eDiscovery calls second-level review, a human check on whether the system's output actually holds up.

What is still missing from most RAG evaluation pipelines is the statistical rigor that eDiscovery treats as baseline. Cormack and Grossman's 2016 paper on engineering quality and reliability in technology-assisted review argued that retrieval systems should be engineered for measurable quality, not tuned by intuition and then trusted (Cormack and Grossman, 2016). Random sampling with confidence intervals, documented error rates, and formal validation protocols are standard in eDiscovery and still rare in RAG deployment. A scoring rubric borrowed from legal QA might look like this for a RAG system under evaluation: sample a statistically significant subset of generated answers, score each on groundedness (is every claim traceable to a retrieved chunk), faithfulness (does the answer avoid contradicting its sources), completeness (did retrieval surface the material a human reviewer would consider necessary), and false-negative risk (what was missed, and how costly would that omission be in context). Reporting these as point estimates with confidence intervals, rather than as a single aggregate accuracy score, would bring RAG evaluation much closer to the standard eDiscovery already treats as routine.

Ten Lessons RAG Developers Should Borrow

1. Retrieval must be measurable. Intuition is not a validation strategy.

2. Every answer needs provenance. A citation that cannot be traced to its source chunk is not a citation.

3. Retrieval decisions should be explainable, not just accurate.

4. Sampling is mandatory. No system should be trusted on the basis of spot checks alone.

5. Human validation never disappears. Automation reduces review burden; it does not eliminate the need for oversight.

6. Metadata matters. Knowing where a chunk came from, and how it was processed, is as important as its content.

7. False negatives are often more dangerous than false positives. A missed document, or a missed fact, can be more damaging than an irrelevant one.

8. Evaluation should be continuous, not a one-time benchmark run before launch.

9. Retrieval quality should be statistically measured, with error rates and confidence intervals, not asserted.

10. Trust is earned through transparency, not claimed through marketing.

Conclusion

AI did not invent the challenge of trustworthy retrieval. Long before vector databases and large language models existed, courts demanded retrieval systems capable of finding, documenting, defending, and reproducing evidence, and an entire professional discipline grew up around meeting that demand. Retrieval-Augmented Generation (RAG) represents a genuine technological advance. It is not, however, an entirely new discipline, and treating it as one means reinventing quality assurance practices, defensibility standards, and precision-recall tradeoffs that the legal profession has already spent decades refining under adversarial, high-stakes conditions.

The future of trustworthy AI retrieval may depend less on inventing new evaluation methods from scratch than on rediscovering, and adapting, the hard-earned lessons of eDiscovery. The tools will keep changing. The underlying discipline, measurable, provenance-driven, statistically validated, and honest about its failure modes, does not need to be reinvented. It needs to be borrowed.

---

Notes

Cho, H., & Lee, J.-Y. (2026). RARE: Redundancy-aware retrieval evaluation framework for high-similarity corpora. ACL 2026. https://aclanthology.org/2026.acl-long.923

Cormack, G. V., & Grossman, M. R. (2015). Autonomy and reliability of continuous active learning for technology-assisted review. https://arxiv.org/abs/1504.06868

Cormack, G. V., & Grossman, M. R. (2016). Engineering quality and reliability in technology-assisted review. Proceedings of SIGIR 2016. https://plg.uwaterloo.ca/~gvcormac/PAPERSx.html

Elkiran, H., & Rasheed, J. (2026). Evaluating retriever reranker pairings in RAG based on quality and efficiency trade-offs. Discover Computing. https://doi.org/10.1007/s10791-026-10156-3

Knollmeyer, S., et al. (2026). Evaluating retrieval augmented generation: A comprehensive review of evaluation dimensions, question types, and application. SN Computer Science. https://link.springer.com/article/10.1007/s42979-026-05134-x

Redgrave LLP (2026). "Putting Gen AI to the Test: A Document Review Accuracy Study. Relativity aiR for Review vs. Active Learning". https://resources.relativity.com/document-review-accuracy-analysis-study-lp.html

The Sedona Conference. (2023). The Sedona Conference TAR case law primer (2nd ed.). https://www.thesedonaconference.org/node/10365

Sotic, B. N., & Kamps, J. (2026). Information seeking behavior in LLM-based RAG: Mental models and missing information. Proceedings of ACM SIGIR 2026. https://doi.org/10.1145/3805712.3808537

Tan, X., et al. (2026). PRA-RAG: Provably robust aggregation in retrieval-augmented generation against retrieval corruption. Findings of ACL 2026. https://aclanthology.org/2026.findings-acl.1794

Wu, S., et al. (2026). Retrieval-augmented generation for natural language processing: A survey. Artificial Intelligence Review. https://link.springer.com/article/10.1007/s10462-026-11605-7

Zhang, H., et al. (2026). Optimising clinical information extraction: A comparative study of retrieval-augmented generation techniques in clinical notes. Journal of Biomedical Informatics. https://www.sciencedirect.com/science/article/pii/S1532046426000778