Showing posts with label FOIA. Show all posts
Showing posts with label FOIA. Show all posts

Monday, August 10, 2026

A Rubric for Assessing AI-Generated FOIA Summaries

by Alvin Blackshear | Historian & Researcher


Intro: A Gap Hiding in Plain Sight

Two literatures are growing quickly in 2026 and, so far, growing apart. The first is the literature on large language model (LLM) evaluation: rubric-driven frameworks for scoring factuality, faithfulness, hallucination, and task completion across generative outputs. The second is the literature on FOIA (Freedom of Information Act) automation: case studies and vendor analyses describing how agencies are using AI to triage requests, search collections, and draft redactions faster than manual review allows. Both bodies of work are maturing on their own terms. Neither has, in any formal sense, met the other.

That absence is notable given the setting. FOIA is not an ordinary text-summarization domain. A FOIA summary sits downstream of a legal disclosure decision. It distills a record that has already been searched, reviewed, exempted, and redacted, and it typically becomes part of what a requester, a court, an oversight body, or a journalist relies on to understand what an agency did and why. If an AI-generated summary drops a redacted date, blurs the line between what was withheld and what was released, or introduces even a plausible-sounding inference not supported by the underlying record, the error does not stay contained to a chatbot transcript. It can shape a public record, an appeal, or a legal filing.

Despite that stakes profile, a search of the current literature turns up remarkably little that combines the two threads. Papers on LLM evaluation rubrics rarely mention FOIA. Papers on FOIA automation rarely propose a scoring framework specific to summary quality. Almost none attempt a rubric purpose-built for the FOIA summarization task, one that treats redaction fidelity, exemption logic, and traceability back to the released record as first-class evaluation criteria, on par with the factuality and hallucination metrics borrowed from general-purpose LLM evaluation. That gap is the occasion for this article: a proposed rubric for assessing AI-generated FOIA summaries, and a look at how the leading purpose-built FOIA technology on the market today, Relativity FOIA and Relativity aiR, fits into, and exposes, that same gap.

Why FOIA Summarization Needs Its Own Rubric

General-purpose summarization rubrics, the ones built for news articles, meeting transcripts, or customer support tickets, optimize for coherence, conciseness, and topical coverage. Those criteria matter for FOIA summaries too, but they are not sufficient. A FOIA summary carries legal weight that a meeting-notes summary does not. Three features of the FOIA context justify a dedicated evaluation instrument:

·         Redaction is structural, not incidental. A FOIA record is rarely handed over whole. Exemptions under the statute (and analogous state public-records laws) remove specific categories of information, such as personal privacy, law enforcement techniques, or deliberative process, while leaving the rest disclosable. A summary that fails to distinguish "this is what was released" from "this is what was withheld and why" misrepresents the disclosure itself, not just the document.

·         Legal meaning is precise and often reader-facing. Terms like "responsive," "exempt," "segregable," and the specific statutory exemption cited (e.g., Exemption 5, Exemption 7(C)) carry defined legal consequences. A summary that paraphrases those terms loosely can change what a reader believes the agency determined.

·         The audience includes adversarial and oversight readers. Unlike an internal meeting summary, a FOIA summary may be read by litigators, journalists, or Inspectors General specifically looking for inconsistency between the summary, the redactions, and the underlying record. The evaluation bar has to anticipate that scrutiny.

These features are why importing an off-the-shelf LLM summarization rubric is inadequate on its own, and why FOIA-specific criteria, including redaction and exemption handling, legal-meaning preservation, and traceability to the released record, need to sit alongside more familiar accuracy and hallucination metrics.

A Proposed Rubric for AI-Generated FOIA Summaries

The rubric below is offered as a starting framework rather than a finished instrument. It borrows its general shape, weighted criteria that are each independently scorable, from the structured evaluation approaches now common in LLM benchmarking, while grounding the specific criteria in FOIA practice.

 Criterion

Weight

 Factual accuracy

25%

 Completeness of disclosed information

20%

 Correct treatment of redactions and exemptions

15%

 Preservation of legal meaning

15%

 Hallucination rate

10%

 Neutrality and absence of editorial bias

5%

 Citation/traceability back to released records

5%

 Readability and organization

5%


Factual accuracy (25%)
carries the largest weight because it is the load-bearing criterion for every downstream use of the summary. Evaluators should check whether every factual claim in the summary is verifiable against the record as released, with particular attention to dates, names of programs or offices (as distinct from names withheld under privacy exemptions), dollar figures, and sequences of events.

Completeness of disclosed information (20%) asks whether the summary captures the material substance of what was released, not merely a plausible-sounding excerpt of it. A summary that is accurate but omits a disclosed finding, a key admission, or a significant data point can be just as misleading as one that fabricates something, because it gives a false impression of a "clean" record.

Correct treatment of redactions and exemptions (15%) is the most FOIA-specific criterion on the list. It asks whether the summary clearly signals where information was withheld, whether it correctly reflects the cited exemption basis where that basis is itself disclosed, and, critically, whether it avoids inferring or reconstructing withheld content from context. That kind of inference is a known failure mode of generative summarization when a document has gaps.

Preservation of legal meaning (15%) evaluates whether statutory and procedural terms are used correctly and whether the summary's characterization of the agency's determination (responsive, non-responsive, partially exempt, referred to another agency, etc.) matches the record's actual disposition.

Hallucination rate (10%) is scored separately from general factual accuracy because it targets a distinct failure mode: content that has no basis anywhere in the source record, rather than content that is merely imprecise. In FOIA summarization this includes fabricated exemption rationales, invented dates, or synthesized "likely" content behind a redaction.

Neutrality and absence of editorial bias (5%) checks that the summary does not editorialize about the agency's conduct, insert framing not present in the record, or adopt a tone that implies a conclusion (such as wrongdoing, cover-up, or exoneration) the record itself does not support.

Citation/traceability back to released records (5%) asks whether a reviewer can trace each summary statement back to a specific page, Bates number, or section of the released record. This matters enormously for defensibility if the summary is later challenged.

Readability and organization (5%) is weighted lowest deliberately. It matters for usability, but a highly readable summary that fails the criteria above is a worse outcome than an awkward one that passes them.

Applied consistently, a rubric of this kind would let researchers and agencies compare AI-generated FOIA summaries across vendors, models, and prompt designs using a shared, FOIA-native standard, something the current literature does not yet offer.

Where the Industry Stands: Relativity FOIA and Relativity aiR

Relativity is a useful lens for testing this gap against real-world deployment, because it is one of the few vendors building AI specifically into the FOIA review workflow rather than adapting a general eDiscovery tool after the fact.

Relativity FOIA, announced as generally available inside RelativityOne Government, unifies records request intake, case management, AI-supported review, disclosure, and reporting into a single connected workflow. The product's own framing makes clear that AI is meant to assist rather than replace the reviewer: every prioritized record arrives with a confidence score, a plain-language explanation, a counterpoint, and supporting citations, and it is the human reviewer, not the model, who makes the final call on each document. The accompanying launch announcement strikes the same note, describing the platform as pairing defensible AI and automation with Relativity's established review technology so that agencies can standardize exemption rationale and accelerate review. That is a framing built around assistance and consistency, not autonomous disclosure decisions.

Relativity aiR, the generative AI layer underlying that workflow, does not market a standalone "summarization" feature by that name, but its analysis outputs function as summaries in practice. RelativityOne's government release notes describe a mid-2026 update to aiR for Review that lets teams build custom analyses across both text and images, producing what the release notes themselves describe as "insights, extractions, and summaries." In other words, the document-level explanations aiR for Review generates, the kind of output a FOIA reviewer would read to understand a record before deciding what to release, are a summarization capability in every functional sense, even without that specific product label.

The philosophy behind these design choices is spelled out most directly in Brian Thompson's commentary on government AI infrastructure. Thompson argues that courts and oversight bodies require agency decisions to be explainable, traceable, and defensible, and that AI can only meet that bar when it is purpose-built for legal and public sector work in a way that preserves audit trails and keeps a verifiable link between inputs and outputs, with human validation built in rather than bolted on. That is a strong statement of principle: transparency, defensibility, explainability, and reproducibility as design requirements. But it remains a statement of principle rather than a published measurement framework.

The clearest evidence of both the industry's evaluation ambition and its current limits comes from an independent study Relativity has publicized: Redgrave LLP's head-to-head comparison of aiR for Review against a traditional active-learning managed review. The results were striking on the metrics the study did measure. aiR for Review reached 88 percent recall against 64 percent for the active-learning workflow, and its elusion rate (responsive documents wrongly left in the discard pile) came in at 1 percent versus 3 percent for the manual process, while consuming roughly 18 attorney-hours against an estimated 1,123 hours for the 24-person manual review team. The study is equally candid about tradeoffs: aiR for Review's precision, at 29 percent, trailed the active-learning workflow's 39 percent, a gap the authors attribute in part to the low "richness," meaning the scarcity of responsive documents, in the test population.

That study, however, measured a responsiveness-classification task, finding documents relevant to a legal standard, not a summarization task. It offers precision and recall figures for document identification, not for the fidelity of AI-generated summaries. Reviewed against the rubric above and against the broader FOIA-specific commentary, six gaps stand out in Relativity's current public-facing material:

·         There is no formal, published FOIA summarization evaluation rubric comparable to the one proposed here.

·         There are no benchmark scores specific to AI-generated FOIA summaries, as distinct from the responsiveness-classification benchmarks that do exist.

·         There is no hallucination testing specific to FOIA summaries. The closest published work addresses document-level responsiveness accuracy, not summary-level fidelity.

·         There is no disclosed human evaluation methodology specific to FOIA summarization, meaning no description of how human reviewers score AI summary outputs against a defined standard.

·         There are no precision/recall studies of AI-generated summaries themselves, again as distinct from precision/recall for document classification.

·         There is no academic validation paper describing how Relativity measures FOIA summary quality specifically, as opposed to review-workflow efficiency more broadly.

None of this is a criticism of the product design, which is explicitly built around human validation, citation-backed outputs, and audit trails, features that a rigorous FOIA summarization rubric would actually reward under the "traceability" and "correct treatment of redactions and exemptions" criteria. It is, instead, a description of an open research space. The underlying architecture for defensible AI summarization exists, but the formal instrument for measuring whether it produces good FOIA summaries, as opposed to good responsiveness calls, does not yet exist in the public record.

Implications for Research and Practice

The rubric proposed here is deliberately modest in scope: eight criteria, weighted to reflect FOIA's legal stakes, designed to be usable by researchers, agency FOIA officers, and vendors alike. Its value lies less in the specific weights, which should be stress-tested and adjusted through empirical study, than in establishing that FOIA summarization deserves a dedicated evaluation standard rather than an inherited one.

For researchers, the immediate opportunity is to pair this kind of rubric with a labeled dataset of AI-generated FOIA summaries scored by trained reviewers, producing the kind of benchmark and hallucination-rate figures currently missing from the literature. For agencies and vendors, the opportunity is to publish exactly the kind of validation study that Relativity has already modeled for document responsiveness (blind expert review, ground-truth comparison, transparent methodology) but aimed squarely at summary quality rather than classification accuracy.

As FOIA request volumes climb and agencies lean further into AI-assisted review, the gap between AI that helps reviewers understand records and AI whose summaries have been rigorously measured against a FOIA-specific standard is the gap this research agenda should close.


Sources

1.     Relativity, "Relativity FOIA: Purpose-Built FOIA Software for Federal and State Agencies," June 2026. https://relativity.com/blog/relativity-foia-purpose-built-foia-software-for-federal-and-state-agencies

2.     Relativity, "Relativity Launches Relativity FOIA to Streamline Public Disclosure Operations for Government Agencies," June 1, 2026. https://relativity.com/news-events/relativity-launches-relativity-foia-to-streamline-public-disclosure-operations-for-government-agencies

3.     Relativity, "What's New in RelativityOne Government (FOIA)." https://help.relativity.com/RelativityOne/Content/What_s_New/What_s_new_in_RelativityOne_Government.htm

4.     Brian Thompson, "AI, Operational Infrastructure, and the Future of Government Legal Work," Relativity Blog, April 8, 2026. https://www.relativity.com/blog/ai-operational-infrastructure-and-the-future-of-government-legal-work

5.     Robert Keeling and Ray Mangum (Redgrave LLP), "Results Are In: 5 Lessons from an Independent Study of aiR for Review," Relativity Blog, June 9, 2026. https://www.relativity.com/blog/results-are-in-5-lessons-from-an-independent-study-of-air-for-review