Intro: A Gap Hiding in Plain Sight
Two literatures are growing quickly in 2026 and, so far,
growing apart. The first is the literature on large language model (LLM)
evaluation: rubric-driven frameworks for scoring factuality, faithfulness,
hallucination, and task completion across generative outputs. The second is the
literature on FOIA (Freedom of Information Act) automation: case studies and
vendor analyses describing how agencies are using AI to triage requests, search
collections, and draft redactions faster than manual review allows. Both bodies
of work are maturing on their own terms. Neither has, in any formal sense, met
the other.
That absence is notable given the setting. FOIA is not an
ordinary text-summarization domain. A FOIA summary sits downstream of a legal
disclosure decision. It distills a record that has already been searched,
reviewed, exempted, and redacted, and it typically becomes part of what a
requester, a court, an oversight body, or a journalist relies on to understand
what an agency did and why. If an AI-generated summary drops a redacted date,
blurs the line between what was withheld and what was released, or introduces
even a plausible-sounding inference not supported by the underlying record, the
error does not stay contained to a chatbot transcript. It can shape a public
record, an appeal, or a legal filing.
Despite that stakes profile, a search of the current
literature turns up remarkably little that combines the two threads. Papers on
LLM evaluation rubrics rarely mention FOIA. Papers on FOIA automation rarely
propose a scoring framework specific to summary quality. Almost none attempt a
rubric purpose-built for the FOIA summarization task, one that treats redaction
fidelity, exemption logic, and traceability back to the released record as
first-class evaluation criteria, on par with the factuality and hallucination
metrics borrowed from general-purpose LLM evaluation. That gap is the occasion
for this article: a proposed rubric for assessing AI-generated FOIA summaries,
and a look at how the leading purpose-built FOIA technology on the market
today, Relativity FOIA and Relativity aiR, fits into, and exposes, that same
gap.
Why FOIA
Summarization Needs Its Own Rubric
General-purpose summarization rubrics, the ones built for
news articles, meeting transcripts, or customer support tickets, optimize for
coherence, conciseness, and topical coverage. Those criteria matter for FOIA
summaries too, but they are not sufficient. A FOIA summary carries legal weight
that a meeting-notes summary does not. Three features of the FOIA context
justify a dedicated evaluation instrument:
·
Redaction is structural, not incidental.
A FOIA record is rarely handed over whole. Exemptions under the statute (and
analogous state public-records laws) remove specific categories of information,
such as personal privacy, law enforcement techniques, or deliberative process,
while leaving the rest disclosable. A summary that fails to distinguish
"this is what was released" from "this is what was withheld and
why" misrepresents the disclosure itself, not just the document.
·
Legal meaning is precise and often
reader-facing. Terms like "responsive," "exempt,"
"segregable," and the specific statutory exemption cited (e.g.,
Exemption 5, Exemption 7(C)) carry defined legal consequences. A summary that
paraphrases those terms loosely can change what a reader believes the agency
determined.
·
The audience includes adversarial and
oversight readers. Unlike an internal meeting summary, a FOIA summary may
be read by litigators, journalists, or Inspectors General specifically looking
for inconsistency between the summary, the redactions, and the underlying
record. The evaluation bar has to anticipate that scrutiny.
These features are why importing an off-the-shelf LLM
summarization rubric is inadequate on its own, and why FOIA-specific criteria,
including redaction and exemption handling, legal-meaning preservation, and
traceability to the released record, need to sit alongside more familiar
accuracy and hallucination metrics.
A
Proposed Rubric for AI-Generated FOIA Summaries
The rubric below is offered as a starting framework rather
than a finished instrument. It borrows its general shape, weighted criteria
that are each independently scorable, from the structured evaluation approaches
now common in LLM benchmarking, while grounding the specific criteria in FOIA
practice.
|
Criterion |
Weight |
|
Factual accuracy |
25% |
|
Completeness of
disclosed information |
20% |
|
Correct treatment
of redactions and exemptions |
15% |
|
Preservation of
legal meaning |
15% |
|
Hallucination rate |
10% |
|
Neutrality and
absence of editorial bias |
5% |
|
Citation/traceability
back to released records |
5% |
|
Readability and
organization |
5% |
Factual accuracy (25%) carries the largest weight because it is the
load-bearing criterion for every downstream use of the summary. Evaluators
should check whether every factual claim in the summary is verifiable against
the record as released, with particular attention to dates, names of programs
or offices (as distinct from names withheld under privacy exemptions), dollar
figures, and sequences of events.
Completeness of disclosed information (20%) asks
whether the summary captures the material substance of what was released, not
merely a plausible-sounding excerpt of it. A summary that is accurate but omits
a disclosed finding, a key admission, or a significant data point can be just
as misleading as one that fabricates something, because it gives a false
impression of a "clean" record.
Correct treatment of redactions and exemptions (15%)
is the most FOIA-specific criterion on the list. It asks whether the summary
clearly signals where information was withheld, whether it correctly reflects
the cited exemption basis where that basis is itself disclosed, and,
critically, whether it avoids inferring or reconstructing withheld content from
context. That kind of inference is a known failure mode of generative
summarization when a document has gaps.
Preservation of legal meaning (15%) evaluates whether
statutory and procedural terms are used correctly and whether the summary's
characterization of the agency's determination (responsive, non-responsive,
partially exempt, referred to another agency, etc.) matches the record's actual
disposition.
Hallucination rate (10%) is scored separately from
general factual accuracy because it targets a distinct failure mode: content
that has no basis anywhere in the source record, rather than content that is
merely imprecise. In FOIA summarization this includes fabricated exemption
rationales, invented dates, or synthesized "likely" content behind a
redaction.
Neutrality and absence of editorial bias (5%) checks
that the summary does not editorialize about the agency's conduct, insert
framing not present in the record, or adopt a tone that implies a conclusion
(such as wrongdoing, cover-up, or exoneration) the record itself does not
support.
Citation/traceability back to released records (5%)
asks whether a reviewer can trace each summary statement back to a specific
page, Bates number, or section of the released record. This matters enormously
for defensibility if the summary is later challenged.
Readability and organization (5%) is weighted lowest
deliberately. It matters for usability, but a highly readable summary that
fails the criteria above is a worse outcome than an awkward one that passes
them.
Applied consistently, a rubric of this kind would let
researchers and agencies compare AI-generated FOIA summaries across vendors,
models, and prompt designs using a shared, FOIA-native standard, something the
current literature does not yet offer.
Where the
Industry Stands: Relativity FOIA and Relativity aiR
Relativity is a useful lens for testing this gap against
real-world deployment, because it is one of the few vendors building AI
specifically into the FOIA review workflow rather than adapting a general
eDiscovery tool after the fact.
Relativity FOIA, announced as generally available
inside RelativityOne Government, unifies records request intake, case
management, AI-supported review, disclosure, and reporting into a single
connected workflow. The product's own framing makes clear that AI is meant to
assist rather than replace the reviewer: every prioritized record arrives with
a confidence score, a plain-language explanation, a counterpoint, and
supporting citations, and it is the human reviewer, not the model, who makes
the final call on each document. The accompanying launch announcement strikes
the same note, describing the platform as pairing defensible AI and automation
with Relativity's established review technology so that agencies can
standardize exemption rationale and accelerate review. That is a framing built
around assistance and consistency, not autonomous disclosure decisions.
Relativity aiR, the generative AI layer underlying
that workflow, does not market a standalone "summarization" feature
by that name, but its analysis outputs function as summaries in practice.
RelativityOne's government release notes describe a mid-2026 update to aiR for
Review that lets teams build custom analyses across both text and images,
producing what the release notes themselves describe as "insights,
extractions, and summaries." In other words, the document-level
explanations aiR for Review generates, the kind of output a FOIA reviewer would
read to understand a record before deciding what to release, are a
summarization capability in every functional sense, even without that specific
product label.
The philosophy behind these design choices is spelled out
most directly in Brian Thompson's commentary on government AI infrastructure.
Thompson argues that courts and oversight bodies require agency decisions to be
explainable, traceable, and defensible, and that AI can only meet that bar when
it is purpose-built for legal and public sector work in a way that preserves
audit trails and keeps a verifiable link between inputs and outputs, with human
validation built in rather than bolted on. That is a strong statement of
principle: transparency, defensibility, explainability, and reproducibility as
design requirements. But it remains a statement of principle rather than a
published measurement framework.
The clearest evidence of both the industry's evaluation
ambition and its current limits comes from an independent study Relativity has
publicized: Redgrave LLP's head-to-head comparison of aiR for Review against a
traditional active-learning managed review. The results were striking on the
metrics the study did measure. aiR for Review reached 88 percent recall against
64 percent for the active-learning workflow, and its elusion rate (responsive
documents wrongly left in the discard pile) came in at 1 percent versus 3
percent for the manual process, while consuming roughly 18 attorney-hours
against an estimated 1,123 hours for the 24-person manual review team. The
study is equally candid about tradeoffs: aiR for Review's precision, at 29
percent, trailed the active-learning workflow's 39 percent, a gap the authors
attribute in part to the low "richness," meaning the scarcity of
responsive documents, in the test population.
That study, however, measured a responsiveness-classification
task, finding documents relevant to a legal standard, not a summarization
task. It offers precision and recall figures for document identification,
not for the fidelity of AI-generated summaries. Reviewed against the rubric
above and against the broader FOIA-specific commentary, six gaps stand out in
Relativity's current public-facing material:
·
There is no formal, published FOIA summarization
evaluation rubric comparable to the one proposed here.
·
There are no benchmark scores specific to
AI-generated FOIA summaries, as distinct from the responsiveness-classification
benchmarks that do exist.
·
There is no hallucination testing specific to
FOIA summaries. The closest published work addresses document-level
responsiveness accuracy, not summary-level fidelity.
·
There is no disclosed human evaluation
methodology specific to FOIA summarization, meaning no description of how human
reviewers score AI summary outputs against a defined standard.
·
There are no precision/recall studies of
AI-generated summaries themselves, again as distinct from precision/recall for
document classification.
·
There is no academic validation paper describing
how Relativity measures FOIA summary quality specifically, as opposed to
review-workflow efficiency more broadly.
None of this is a criticism of the product design, which is
explicitly built around human validation, citation-backed outputs, and audit
trails, features that a rigorous FOIA summarization rubric would actually
reward under the "traceability" and "correct treatment of
redactions and exemptions" criteria. It is, instead, a description of an
open research space. The underlying architecture for defensible AI
summarization exists, but the formal instrument for measuring whether it
produces good FOIA summaries, as opposed to good responsiveness calls, does not
yet exist in the public record.
Implications
for Research and Practice
The rubric proposed here is deliberately modest in scope:
eight criteria, weighted to reflect FOIA's legal stakes, designed to be usable
by researchers, agency FOIA officers, and vendors alike. Its value lies less in
the specific weights, which should be stress-tested and adjusted through
empirical study, than in establishing that FOIA summarization deserves a
dedicated evaluation standard rather than an inherited one.
For researchers, the immediate opportunity is to pair this
kind of rubric with a labeled dataset of AI-generated FOIA summaries scored by
trained reviewers, producing the kind of benchmark and hallucination-rate
figures currently missing from the literature. For agencies and vendors, the
opportunity is to publish exactly the kind of validation study that Relativity
has already modeled for document responsiveness (blind expert review,
ground-truth comparison, transparent methodology) but aimed squarely at summary
quality rather than classification accuracy.
As FOIA request volumes climb and agencies lean further into
AI-assisted review, the gap between AI that helps reviewers understand records
and AI whose summaries have been rigorously measured against a FOIA-specific
standard is the gap this research agenda should close.
Sources
1. Relativity,
"Relativity FOIA: Purpose-Built FOIA Software for Federal and State
Agencies," June 2026. https://relativity.com/blog/relativity-foia-purpose-built-foia-software-for-federal-and-state-agencies
2. Relativity,
"Relativity Launches Relativity FOIA to Streamline Public Disclosure
Operations for Government Agencies," June 1, 2026. https://relativity.com/news-events/relativity-launches-relativity-foia-to-streamline-public-disclosure-operations-for-government-agencies
3. Relativity,
"What's New in RelativityOne Government (FOIA)." https://help.relativity.com/RelativityOne/Content/What_s_New/What_s_new_in_RelativityOne_Government.htm
4. Brian
Thompson, "AI, Operational Infrastructure, and the Future of Government
Legal Work," Relativity Blog, April 8, 2026. https://www.relativity.com/blog/ai-operational-infrastructure-and-the-future-of-government-legal-work
5. Robert
Keeling and Ray Mangum (Redgrave LLP), "Results Are In: 5 Lessons from
an Independent Study of aiR for Review," Relativity Blog, June 9,
2026. https://www.relativity.com/blog/results-are-in-5-lessons-from-an-independent-study-of-air-for-review
