Sunday, August 23, 2026

Who was Jason Arday?


When Scrutiny Becomes a Weapon: The Case Against How Jason Arday Was Treated

by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

Jason Arday spent his career studying inequality. In the final weeks of his life, he became a case study in it.

The 41-year-old sociologist, who in 2023 became the youngest Black person ever appointed to a professorship at Cambridge, resigned his post in early August 2026 after weeks of intensifying public scrutiny over alleged plagiarism in his 2015 doctoral thesis. Nine days later, he was found dead at his home in Battersea, south London. Police said the death was unexpected but not suspicious. His family later said he had endured a campaign of sustained abuse that had, in their telling, stretched on for more than three years before it finally consumed him. They asked the press to leave them alone.

There is a version of this story that treats it as a simple morality tale about academic fraud finally catching up with someone. That version is too convenient, and it leaves out almost everything that made this case what it became.

The allegations were never just about a thesis

Start with what actually happened once Nathan Cofnas, a self-described "race realist" who had already been dismissed from his own Cambridge affiliation over his writings on race, put his findings on Substack. The Times of London followed with an analysis of Arday's decade-old PhD thesis, and from there the story metastasized. Within a matter of weeks, one count found 289 separate articles about Arday; another tally put the number at 249 pieces across 15 national outlets. That is not the volume of coverage a plagiarism finding typically generates. It is the volume of coverage reserved for a public crucifixion.

And the coverage did not stay on the thesis. It moved into his childhood, his autism diagnosis, the charitable fundraising he had done, and the broader narrative of his life story. A man whose personal history of not speaking until age eleven and not learning to read until adulthood had once been treated as an inspiring footnote to his achievements was now having that same history picked apart as though it were evidence of a con. That shift, from auditing a document to auditing a person's entire biography and identity, is the tell. Legitimate scholarly accountability does not require re-litigating someone's childhood.

Consider who was doing the accusing, and why it matters

Cofnas did not arrive at this case as a neutral fact-checker. He has built a public career arguing that racial groups differ in intelligence, and he has written explicitly that under what he considers genuine meritocracy, institutions like Harvard and Cambridge would have very few Black professors. That is not incidental biography. It is the lens through which he approached Arday's record in the first place, and it is worth asking whether a white academic with an unremarkable public profile and the same thesis-level irregularities would have faced anything close to this level of exposure, or whether the intensity of the campaign was inseparable from the fact that Arday had become, in the public imagination, a symbol of Black academic advancement and diversity initiatives that Cofnas has spent years arguing against.

Ghent University itself seemed to sense this tension. When it suspended Cofnas as a precautionary measure and opened a preliminary disciplinary investigation, it did not challenge the underlying plagiarism claims. It said, carefully, that academic freedom carries responsibilities and can be limited when necessary to protect the rights of others, and it pointed specifically to its commitment to human dignity and its opposition to discrimination. A university does not open a discrimination investigation into someone for successfully flagging a legitimate citation problem. It opens one when the pattern of conduct surrounding the accusation looks like something else.

The institution that promoted him disappeared from the story

There is also the question of where Cambridge was in all of this. The university appointed Arday, promoted him, and held him up publicly as a landmark hire. If his academic record genuinely contained the problems now alleged, that is at least as much an indictment of Cambridge's own vetting, credentialing, and peer-review processes as it is of Arday personally. Instead, the institution that failed to catch these issues for a decade was allowed to recede into the background while a single Black professor absorbed the entirety of the public reckoning, his resignation treated as the natural end of the story rather than the beginning of a harder question about how he got there in the first place.

Two things can be true

None of this requires pretending Arday's scholarship was beyond reproach. He himself acknowledged that mistakes had occurred, even as he denied deliberately plagiarizing material. Investigating serious allegations against any academic, regardless of race, is legitimate, and being a symbol of diversity should never function as immunity from normal scholarly standards.

But proportionality matters, and proportionality is precisely what collapsed here. A citation problem in a ten-year-old thesis does not, on its own, generate hundreds of articles, a public autopsy of someone's disability and childhood, and a campaign led by a man who has built his reputation arguing that people who look like the accused do not belong in these positions to begin with. When the punishment for an alleged academic lapse becomes an unrelenting, identity-focused public campaign, the campaign itself becomes the story, and a serious institution has an obligation to ask whether race shaped not the fact of the scrutiny, but its ferocity.

---

This article makes the case for one interpretation of a genuinely contested and unresolved situation. Others, including Cofnas himself, argue that the plagiarism findings were substantive and long known to Cambridge insiders, that flagging misconduct is not disproportionate simply because the subject is well known, and that treating the accusation as racially motivated risks shielding real academic fraud from scrutiny. Ghent University's investigation into Cofnas has not concluded, and no independent inquiry has yet determined whether race materially affected how the case was pursued. Readers weighing this case should consider both sets of claims as the record develops.

My Next Four Reads


Yes, I know I have more than thirty books next to my nightstand waiting to be read. Here are four more that I am adding to the stack.

A Guardian and a Thief by Megha Majumdar, is the timeliest pick: a taut, 2025 National Book Award finalist set in a near-future Kolkata ravaged by climate change and food scarcity, following two families forced into conflict as they try to protect their children. It's also only 200 pages of tightly wound tension, so it won't demand a long commitment.

Bridge of Sighs, by Richard Russo, moves at the opposite pace: a patient, generous portrait of a small upstate New York town and the decades-long friendships that shape it.

Stone Yard Devotional by Charlotte Wood, rewards that same patience. It's a Booker-shortlisted Australian novel about a woman who abruptly leaves her life for a cloistered religious community, told through a journal-like structure as she reckons with unresolved grief and guilt. Quiet, but it stays with you.

Cursed Daughters is the most propulsive of the four, Oyinkan Braithwaite's follow-up to My Sister, the Serial Killer, about three generations of Nigerian women bound by a family curse and a daughter believed to be her dead aunt's reincarnation.

Together they cover very different registers: thriller-paced literary fiction, slow-burn small-town saga, contemplative memoir-in-fiction, and family-curse drama. Worth adding to the to the nightstand stack.

Saturday, August 22, 2026

Evaluating Historical Accuracy in LLM: Reconstruction, Jim Crow, and the Civil Rights Movement as High-Stakes Test Cases


by Alvin Blackshear  |  Historian & Researcher   <ablackshear@gmail.com>

Ask a large language model a seemingly straightforward question: "Who was the first African American elected to statewide office after Reconstruction?" The answer arrives instantly, fluently, and with the trappings of scholarship. It may cite books. It may offer dates with confident precision. It sounds, in short, like the work of someone who knows the field.

Historians know better. A question that appears simple on its surface is, in practice, layered with complications. What counts as "statewide office"? Does an elected superintendent of education count the same as a lieutenant governor? Does the answer depend on which state, since Reconstruction unfolded unevenly across the former Confederacy and border states? Does "first" mean first elected, first to take office, or first to serve a full term amid the violence and fraud that often interrupted Reconstruction-era elections? A model that produces a single, tidy name has already made a series of interpretive choices, usually without disclosing them, and sometimes without any grounding in verifiable sources at all.

This is the central argument of this article: historical accuracy cannot be measured merely by whether an isolated fact happens to be correct. Accuracy in historical writing is a compound of chronology, sourcing, context, and interpretation, and any evaluation of AI-generated history that stops at fact-checking a name or a date has already missed most of what makes history rigorous. Reconstruction, African American history, and the Civil Rights Movement make an unusually demanding proving ground for this kind of evaluation, and examining them closely offers historians a template for judging AI output more broadly.

Why These Three Fields Are an Ideal Stress Test for LLMs

Several features of these fields combine to expose the weaknesses of language models more sharply than most other historical subjects.

Scholarship on Reconstruction and the Civil Rights Movement has changed rapidly over the past two generations, as historians have moved from older, dismissive interpretations toward accounts that take Black political agency, local organizing, and grassroots activism seriously. A model trained on a broad mixture of older and newer sources may blend outdated framings with current scholarship without signaling the shift, producing a synthesis that never actually existed in the historiography.

These fields are also politically contested in ways that shape how information about them circulates online, in textbooks, and in public commentary. Terminology is inconsistent across time and place: the language used for racial categories, political offices, and organizations has shifted across the nineteenth and twentieth centuries, and even among historians writing in the same decade. Local, state, and federal histories overlap and sometimes contradict one another, particularly during Reconstruction, when a county-level fusion government might operate differently from the state legislature above it. Archival records are incomplete, and the voices of the enslaved, the formerly enslaved, and the disenfranchised are systematically underrepresented in the documentary record that survives.

Together, these characteristics force a model to reason about evidence rather than simply retrieve settled facts. That is precisely why they make such a useful stress test. A model that performs well on a straightforward question about, say, the dates of a well-documented battle may perform far worse when the underlying historical reality is genuinely contested, sparsely documented, or subject to ongoing revision.


What Does "Historical Accuracy" Actually Mean?


Before any evaluation can proceed, it helps to unbundle the concept of accuracy into its component parts.

Factual accuracy is the most familiar category and covers names, dates, and places. It is necessary but far from sufficient. Chronological accuracy concerns the sequence of events, their timing relative to one another, and the crucial distinction between cause and consequence, a distinction that language models frequently blur when they narrate history as a smooth chain of inevitabilities rather than a contested and contingent process.

Citation accuracy asks two separate questions: do the cited works actually exist, and, if they do, do they actually support the claim attributed to them? A hallucinated citation is an obvious failure, but a real citation attached to a claim it does not support is a subtler and in some ways more dangerous one, since it can survive a cursory check.

Contextual accuracy asks whether the historical setting has been correctly explained, not merely gestured at. Historiographical accuracy asks whether the answer acknowledges that historians disagree, and whether it fairly represents the range of interpretation rather than presenting one school of thought as settled consensus. Evidentiary accuracy, finally, asks whether every important claim can be traced back to documentary evidence, the standard that separates history from well-informed narrative.

A model can satisfy the first of these categories while failing all the others. That gap is where most of the risk in AI-generated history lives.


Common Failure Modes in African American History


Rather than cataloguing hallucinations as isolated errors, it is more useful to organize failures by the type of historical reasoning they violate.

Invented primary sources and fabricated quotations are the most conspicuous failures, and they are especially damaging in African American history, where authentic primary source material is already scarce and precious. A model that invents a plausible-sounding letter or speech does more than get a fact wrong; it pollutes a record that is already thin.

Composite biographies are a subtler failure, in which a model merges the experiences or achievements of two or more historical figures into a single, internally inconsistent account. Chronological compression, in which distinct decades are folded together as though they were a single moment, is common in narratives that span the long civil rights era, a period historians increasingly define as stretching from the 1930s through the 1970s rather than the traditional decade between 1955 and 1965.

Misidentified officeholders and geographic confusion, particularly the conflation of state and local officials, recur often enough in Reconstruction-era queries to be treated as a distinct category. False causal claims, in which one event is assumed to have directly produced another without the intervening political and social process that actually connected them, distort the texture of change over time. Presentism, the projection of modern assumptions and categories backward onto historical actors who did not share them, is perhaps the most pervasive failure of all, since it can infect an otherwise factually accurate answer with an interpretive frame the historical subjects themselves would not have recognized.


Reconstruction as an AI Evaluation Benchmark


Reconstruction is an unusually demanding benchmark because it compresses so much complexity into roughly a dozen years. A model addressing this period correctly must integrate multiple constitutional amendments, uneven state-by-state implementation, rapid and often reversed political change, contested definitions of citizenship and suffrage, shifting racial classifications, waves of violent resistance, and a series of short-lived governments that rose and fell within a single decade.

The documentary base compounds the difficulty. Freedmen's Bureau records, election returns, and the proceedings of state constitutional conventions are voluminous but unevenly digitized, unevenly indexed, and often contradictory on points of basic fact, such as the precise vote count in a contested election. Economic historians examining Reconstruction have shown just how tangled the causal picture becomes when one moves past the political narrative into the socio-economic effects of emancipation and federal enforcement, where multiple datasets, regional variation, and the halting reach of federal authority all interact in ways that resist simple summary. A model asked to explain "the economic effects of Reconstruction" without acknowledging this complexity is not simplifying for clarity; it is producing a distortion.


The Civil Rights Movement as a Verification Challenge


The Civil Rights Movement presents a different but related problem: the temptation toward a tidy, celebratory narrative organized around a small number of famous leaders. Language models trained on the volume of public-facing material about the movement tend to reduce decades of organizing to a handful of familiar names and set-piece events, at the expense of the local activism, particularly the labor organizing, legal work, and grassroots mobilization carried out largely by women and by people whose names never entered the national press.

This tendency shows up as compressed decades folded into a single narrative arc, confusion between overlapping organizations with similar acronyms and overlapping memberships, legal decisions misplaced in time or attributed to the wrong court, invented speeches, and misquoted newspapers. What makes this category of error especially important for historians to flag is that omission functions here as a form of distortion. A summary that mentions only the most famous figures and set-piece protests is not merely incomplete; it actively misrepresents how change actually happened, and it can be just as misleading as an outright fabrication, even though it contains no single false statement.


A Historian's Evaluation Rubric for LLM Output


Translating these concerns into a working tool for evaluation yields a rubric that can be applied consistently across outputs:

Criterion

Questions

Citation authenticity

Does every cited work exist?

Primary-source fidelity

Are quotations verifiable?

Chronology

Is the timeline internally consistent? 

Context

Is sufficient historical context provided?

Historiography

Are competing interpretations acknowledged?

Provenance

Can factual claims be traced?

Geographic precision 

Are jurisdictions correctly distinguished?

Uncertainty

Does the model acknowledge ambiguity?

Reproducibility

Can another historian repeat the verification?

Applied consistently, this rubric shifts the evaluator's attention away from spot-checking isolated facts and toward the structural qualities that make a historical account trustworthy. A response can score reasonably well on factual accuracy and still fail on historiography, provenance, or uncertainty, and it is precisely those failures that a fact-focused benchmark would miss.

Measuring Historical Accuracy Beyond Benchmarks

Most existing AI benchmarks evaluate factual recall, question answering, and multiple-choice accuracy, formats borrowed largely from standardized testing rather than from historical practice. Recent scholarly efforts have begun to push against this default. One 2026 benchmark built around the Chinese imperial examination tradition was designed specifically to test historical reasoning rather than simple retrieval, and its authors found that even the most capable models struggled once the task moved beyond fact lookup into the kind of analytical work professional historians actually perform. That finding is instructive: models can appear highly competent on narrow factual questions while still falling well short of the reasoning historians rely on day to day.


Historians evaluate evidentiary reasoning, provenance, interpretation, source criticism, contextualization, and the state of competing scholarship. These are fundamentally different evaluation targets than the ones most AI benchmarks were built to measure, and the gap between the two helps explain why a model can pass conventional accuracy tests while still producing what one recent essay memorably termed "stochastic history," an account that carries the surface texture of scholarship without the interpretive reasoning that actually produces it. Closing that gap will require benchmarks designed by historians, for historical reasoning specifically, rather than benchmarks adapted from unrelated domains.


Human-in-the-Loop Historical Verification


None of this argues for abandoning AI as a research aid. It argues for a disciplined workflow in which the historian remains the final authority. A workable sequence looks like this: the model generates a draft; the historian validates every citation; primary sources are independently verified; chronology is checked against the documentary record; competing interpretations are compared against current historiography; the historiographical framing itself is reviewed for balance; and only then does the material move toward publication.

In this workflow, the historian is not a passive consumer of AI output but an active evaluator standing between a plausible draft and a trustworthy account. That distinction, subtle as it sounds, is the difference between treating AI output as evidence requiring evaluation and treating it, mistakenly, as evidence in its own right.


Future Directions


Several developments now underway may narrow the gap between AI fluency and historical rigor. Retrieval-augmented generation systems, which ground model output in retrieved documents rather than internalized patterns, offer one promising path, as do archival retrieval systems built specifically for historical collections. Provenance-aware AI, designed to track and disclose the origin of every claim it makes, addresses the citation problem directly rather than leaving it to after-the-fact verification. Uncertainty estimation, which would allow a model to signal when the historical record is genuinely contested rather than presenting contested claims with false confidence, speaks directly to one of the recurring failures described above. Citation-grounded generation, historian-designed benchmarks, and dedicated historical evaluation datasets round out a research agenda that treats historical reasoning as its own discipline rather than a subset of general knowledge retrieval.

None of these tools eliminate the historian's role. If anything, they raise the bar for what that role requires: historians will need to evaluate not only the answers a retrieval-augmented system produces but the quality, provenance, and interpretive framing of the evidence it retrieves in the first place. A flawed archive fed into a well-engineered retrieval system still produces a flawed history.

Until AI systems consistently meet the standard historians already hold themselves to, historical accuracy will remain not a property inherent to the model, but the product of rigorous, sustained human evaluation.


Sources


"Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination." ACL 2026. https://aclanthology.org/2026.acl-long.1378

"Using Generative AI in Historical Practice." Cambridge University Press, 2026. https://www.cambridge.org/core/elements/using-generative-ai-in-historical-practice/7C1392A6E9DBD379FAA42E2D16A6D45B

"Writing Against the Machine: Computational Authorship and Historical Writing." History (Wiley), 2026. https://onlinelibrary.wiley.com/doi/10.1111/1468-229x.70092

"The Algorithm of Silence: Artificial Intelligence, Archival Bias and the Ethical Reconstruction of Digital Memory." Journal of Documentation, 2026. https://doi.org/10.1108/JD-08-2025-0249

"Automating the Past: Artificial Intelligence and the Next Frontiers of Digital History." International Journal of Humanities and Arts Computing, 2026. https://doi.org/10.3366/ijhac.2026.0361

"Reframing Historical Text Extraction: A Cross-Pathway Validation of OCR, LLM-Assisted Correction, and Direct Multimodal Transcription." Information (MDPI), 2026. https://www.mdpi.com/2078-2489/17/8/722

"Was Freedom Road a Dead End? Socio-economic Effects of Reconstruction in the American South." Economic History Review, 2026. https://onlinelibrary.wiley.com/doi/10.1111/ehr.70085

"Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models." 2026. https://arxiv.org/abs/2602.17433

Friday, August 21, 2026

What Exactly Is WISeR?


by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com> 

A Deeper Look at Medicare's AI Prior Authorization Pilot

A lot of you wrote in after “Is AI Deciding Who Gets Healthcare?” asking the same question: okay, but where, and for what?  Fair ask. Here are the details.

The Basics

WISeR stands for Wasteful and Inappropriate Service Reduction. It's a six-year model run by Centers for Medicare & Medicaid Services' (CMS) Innovation Center, announced in June 2025 and live since January 1, 2026. Unlike Medicare Advantage, where prior authorization has been standard for years, WISeR marks one of the first real tests of AI-assisted prior authorization inside original (fee-for-service) Medicare, a program that has historically had almost none of it. CMS frames the goal as trimming fraud, waste, and abuse. Critics frame it as outsourcing coverage denials to vendors whose paychecks depend on how much they deny.

Which States

WISeR is currently running in six states, split across four Medicare Administrative Contractor jurisdictions:

        New Jersey

        Ohio

        Oklahoma and Texas

        Arizona and Washington

CMS chose these based partly on which states' MAC jurisdictions and existing local coverage rules made the model easiest to test cleanly. There's no public timeline yet for expanding beyond these six, but several health-finance analysts expect that if the pilot runs smoothly, CMS won't wait out the full six years before widening it.

Which Procedures

For the initial phase, WISeR covers 14 selected Medicare Part B services, all elective rather than emergency or first-line treatments. The list centers on categories CMS has flagged as historically vulnerable to overbilling or overuse, including:

        Skin and tissue substitutes (grafts used in wound care)

        Electrical nerve stimulator implants

        Epidural steroid injections for pain management

        Knee arthroscopy for osteoarthritis

CMS has said it may add services to this list as the model progresses. Notably excluded: inpatient-only services, emergency services, and anything CMS considers too risky to delay.

Who's Actually Doing the Reviewing

This is the detail that surprises people most: the “model participants” aren't providers or insurers; they're technology companies. CMS selected six of them to operate the AI/ML-driven review process across the six states: Cohere, Genzeon, Humata, Innovaccer, Virtix, and Zyter.

Here's roughly how it works. A medical provider can either submit a request for prior authorization before performing the service, or skip that step and submit the claim for a post-service, pre-payment review instead. If they go the prior-authorization route, the AI vendor is supposed to return a determination (an affirmation or non-affirmation) within about three days. Any non-affirmation (denial) has to go through a licensed clinician before it's finalized; CMS says the AI itself can't issue the final “no.”

The Part That Worries Critics

The vendors' compensation is tied, at least in part, to the savings their reviews generate, meaning the more claims they flag as non-payable, the more they can earn, adjusted against other metrics like accuracy and provider experience. Revenue-cycle experts have publicly guessed that a meaningful share of claims routed through WISeR (some estimates run around 25%) could end up denied. That incentive structure is central to why members of Congress introduced the Ban AI Denials in Medicare Act, aiming to shut the model down before it expands further.

CMS is also piloting a “gold-carding” feature, expected to roll out further details later in 2026, that would exempt providers with consistently clean approval histories from future prior-authorization requirements, an attempt to reward compliance rather than blanket every claim with review.

The Takeaway

WISeR is narrower than my original article may have implied: six states, 14 services, six named vendors, with a human clinician required to sign off on any denial. But the financial incentives baked into the model, and the fact that this is the first large-scale AI prior-auth experiment inside traditional Medicare, are exactly why it's worth watching closely as it expands.

Sources

CMS, “WISeR (Wasteful and Inappropriate Service Reduction) Model”: https://www.cms.gov/priorities/innovation/innovation-models/wiser

CMS, “WISeR Model Frequently Asked Questions”: https://www.cms.gov/priorities/innovation/files/document/wiser-model-frequently-asked-questions

CMS, “WISeR Fact Sheet”: https://www.cms.gov/files/document/wiser-fact-sheet.pdf

CMS, “WISeR Model Provider and Supplier Operational Guide”: https://www.cms.gov/priorities/innovation/files/wiser-provider-supplier-guide.pdf

Nick Hut, “The WISeR prior authorization model for Medicare is set to pose challenges for hospitals,” HFMA (Dec. 23, 2025): https://www.hfma.org/revenue-cycle/the-wiser-prior-authorization-model-for-medicare-is-set-to-pose-challenges-for-hospitals/

ASRA Update, “CMS Provides More Details on WISeR Prior Authorization Model” (Oct. 24, 2025): https://asra.com/news-publications/asra-update-item/asra-updates/2025/10/24/cms-provides-more-details-on-wiser-prior-authorization-model

Dinsmore & Shohl, “Is AI WISeR? CMS Models AI-based Prior Authorization Process in Six States” (Jan. 30, 2026): https://www.dinsmore.com/publications/is-ai-wiser-cms-models-ai-based-prior-authorization-process-in-six-states/

Moss Adams, “Medicare WISeR Model Requires Prior Authorizations in Six States” (Jan. 13, 2026): https://www.mossadams.com/articles/2026/01/medicare-wiser-model

DLA Piper, “CMS launches WISeR Model: What providers need to know” (2026): https://www.dlapiper.com/en/insights/publications/2026/01/cms-wiser-model

Center for Medicare Advocacy, “Early Reports on WISeR Model Are Troubling” (Mar. 26, 2026): https://medicareadvocacy.org/early-reports-on-wiser-model-are-troubling/

Congressional Research Service via Congress.gov, “Overview of the Medicare Wasteful and Inappropriate Service Reduction (WISeR) Model” (May 12, 2026): https://www.congress.gov/crs-product/IF13133

Medical Economics, “WISeR spending or unneeded delays in health care? Prior authorizations, AI in Medicare prompt concerns” (July 15, 2026): https://www.medicaleconomics.com/view/wiser-spending-or-unneeded-delays-in-health-care-prior-authorizations-ai-in-medicare-prompt-concerns

Thursday, August 20, 2026

Is AI Deciding Who Gets Healthcare?


by Alvin Blackshear  |  Historian & Researcher   <ablackshear@gmail.com>

AI is now deciding who gets healthcare in six US states.” It is the kind of sentence built to travel fast, and for older Americans who have spent decades trusting that traditional Medicare would simply pay for what their doctor ordered, it lands like an alarm. So is it true? Is artificial intelligence actually deciding whether Medicare patients receive medical treatment, or does that framing exaggerate what the federal government is actually doing?

The honest answer sits in uncomfortable middle ground. The underlying story is substantially real. Medicare has launched an important and controversial experiment involving AI-assisted prior authorization. But describing this simply as “AI deciding who gets healthcare” leaves out critical details about human review, existing Medicare coverage standards, the narrow set of services involved, and the difference between a coverage determination and a physician’s decision to provide care. At the same time, legitimate concerns about delays, financial incentives, transparency, and algorithmic influence should not be waved away merely because a headline overstates the case.

What Exactly Is WISeR?

The program at the center of the controversy is the Wasteful and Inappropriate Service Reduction Model, known as WISeR, created by the Centers for Medicare & Medicaid Services (CMS). It began operating on January 1, 2026, with prior authorization requests starting January 5 and application to relevant services beginning January 15. The model is scheduled to run through December 31, 2031. It applies to selected services under Original Medicare, not Medicare Advantage, and it operates in six states: Arizona, New Jersey, Ohio, Oklahoma, Texas, and Washington. CMS says the purpose is reducing fraud, waste, abuse, and low-value medical care, and the program targets a defined list of procedures and services rather than all Medicare-covered healthcare. In July 2026, the Senate voted against an effort to overturn WISeR, allowing the pilot to continue.

This distinction matters: WISeR is not an AI system reviewing every Medicare patient’s medical care. It is a technology-assisted prior-authorization and prepayment-review system covering a specific set of services that CMS considers vulnerable to inappropriate utilization, fraud, waste, or abuse.

Is AI Actually Making the Decision?

This is the question at the heart of any honest fact check. CMS explicitly permits WISeR contractors to use technologies including artificial intelligence and machine learning. But CMS also states that appropriately trained clinicians participate in verifying whether requests satisfy Medicare’s coverage requirements. Three claims deserve to be separated. It is accurate to say AI is being used as part of Medicare’s prior-authorization process. It is potentially misleading to say AI independently determines whether a beneficiary may receive healthcare. The more precise formulation is that private contractors use AI and machine learning-assisted systems together with clinical review to determine whether certain services meet Medicare coverage requirements.

There is also a distinction often lost in the headlines: a Medicare coverage decision is not identical to a medical decision. A physician can recommend a procedure while Medicare separately determines it will not pay for it. For many patients, however, that distinction can feel largely theoretical. If Medicare declines to pay for an expensive treatment, the practical result is often that the patient cannot obtain it at all.

Why CMS Built the Program, and Why Critics Are Worried

CMS argues that wasteful healthcare spending represents a significant share of American medical expenditures, and that the services selected for WISeR already have established coverage criteria, generally involve elective rather than emergency treatment, carry safety concerns when performed unnecessarily, and have histories of fraud or inappropriate use. Emergency and inpatient-only services fall outside the program entirely. The government’s argument, in short, is not that AI should practice medicine. It is that technology can more efficiently determine whether a procedure satisfies existing Medicare rules before taxpayers pay for it.

Critics raise four distinct concerns. Prior authorization already generates substantial complaints within Medicare Advantage and private insurance, so extending it into traditional Medicare marks a meaningful policy shift. Proprietary AI systems create questions of explainability, since physicians and patients may not know precisely how a case was evaluated. The payment structure has also drawn scrutiny because contractors can benefit financially from savings generated by preventing expenditures, which critics argue could incentivize denials, even as CMS maintains that contractors are simply rewarded for catching genuinely wasteful spending. Finally, human review does not automatically eliminate algorithmic influence. Even when a clinician confirms a determination, the AI system may shape which cases get flagged and what information the reviewer ultimately sees, which raises a harder question: when does AI merely assist a decision, and when does it effectively shape it?

Are Patients Already Being Harmed?

This claim deserves particular caution. There are documented reports of authorization delays, and Senator Maria Cantwell (Washington-D) released a 2026 report citing Washington hospital data suggesting that patients undergoing WISeR-covered procedures experienced substantially longer processing times. But documented delays, denied payment, allegations of denied medically necessary care, and proof that AI itself caused measurable patient harm are not interchangeable claims. A responsible account should avoid concluding that AI is harming Medicare patients unless evidence establishes the full chain from algorithm to incorrect determination to delayed or denied treatment to actual harm. That evidence remains under active investigation.

The Transparency Problem

The Electronic Frontier Foundation has filed a federal Freedom of Information Act lawsuit seeking information about WISeR’s underlying AI systems, including details on algorithms, vendor arrangements, accuracy testing, bias testing, hallucinations, and ongoing monitoring. This creates an unusual situation: the federal government is testing AI-assisted healthcare decision-making while the information needed to independently evaluate those systems may not yet be public. That does not prove the algorithms are unsafe. It establishes that independent verification remains difficult.

Misinformation, or an Oversimplified Warning?

Calling the original claim simply “misinformation” would go too far. The essential facts are real: Medicare is testing technology-enabled prior authorization involving AI and machine learning in six states, private contractors are participating, and coverage determinations can affect patients’ access to particular procedures. But “AI is deciding who gets healthcare” is more sweeping than the evidence supports.

The more consequential story may be the quieter one. For decades, Americans have debated whether insurers and government bureaucracies should second-guess physicians’ medical judgment. Artificial intelligence has now introduced a new actor into that old dispute. The sharper question is not simply whether AI is denying healthcare. It is how much influence an algorithm should have over a decision that can determine whether a patient can actually afford to follow their doctor’s advice, and what standard of evidence the government should have to meet before allowing an opaque system to shape that decision at all.

References:

CMS, WISeR Provider and Supplier Operational Guide. https://www.cms.gov/priorities/innovation/files/wiser-provider-supplier-guide.pdf

CMS, WISeR Request for Applications. https://www.cms.gov/files/document/wiser-model-rfa.pdf

KFF, Examining the Potential Impact of Medicare’s New WISeR Model, February 10, 2026. https://www.kff.org/medicare/examining-the-potential-impact-of-medicares-new-wiser-model/

Congressional Research Service, “Overview of the Medicare Wasteful and Inappropriate Service Reduction (WISeR) Model,” updated May 12, 2026. https://www.everycrsreport.com/reports/IF13133.html

Federal Register, CMS, original WISeR implementation notice, July 1, 2025. https://www.federalregister.gov/documents/2025/07/01/2025-12195/medicare-program-implementation-of-prior-authorization-for-select-services-for-the-wasteful-and

Sen. Maria Cantwell, WISeR Snapshot Report, 2026. https://www.cantwell.senate.gov/imo/media/doc/wiser_snapshot_report.pdf

EFF v. CMS, federal complaint, filed March 25, 2026.