Tuesday, September 08, 2026

Lonnie Bunch Is Retiring - Look at What Is Happening to Black Leadership in Washington


The Smithsonian secretary says Trump did not drive him from office. But his departure raises a larger question about the administration's campaign against DEI, Black leadership and representative government.

by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

On September 8, Lonnie G. Bunch III announced he will retire as secretary of the Smithsonian at the end of 2026, closing out nearly 38 years at the institution. Bunch became secretary in 2019, the first African American and first historian to hold the Smithsonian's highest office. Before that, he built the National Museum of African American History and Culture from almost nothing, guiding it to its 2016 opening on the National Mall.

His retirement lands at a remarkable moment. Since early 2025, the Trump administration has pushed to strip what it calls "improper ideology" from the Smithsonian, challenging exhibits on race, slavery and immigration, pressuring the institution's leadership, and, just days before Bunch's announcement, threatening to withdraw federal-agency support altogether. Bunch insists that pressure isn't why he's leaving. But he also says the moment is one "we're at a time when people are challenging that independence," and he has pledged to keep fighting for the Smithsonian's independence until his last day.

Take him at his word. His departure isn't proof of anything by itself. But it's an occasion to ask a bigger question: what is happening to Black leadership inside the institutions of the federal government?

A pattern, not an anecdote

Start counting names. Gen. Charles Q. Brown Jr., the second Black chairman of the Joint Chiefs of Staff, fired. Carla Hayden, the first Black person and first woman to serve as Librarian of Congress, fired. Gwynne Wilcox, the first Black woman on the National Labor Relations Board, removed. Alvin Brown, the only Black member of the National Transportation Safety Board, removed. Robert Primus, a Black member of the Surface Transportation Board, removed. Peggy Carr, a 35-year Education Department veteran and Black commissioner of the National Center for Education Statistics, removed. Lisa Cook, the first Black woman to serve as a Federal Reserve governor, targeted for removal, a move now tied up in litigation.

Rachel Maddow and other reporters have compiled versions of this list since the spring, and in one federal complaint, attorneys for a fired official argued that roughly three-quarters of Black officials serving at independent federal agencies had been removed under this administration. How many names does it take before we stop treating each firing as its own isolated story and start asking whether there's a pattern?

DEI as the operational language of the personnel story

The administration's public rationale is that diversity, equity and inclusion programs themselves constitute discrimination, and that dismantling them restores merit-based, race-neutral government. That's the stated case, and it should be reported accurately. But it's worth asking what "DEI" means operationally once that label gets applied to institutions, personnel decisions and historical exhibits alike. The real question isn't whether the acronym is popular or unpopular. It's this: when an administration defines efforts to broaden representation of historically excluded Americans as discrimination, who loses when those efforts disappear?

Beyond the famous names

The prominent firings may end up being less significant than what's happening to ordinary Black federal employees. Black Americans have long been overrepresented in the federal workforce relative to their share of the population, and not by accident: federal jobs offered access to professional careers, stable pay, pensions and advancement at a time when much of the private economy discriminated against them.

Now consider the scale of the current downsizing. According to the Government Accountability Office, the federal workforce across 22 major agencies fell by nearly 256,000 employees (from about 2.27 million to 2.01 million) between December 2024 and January 2026, the product of roughly 378,000 separations against about 127,000 new hires. That reframes the story. It's no longer just "Trump fired Black leaders." It's a question about what's happening to Black participation in the federal government as a whole.

Inside the Pentagon

The military deserves its own look, because the evidence there is distinct. Gen. Brown's removal as Joint Chiefs chairman was followed by other senior Black military departures and a broader anti-DEI restructuring at the Pentagon. Reporting has also described Black and female officers being disproportionately dropped from promotion lists, and a Pentagon purge of DEI-related content that swept up material on the Tuskegee Airmen, Jackie Robinson and other minority military figures, some of it later restored after public criticism. That raises a sharper question than "Is the Pentagon eliminating DEI programs?" It's whether the campaign is reshaping who advances, who leads, and which chapters of American military history the government is willing to commemorate.

An old American question

America has been here before. When Woodrow Wilson took office in 1913, Black Americans had built a real foothold in federal employment. His administration segregated federal workplaces; Black employees were separated, reassigned, demoted and dismissed. The point isn't that Trump is Wilson. The point is that federal employment can be expanded or contracted as an avenue of Black opportunity through administrative power alone, with no new law required.

Representative government, not just a demographic count

Representation isn't only about whether Congress looks like the country. It's also about whether the institutions that exercise government power (the military, the Federal Reserve, regulatory commissions, the federal courts, the civil service, libraries, museums, scientific and education agencies, national cultural institutions) reflect the population they govern. What happens to representative government when the people making decisions inside it increasingly stop resembling the people governed by it? That's a question about legitimacy and institutional access, not just headcounts.

Back to Bunch, and who tells the story

Which brings us back to Bunch, who is more than an opening anecdote. The fight over Smithsonian exhibits dealing with slavery and race isn't separate from the personnel story; it's the other half of it. The administration has accused the institution of presenting an insufficiently celebratory account of the country; Bunch has defended scholarship that presents accomplishments alongside failures. There are two questions running through this moment: who gets to hold positions of federal authority, and who gets to tell America's history once they're there? Bunch, the Smithsonian's first Black secretary and the historian who built its African American history museum, sits at the intersection of both.

What the evidence supports

I don't believe these developments are coincidental, and I don't think "DEI" fully explains what we're witnessing. When Black leaders are repeatedly removed, Black federal workers disproportionately absorb the consequences of downsizing, programs meant to broaden participation get relabeled as discriminatory, Black military advancement is disrupted, and institutions built to tell Black history come under federal pressure, racial motivation has to be considered as a possible explanation for the pattern. That's a different claim than saying any individual official was fired because he or she is Black; it says the cumulative project and its effects are evidence from which that motivation can be debated. The precise extent to which racial animus drove any single decision remains uncertain. That's not the same as saying there's no evidence to examine.

There's a further question worth sitting with: if the administration succeeds in linking Black advancement itself to "DEI," does every Black official become vulnerable to the assumption that they represent diversity rather than merit? That may be one of the most consequential effects of this campaign, and it deserves investigation rather than assumption.

Lonnie Bunch says Donald Trump didn't push him out, and there's no reason to doubt him. But his retirement is a fair moment to look around Washington and take stock. Who occupies the government's senior offices now? Who is leaving the federal workforce? Who advances through the military? Which histories can still be told, and who decides? Every president dismisses officials; that's not the question. The question is whether the federal government is being systematically reshaped in ways that diminish Black representation, Black institutional authority and Black influence over the American story.

Monday, September 07, 2026

The Rule Doesn't Have to Become Law to Change Higher Education


How the threat of federal punishment can dismantle programs for minority students before a court ever decides whether the government has the power to do it.

by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

Here is a fact that should be reassuring: on September 4, 2026, the Treasury Department and the IRS published a proposed regulation, not a final one. Nothing has actually changed. Nothing is yet required.

Now turn that fact upside down. Colleges do not have the luxury of pretending REG-119986-25 doesn't exist. The administration has explicitly warned that private educational institutions maintaining what it regards as racially discriminatory programs could lose their federal tax-exempt status, a penalty severe enough to end most private schools outright. Treasury itself estimates that as many as 18,000 private educational institutions could be affected. The rule does not have to become law to change higher education. The threat may be enough.

What is being proposed

The proposal covers admissions, scholarships and loans, athletics, and other school-supported programs. It would treat any use of race, color, or national or ethnic origin in distributing benefits as disqualifying, even when the purpose is explicitly remedial. Schools could still target assistance using income, geography, first-generation status, individual hardship, and other race-neutral criteria. Crucially, the regulation would apply only to taxable years beginning after May 31, 2027. That gap between now and then is where the real story lives.

Anticipatory compliance

Imagine yourself as a university president, trustee, or general counsel. Do you maintain a scholarship established specifically for Black students and risk an eventual confrontation with the IRS? Or do you quietly rewrite the eligibility requirements now, while no one is forcing you to?

For a risk-averse institution, the rational response is to comply before anyone has required compliance. A Black student scholarship becomes an "economically disadvantaged" scholarship. A minority mentoring program opens to everyone. Targeted recruitment changes. Donors are gently discouraged from establishing race-conscious funds in the first place. No IRS agent has to show up. No exemption has to be revoked. No judge has to rule on anything. The program simply disappears, quietly, as a matter of institutional self-preservation.

How power like this operates

It is a fact that the administration proposes treating race-conscious remedial programs as disqualifying discrimination. Whether that reflects hostile intent is a separate question, one this argument doesn't need to resolve. What matters is how the power functions: government need not command a result directly to produce it. It can identify a category of conduct as risky, attach an extraordinary financial consequence to it, and let institutions, lawyers, and administrators do the rest. American racial policy has often worked this way, through funding conditions, tax treatment, and the threat of losing government benefits, long before any court weighs in. The pressing question isn't only "will this regulation survive review?" It's "what will disappear while we're waiting to find out?"

The strongest counterargument

Supporters can fairly respond that the administration isn't banning help for disadvantaged students. Schools can still aid poor, first-generation, and geographically disadvantaged students; Treasury's own framing insists on this. What's demanded is that aid follow individual disadvantage, not race. That's a coherent principle. But it invites a historical question worth sitting with: can a race-neutral remedy fully repair an inequality that was created through explicitly race-conscious discrimination? Reasonable people disagree, and this piece won't settle it for them.

The question nobody can answer

Congress has noticed. Representatives Lloyd Doggett and Terri Sewell have introduced the PROOF Act, aimed at guaranteeing due process before the IRS can strip an organization's tax-exempt status. It's a meaningful check, but it doesn't touch the underlying rule, and it does nothing to stop an administrator today from asking, "why risk an IRS examination at all? Just change the program now."

Which returns us to the real stakes. Courts may eventually uphold this regulation. They may strike it down. But by then, the more important consequence may have already occurred: scholarships rewritten, programs eliminated, recruitment quietly redirected, donors steered elsewhere. If universities voluntarily dismantle these programs in anticipation of a rule, and courts later strike that rule down, how many of those programs will ever come back?


For further reading:

Federal Register, REG-119986-25, "Racial Nondiscrimination in Private Schools" (Sept. 4, 2026).
https://www.federalregister.gov/documents/2026/09/04/2026-18127/racial-nondiscrimination-in-private-schools

U.S. Department of the Treasury, press release on the proposed rule (Sept. 3, 2026).  https://home.treasury.gov/news/press-releases/sb0621/

H.R. 10258, the PROOF Act, introduced Sept. 3, 2026.
https://www.congress.gov/bill/119th-congress/house-bill/10258

Wednesday, September 02, 2026

After Affirmative Action: What the Data Says About Who Really Gets an Edge


by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

A Sad Day, Not a Surprising One

When the Supreme Court struck down race conscious admissions in Students for Fair Admissions v. Harvard on June 29, 2023, the reaction from educators and advocates was less shock than grief. As one campus diversity officer put it in the ruling's immediate aftermath, it was "a sad day in America, but not a surprising day." That framing captures something real. Affirmative action's opponents had been building toward this moment for a decade, and those working in higher education had time to brace for it, even if bracing didn't make the outcome easier to absorb.

The Argument That Won't Go Away. Legacy, Donors, and Athletes

The sharpest counterpunch from critics of the ruling is that it eliminated one narrow preference while leaving much larger ones intact. A 2026 study published in Education Finance and Policy by Vanderbilt researchers Brent Evans and Cody Christensen examined seven institutions and one statewide policy that banned legacy admissions preferences, including Amherst College, Johns Hopkins, the University of California system, and the state of Colorado. Their finding complicates the simple version of this argument. Banning legacy preferences alone did not consistently increase student diversity. Some institutions saw real gains, others saw little to no change, in part because many state policies ban legacy preference but leave donor preference fully intact and rarely include enforcement mechanisms. The takeaway is not that legacy and donor preferences are harmless. It's that removing them is not, by itself, a substitute for the access affirmative action provided. Older data still underscores the scale of the underlying preference. A widely cited, but dated, Harvard admissions analysis found that more than 43 percent of white students admitted between 2009 and 2014 were recruited athletes, legacies, on the dean's interest list, or children of faculty and staff, compared with under 16 percent for Black, Asian American, and Hispanic admits.

The Precedent Nobody Should Ignore. What Happened After California Banned Race Conscious Admissions

Critics point to California as a preview of the nationwide ruling's likely effects, and researchers are still tracking the fallout from the newer, national version of that experiment. In a 2026 Brookings research brief, University of Maryland professor Julie Park (whose book Race, Class, and Affirmative Action was published by Harvard Education Press this year and reviewed in the peer-reviewed journal Education Review) documents what she calls a "cascade effect." Underrepresented students turned away from elite institutions are enrolling instead at state flagships, and students previously headed to flagships are being displaced further down the selectivity ladder. Park's analysis found that Black student enrollment fell at a majority of the 29 elite institutions she tracked, with 16 public flagships reporting a net loss of Black students. At the same time, 83 percent of public flagship institutions saw overall underrepresented minority enrollment rise, driven largely by Latino students rather than Black students. This is not a story of straightforward decline. It's a redistribution, and Park's research draws heavily on earlier causal work by Princeton economist Zachary Bleemer showing that students who lose access to more selective institutions tend to have measurably worse graduation rates, grades, and postgraduate earnings than they would have had otherwise. That earlier research, on California's Proposition 209, found the ban deterred more than 1,000 underrepresented minority applicants a year from even applying to the UC system and cut Black and Latino enrollment at UC Berkeley by roughly 40 percent.

Race Neutral Alternatives Aren't Neutral in Effect

The practical question facing universities now is what to do instead. Park's 2026 analysis notes that class based, income based alternatives to race conscious admissions do not reliably produce the same racial diversity gains, even when they succeed at increasing economic diversity. That gap is the empirical heart of the argument that "colorblind" admissions policies are not simply neutral substitutes. They tend to produce smaller and less consistent results.

Where This Leaves the Debate

None of this settles the constitutional question the Court already resolved. But it sharpens the practical one facing universities today. If the preferences that most favor wealthy, disproportionately white applicants remain harder to dislodge than expected, and if the race neutral alternatives on offer are inconsistent at best, the real fight ahead is over legacy and donor admissions reform, income based alternatives, and pipeline investment, not just compliance with the letter of the ruling.

------------------

Notes

Arcidiacono, P., Kinsler, J., & Ransom, T. (2022). "Legacy and Athlete Preferences at Harvard." Journal of Labor Economics, 40(1), 133–156. https://www.journals.uchicago.edu/doi/abs/10.1086/713744

Evans, B. J., & Christensen, C. L. (2026). "The Evolving Landscape of Legacy Preference Bans in Postsecondary Admissions. Evidence and Policy Implications from Case Studies." Education Finance and Policy, 21(3), 564–585. https://doi.org/10.1162/EDFP.a.433

Park, J. J. (2026). Race, Class, and Affirmative Action. College Admissions in a New Era. Harvard Education Press. Reviewed by Yingyuan Sun (2026), Education Review, 33. https://doi.org/10.14507/er.v33.4653

Thursday, August 27, 2026

The Storm We Keep Explaining Away - American Totalitarianism

by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

Societies rarely recognize the moment they cross from ordinary politics into something more dangerous, because the transition rarely looks dramatic from the inside. It looks like business continuing, laws being debated, elections being held, until one day the machinery that was supposed to check power has quietly stopped working. That is the situation worth naming plainly: America is drifting toward a fusion of concentrated private wealth and authoritarian political power, and the comfortable assumption that it cannot happen here is itself one of the conditions that allows it to happen.

Hannah Arendt spent much of her career studying how totalitarian regimes actually formed, not just how they governed once entrenched. In The Origins of Totalitarianism, published in 1951, she argued that such movements take root primarily through social conditions that precede any specific ideology. Large numbers of people feel politically homeless, economically insecure, and cut off from any sense that existing institutions represent their interests. A movement, unlike a conventional political party, does not need to offer coherent policy. It offers identity, belonging, and an enemy. People do not have to be persuaded by a doctrine so much as relieved of their isolation.

One of Arendt's more distinctive arguments concerned the relationship between power and truth. She observed that authoritarian control depends less on citizens believing official lies than on citizens losing confidence that any shared truth is knowable at all. When public statements shift constantly and contradict themselves without consequence, people gradually lose the capacity to compare claims against evidence and act accordingly. This, in her view, was more corrosive than any single falsehood. A public that has given up trying to distinguish fact from fiction becomes governable by whoever asserts things with the greatest confidence, regardless of accuracy.

This is where Arendt's framework, built to explain state-driven totalitarianism in the twentieth century, illuminates something more particular to the present moment: a pathway toward authoritarian outcomes that runs through concentrated economic power rather than a uniformed movement or a single dominant party. The pattern is structural rather than personal. Media ownership has consolidated into fewer hands, narrowing the range of information most people encounter day to day. Political and commercial interests increasingly overlap, making it harder to locate who is actually accountable for a given decision. Public loyalty organizes itself around individuals and brands more readily than around laws or institutions, which erodes the slower, less satisfying work of civic accountability. None of this requires tanks in the street. It only requires enough people deciding that noticing the pattern is either too exhausting or too partisan to bother with.

It is worth taking seriously the objection that comparisons to twentieth century totalitarianism can be overused, flattening real historical differences into a rhetorical alarm bell. American constitutional structures, an independent judiciary, federalism, and a free press still function in meaningfully different ways than their counterparts did in Germany by the early 1930s. There is also a real risk that alarm becomes its own kind of performance, a way of feeling politically engaged without doing the harder work of participating in civic life.

Even granting those differences, though, Arendt's deeper argument was never really about matching one historical regime against another point for point. It was about the psychological and social conditions that make populations available for authoritarian capture in the first place: isolation, exhaustion, and a creeping distrust that shared reality exists at all. Those conditions do not require the collapse of formal democracy to take hold. They only require that people stop noticing them.

Arendt's own answer to this danger was not despair but a fairly unglamorous kind of faith. She continued to praise small, self-organized civic bodies, ordinary people speaking and acting together in public, as the most reliable defense against drift. It is not a dramatic remedy, and it offers no guarantees. But it remains, as it did for her, the only remedy grounded in something more durable than hope.

A recent article by David Denby, offers a prescient conclusion, in “The Origins of Totalitarianism,” Arendt observes that bureaucratic despotism hollows out legislatures: parliaments cease making laws, and executives rule by decree. She describes how inflation, unemployment, and the breakdown of stable classes leave masses uprooted, isolated, and possessed by a pervasive hatred of the existing world. A movement—not a party—organizes their discontent, offering fictitious total explanations for real and unresolved problems. At its center stands an infallible leader: his words cannot be questioned, and his lies are turned into a functioning reality. Members must accommodate themselves to its fictions, with varying mixtures of gullibility and cynicism; followers shown proof that the leader has lied may admire him all the more for his tactical cleverness. He and his cadre insult and dehumanize outsiders, who are then stripped of their rights, rendered stateless, and herded into internment camps.

The ultimate object of these fictions is disorientation, more than belief. By continually altering reality, a government can disable the faculties of judgment and action on which political freedom depends. In her last known interview, recorded for French television in 1973, Arendt returned to the subject of public lying. According to a transcript published after her death, here is what she said:

If everyone always lies to you, the consequence is not that you believe the lies, but that no one believes anything at all anymore—and rightly so, because lies, by their very nature, have to be changed, to be “re-lied,” so to speak. So a lying government which pursues different goals at different times has constantly to rewrite its own history. That means that the people are deprived not only of their capacity to act, but also of their capacity to think and to judge. And with such a people you can then do what you please.

------------------

Further Reading

Hannah Arendt, The Origins of Totalitarianism (1951), for the full argument on how totalitarian movements form and how truth erodes under authoritarian pressure.
https://www.penguinrandomhouse.com/books/770207/hannah-arendt-the-origins-of-totalitarianism-expanded-edition-loa-389-by-hannah-arendt--jerome-kohn-and-thomas-wild-editors

Hannah Arendt, Eichmann in Jerusalem: A Report on the Banality of Evil (Viking, 1963), for her account of how ordinary participation, more than ideological fervor, sustains atrocity.

Elisabeth Young-Bruehl, Hannah Arendt: For Love of the World (Yale University Press, 1982), the standard biography, for readers who want the life behind the theory.

Jeremy Waldron, "What Would Hannah Say?" (2007), for a serious scholarly pushback against reflexively invoking Arendt for every political crisis.
https://www.nybooks.com/articles/2007/03/15/what-would-hannah-say

Denby, David. "Hannah Arendt's American Education." The New Yorker, August 17, 2026.
https://www.newyorker.com/magazine/2026/08/17/hannah-arendt-life-of-the-mind-thomas-meyer-book-review-an-admirable-woman-arthur-cohen

Sunday, August 23, 2026

Who was Jason Arday?


When Scrutiny Becomes a Weapon: The Case Against How Jason Arday Was Treated

by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com>

Jason Arday spent his career studying inequality. In the final weeks of his life, he became a case study in it.

The 41-year-old sociologist, who in 2023 became the youngest Black person ever appointed to a professorship at Cambridge, resigned his post in early August 2026 after weeks of intensifying public scrutiny over alleged plagiarism in his 2015 doctoral thesis. Nine days later, he was found dead at his home in Battersea, south London. Police said the death was unexpected but not suspicious. His family later said he had endured a campaign of sustained abuse that had, in their telling, stretched on for more than three years before it finally consumed him. They asked the press to leave them alone.

There is a version of this story that treats it as a simple morality tale about academic fraud finally catching up with someone. That version is too convenient, and it leaves out almost everything that made this case what it became.

The allegations were never just about a thesis

Start with what actually happened once Nathan Cofnas, a self-described "race realist" who had already been dismissed from his own Cambridge affiliation over his writings on race, put his findings on Substack. The Times of London followed with an analysis of Arday's decade-old PhD thesis, and from there the story metastasized. Within a matter of weeks, one count found 289 separate articles about Arday; another tally put the number at 249 pieces across 15 national outlets. That is not the volume of coverage a plagiarism finding typically generates. It is the volume of coverage reserved for a public crucifixion.

And the coverage did not stay on the thesis. It moved into his childhood, his autism diagnosis, the charitable fundraising he had done, and the broader narrative of his life story. A man whose personal history of not speaking until age eleven and not learning to read until adulthood had once been treated as an inspiring footnote to his achievements was now having that same history picked apart as though it were evidence of a con. That shift, from auditing a document to auditing a person's entire biography and identity, is the tell. Legitimate scholarly accountability does not require re-litigating someone's childhood.

Consider who was doing the accusing, and why it matters

Cofnas did not arrive at this case as a neutral fact-checker. He has built a public career arguing that racial groups differ in intelligence, and he has written explicitly that under what he considers genuine meritocracy, institutions like Harvard and Cambridge would have very few Black professors. That is not incidental biography. It is the lens through which he approached Arday's record in the first place, and it is worth asking whether a white academic with an unremarkable public profile and the same thesis-level irregularities would have faced anything close to this level of exposure, or whether the intensity of the campaign was inseparable from the fact that Arday had become, in the public imagination, a symbol of Black academic advancement and diversity initiatives that Cofnas has spent years arguing against.

Ghent University itself seemed to sense this tension. When it suspended Cofnas as a precautionary measure and opened a preliminary disciplinary investigation, it did not challenge the underlying plagiarism claims. It said, carefully, that academic freedom carries responsibilities and can be limited when necessary to protect the rights of others, and it pointed specifically to its commitment to human dignity and its opposition to discrimination. A university does not open a discrimination investigation into someone for successfully flagging a legitimate citation problem. It opens one when the pattern of conduct surrounding the accusation looks like something else.

The institution that promoted him disappeared from the story

There is also the question of where Cambridge was in all of this. The university appointed Arday, promoted him, and held him up publicly as a landmark hire. If his academic record genuinely contained the problems now alleged, that is at least as much an indictment of Cambridge's own vetting, credentialing, and peer-review processes as it is of Arday personally. Instead, the institution that failed to catch these issues for a decade was allowed to recede into the background while a single Black professor absorbed the entirety of the public reckoning, his resignation treated as the natural end of the story rather than the beginning of a harder question about how he got there in the first place.

Two things can be true

None of this requires pretending Arday's scholarship was beyond reproach. He himself acknowledged that mistakes had occurred, even as he denied deliberately plagiarizing material. Investigating serious allegations against any academic, regardless of race, is legitimate, and being a symbol of diversity should never function as immunity from normal scholarly standards.

But proportionality matters, and proportionality is precisely what collapsed here. A citation problem in a ten-year-old thesis does not, on its own, generate hundreds of articles, a public autopsy of someone's disability and childhood, and a campaign led by a man who has built his reputation arguing that people who look like the accused do not belong in these positions to begin with. When the punishment for an alleged academic lapse becomes an unrelenting, identity-focused public campaign, the campaign itself becomes the story, and a serious institution has an obligation to ask whether race shaped not the fact of the scrutiny, but its ferocity.

---

This article makes the case for one interpretation of a genuinely contested and unresolved situation. Others, including Cofnas himself, argue that the plagiarism findings were substantive and long known to Cambridge insiders, that flagging misconduct is not disproportionate simply because the subject is well known, and that treating the accusation as racially motivated risks shielding real academic fraud from scrutiny. Ghent University's investigation into Cofnas has not concluded, and no independent inquiry has yet determined whether race materially affected how the case was pursued. Readers weighing this case should consider both sets of claims as the record develops.

My Next Four Reads


Yes, I know I have more than thirty books next to my nightstand waiting to be read. Here are four more that I am adding to the stack.

A Guardian and a Thief by Megha Majumdar, is the timeliest pick: a taut, 2025 National Book Award finalist set in a near-future Kolkata ravaged by climate change and food scarcity, following two families forced into conflict as they try to protect their children. It's also only 200 pages of tightly wound tension, so it won't demand a long commitment.

Bridge of Sighs, by Richard Russo, moves at the opposite pace: a patient, generous portrait of a small upstate New York town and the decades-long friendships that shape it.

Stone Yard Devotional by Charlotte Wood, rewards that same patience. It's a Booker-shortlisted Australian novel about a woman who abruptly leaves her life for a cloistered religious community, told through a journal-like structure as she reckons with unresolved grief and guilt. Quiet, but it stays with you.

Cursed Daughters is the most propulsive of the four, Oyinkan Braithwaite's follow-up to My Sister, the Serial Killer, about three generations of Nigerian women bound by a family curse and a daughter believed to be her dead aunt's reincarnation.

Together they cover very different registers: thriller-paced literary fiction, slow-burn small-town saga, contemplative memoir-in-fiction, and family-curse drama. Worth adding to the to the nightstand stack.

Saturday, August 22, 2026

Evaluating Historical Accuracy in LLM: Reconstruction, Jim Crow, and the Civil Rights Movement as High-Stakes Test Cases


by Alvin Blackshear  |  Historian & Researcher   <ablackshear@gmail.com>

Ask a large language model a seemingly straightforward question: "Who was the first African American elected to statewide office after Reconstruction?" The answer arrives instantly, fluently, and with the trappings of scholarship. It may cite books. It may offer dates with confident precision. It sounds, in short, like the work of someone who knows the field.

Historians know better. A question that appears simple on its surface is, in practice, layered with complications. What counts as "statewide office"? Does an elected superintendent of education count the same as a lieutenant governor? Does the answer depend on which state, since Reconstruction unfolded unevenly across the former Confederacy and border states? Does "first" mean first elected, first to take office, or first to serve a full term amid the violence and fraud that often interrupted Reconstruction-era elections? A model that produces a single, tidy name has already made a series of interpretive choices, usually without disclosing them, and sometimes without any grounding in verifiable sources at all.

This is the central argument of this article: historical accuracy cannot be measured merely by whether an isolated fact happens to be correct. Accuracy in historical writing is a compound of chronology, sourcing, context, and interpretation, and any evaluation of AI-generated history that stops at fact-checking a name or a date has already missed most of what makes history rigorous. Reconstruction, African American history, and the Civil Rights Movement make an unusually demanding proving ground for this kind of evaluation, and examining them closely offers historians a template for judging AI output more broadly.

Why These Three Fields Are an Ideal Stress Test for LLMs

Several features of these fields combine to expose the weaknesses of language models more sharply than most other historical subjects.

Scholarship on Reconstruction and the Civil Rights Movement has changed rapidly over the past two generations, as historians have moved from older, dismissive interpretations toward accounts that take Black political agency, local organizing, and grassroots activism seriously. A model trained on a broad mixture of older and newer sources may blend outdated framings with current scholarship without signaling the shift, producing a synthesis that never actually existed in the historiography.

These fields are also politically contested in ways that shape how information about them circulates online, in textbooks, and in public commentary. Terminology is inconsistent across time and place: the language used for racial categories, political offices, and organizations has shifted across the nineteenth and twentieth centuries, and even among historians writing in the same decade. Local, state, and federal histories overlap and sometimes contradict one another, particularly during Reconstruction, when a county-level fusion government might operate differently from the state legislature above it. Archival records are incomplete, and the voices of the enslaved, the formerly enslaved, and the disenfranchised are systematically underrepresented in the documentary record that survives.

Together, these characteristics force a model to reason about evidence rather than simply retrieve settled facts. That is precisely why they make such a useful stress test. A model that performs well on a straightforward question about, say, the dates of a well-documented battle may perform far worse when the underlying historical reality is genuinely contested, sparsely documented, or subject to ongoing revision.


What Does "Historical Accuracy" Actually Mean?


Before any evaluation can proceed, it helps to unbundle the concept of accuracy into its component parts.

Factual accuracy is the most familiar category and covers names, dates, and places. It is necessary but far from sufficient. Chronological accuracy concerns the sequence of events, their timing relative to one another, and the crucial distinction between cause and consequence, a distinction that language models frequently blur when they narrate history as a smooth chain of inevitabilities rather than a contested and contingent process.

Citation accuracy asks two separate questions: do the cited works actually exist, and, if they do, do they actually support the claim attributed to them? A hallucinated citation is an obvious failure, but a real citation attached to a claim it does not support is a subtler and in some ways more dangerous one, since it can survive a cursory check.

Contextual accuracy asks whether the historical setting has been correctly explained, not merely gestured at. Historiographical accuracy asks whether the answer acknowledges that historians disagree, and whether it fairly represents the range of interpretation rather than presenting one school of thought as settled consensus. Evidentiary accuracy, finally, asks whether every important claim can be traced back to documentary evidence, the standard that separates history from well-informed narrative.

A model can satisfy the first of these categories while failing all the others. That gap is where most of the risk in AI-generated history lives.


Common Failure Modes in African American History


Rather than cataloguing hallucinations as isolated errors, it is more useful to organize failures by the type of historical reasoning they violate.

Invented primary sources and fabricated quotations are the most conspicuous failures, and they are especially damaging in African American history, where authentic primary source material is already scarce and precious. A model that invents a plausible-sounding letter or speech does more than get a fact wrong; it pollutes a record that is already thin.

Composite biographies are a subtler failure, in which a model merges the experiences or achievements of two or more historical figures into a single, internally inconsistent account. Chronological compression, in which distinct decades are folded together as though they were a single moment, is common in narratives that span the long civil rights era, a period historians increasingly define as stretching from the 1930s through the 1970s rather than the traditional decade between 1955 and 1965.

Misidentified officeholders and geographic confusion, particularly the conflation of state and local officials, recur often enough in Reconstruction-era queries to be treated as a distinct category. False causal claims, in which one event is assumed to have directly produced another without the intervening political and social process that actually connected them, distort the texture of change over time. Presentism, the projection of modern assumptions and categories backward onto historical actors who did not share them, is perhaps the most pervasive failure of all, since it can infect an otherwise factually accurate answer with an interpretive frame the historical subjects themselves would not have recognized.


Reconstruction as an AI Evaluation Benchmark


Reconstruction is an unusually demanding benchmark because it compresses so much complexity into roughly a dozen years. A model addressing this period correctly must integrate multiple constitutional amendments, uneven state-by-state implementation, rapid and often reversed political change, contested definitions of citizenship and suffrage, shifting racial classifications, waves of violent resistance, and a series of short-lived governments that rose and fell within a single decade.

The documentary base compounds the difficulty. Freedmen's Bureau records, election returns, and the proceedings of state constitutional conventions are voluminous but unevenly digitized, unevenly indexed, and often contradictory on points of basic fact, such as the precise vote count in a contested election. Economic historians examining Reconstruction have shown just how tangled the causal picture becomes when one moves past the political narrative into the socio-economic effects of emancipation and federal enforcement, where multiple datasets, regional variation, and the halting reach of federal authority all interact in ways that resist simple summary. A model asked to explain "the economic effects of Reconstruction" without acknowledging this complexity is not simplifying for clarity; it is producing a distortion.


The Civil Rights Movement as a Verification Challenge


The Civil Rights Movement presents a different but related problem: the temptation toward a tidy, celebratory narrative organized around a small number of famous leaders. Language models trained on the volume of public-facing material about the movement tend to reduce decades of organizing to a handful of familiar names and set-piece events, at the expense of the local activism, particularly the labor organizing, legal work, and grassroots mobilization carried out largely by women and by people whose names never entered the national press.

This tendency shows up as compressed decades folded into a single narrative arc, confusion between overlapping organizations with similar acronyms and overlapping memberships, legal decisions misplaced in time or attributed to the wrong court, invented speeches, and misquoted newspapers. What makes this category of error especially important for historians to flag is that omission functions here as a form of distortion. A summary that mentions only the most famous figures and set-piece protests is not merely incomplete; it actively misrepresents how change actually happened, and it can be just as misleading as an outright fabrication, even though it contains no single false statement.


A Historian's Evaluation Rubric for LLM Output


Translating these concerns into a working tool for evaluation yields a rubric that can be applied consistently across outputs:

Criterion

Questions

Citation authenticity

Does every cited work exist?

Primary-source fidelity

Are quotations verifiable?

Chronology

Is the timeline internally consistent? 

Context

Is sufficient historical context provided?

Historiography

Are competing interpretations acknowledged?

Provenance

Can factual claims be traced?

Geographic precision 

Are jurisdictions correctly distinguished?

Uncertainty

Does the model acknowledge ambiguity?

Reproducibility

Can another historian repeat the verification?

Applied consistently, this rubric shifts the evaluator's attention away from spot-checking isolated facts and toward the structural qualities that make a historical account trustworthy. A response can score reasonably well on factual accuracy and still fail on historiography, provenance, or uncertainty, and it is precisely those failures that a fact-focused benchmark would miss.

Measuring Historical Accuracy Beyond Benchmarks

Most existing AI benchmarks evaluate factual recall, question answering, and multiple-choice accuracy, formats borrowed largely from standardized testing rather than from historical practice. Recent scholarly efforts have begun to push against this default. One 2026 benchmark built around the Chinese imperial examination tradition was designed specifically to test historical reasoning rather than simple retrieval, and its authors found that even the most capable models struggled once the task moved beyond fact lookup into the kind of analytical work professional historians actually perform. That finding is instructive: models can appear highly competent on narrow factual questions while still falling well short of the reasoning historians rely on day to day.


Historians evaluate evidentiary reasoning, provenance, interpretation, source criticism, contextualization, and the state of competing scholarship. These are fundamentally different evaluation targets than the ones most AI benchmarks were built to measure, and the gap between the two helps explain why a model can pass conventional accuracy tests while still producing what one recent essay memorably termed "stochastic history," an account that carries the surface texture of scholarship without the interpretive reasoning that actually produces it. Closing that gap will require benchmarks designed by historians, for historical reasoning specifically, rather than benchmarks adapted from unrelated domains.


Human-in-the-Loop Historical Verification


None of this argues for abandoning AI as a research aid. It argues for a disciplined workflow in which the historian remains the final authority. A workable sequence looks like this: the model generates a draft; the historian validates every citation; primary sources are independently verified; chronology is checked against the documentary record; competing interpretations are compared against current historiography; the historiographical framing itself is reviewed for balance; and only then does the material move toward publication.

In this workflow, the historian is not a passive consumer of AI output but an active evaluator standing between a plausible draft and a trustworthy account. That distinction, subtle as it sounds, is the difference between treating AI output as evidence requiring evaluation and treating it, mistakenly, as evidence in its own right.


Future Directions


Several developments now underway may narrow the gap between AI fluency and historical rigor. Retrieval-augmented generation systems, which ground model output in retrieved documents rather than internalized patterns, offer one promising path, as do archival retrieval systems built specifically for historical collections. Provenance-aware AI, designed to track and disclose the origin of every claim it makes, addresses the citation problem directly rather than leaving it to after-the-fact verification. Uncertainty estimation, which would allow a model to signal when the historical record is genuinely contested rather than presenting contested claims with false confidence, speaks directly to one of the recurring failures described above. Citation-grounded generation, historian-designed benchmarks, and dedicated historical evaluation datasets round out a research agenda that treats historical reasoning as its own discipline rather than a subset of general knowledge retrieval.

None of these tools eliminate the historian's role. If anything, they raise the bar for what that role requires: historians will need to evaluate not only the answers a retrieval-augmented system produces but the quality, provenance, and interpretive framing of the evidence it retrieves in the first place. A flawed archive fed into a well-engineered retrieval system still produces a flawed history.

Until AI systems consistently meet the standard historians already hold themselves to, historical accuracy will remain not a property inherent to the model, but the product of rigorous, sustained human evaluation.


Sources


"Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination." ACL 2026. https://aclanthology.org/2026.acl-long.1378

"Using Generative AI in Historical Practice." Cambridge University Press, 2026. https://www.cambridge.org/core/elements/using-generative-ai-in-historical-practice/7C1392A6E9DBD379FAA42E2D16A6D45B

"Writing Against the Machine: Computational Authorship and Historical Writing." History (Wiley), 2026. https://onlinelibrary.wiley.com/doi/10.1111/1468-229x.70092

"The Algorithm of Silence: Artificial Intelligence, Archival Bias and the Ethical Reconstruction of Digital Memory." Journal of Documentation, 2026. https://doi.org/10.1108/JD-08-2025-0249

"Automating the Past: Artificial Intelligence and the Next Frontiers of Digital History." International Journal of Humanities and Arts Computing, 2026. https://doi.org/10.3366/ijhac.2026.0361

"Reframing Historical Text Extraction: A Cross-Pathway Validation of OCR, LLM-Assisted Correction, and Direct Multimodal Transcription." Information (MDPI), 2026. https://www.mdpi.com/2078-2489/17/8/722

"Was Freedom Road a Dead End? Socio-economic Effects of Reconstruction in the American South." Economic History Review, 2026. https://onlinelibrary.wiley.com/doi/10.1111/ehr.70085

"Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models." 2026. https://arxiv.org/abs/2602.17433

Friday, August 21, 2026

What Exactly Is WISeR?


by Alvin Blackshear  |  Historian & Researcher  <ablackshear@gmail.com> 

A Deeper Look at Medicare's AI Prior Authorization Pilot

A lot of you wrote in after “Is AI Deciding Who Gets Healthcare?” asking the same question: okay, but where, and for what?  Fair ask. Here are the details.

The Basics

WISeR stands for Wasteful and Inappropriate Service Reduction. It's a six-year model run by Centers for Medicare & Medicaid Services' (CMS) Innovation Center, announced in June 2025 and live since January 1, 2026. Unlike Medicare Advantage, where prior authorization has been standard for years, WISeR marks one of the first real tests of AI-assisted prior authorization inside original (fee-for-service) Medicare, a program that has historically had almost none of it. CMS frames the goal as trimming fraud, waste, and abuse. Critics frame it as outsourcing coverage denials to vendors whose paychecks depend on how much they deny.

Which States

WISeR is currently running in six states, split across four Medicare Administrative Contractor jurisdictions:

        New Jersey

        Ohio

        Oklahoma and Texas

        Arizona and Washington

CMS chose these based partly on which states' MAC jurisdictions and existing local coverage rules made the model easiest to test cleanly. There's no public timeline yet for expanding beyond these six, but several health-finance analysts expect that if the pilot runs smoothly, CMS won't wait out the full six years before widening it.

Which Procedures

For the initial phase, WISeR covers 14 selected Medicare Part B services, all elective rather than emergency or first-line treatments. The list centers on categories CMS has flagged as historically vulnerable to overbilling or overuse, including:

        Skin and tissue substitutes (grafts used in wound care)

        Electrical nerve stimulator implants

        Epidural steroid injections for pain management

        Knee arthroscopy for osteoarthritis

CMS has said it may add services to this list as the model progresses. Notably excluded: inpatient-only services, emergency services, and anything CMS considers too risky to delay.

Who's Actually Doing the Reviewing

This is the detail that surprises people most: the “model participants” aren't providers or insurers; they're technology companies. CMS selected six of them to operate the AI/ML-driven review process across the six states: Cohere, Genzeon, Humata, Innovaccer, Virtix, and Zyter.

Here's roughly how it works. A medical provider can either submit a request for prior authorization before performing the service, or skip that step and submit the claim for a post-service, pre-payment review instead. If they go the prior-authorization route, the AI vendor is supposed to return a determination (an affirmation or non-affirmation) within about three days. Any non-affirmation (denial) has to go through a licensed clinician before it's finalized; CMS says the AI itself can't issue the final “no.”

The Part That Worries Critics

The vendors' compensation is tied, at least in part, to the savings their reviews generate, meaning the more claims they flag as non-payable, the more they can earn, adjusted against other metrics like accuracy and provider experience. Revenue-cycle experts have publicly guessed that a meaningful share of claims routed through WISeR (some estimates run around 25%) could end up denied. That incentive structure is central to why members of Congress introduced the Ban AI Denials in Medicare Act, aiming to shut the model down before it expands further.

CMS is also piloting a “gold-carding” feature, expected to roll out further details later in 2026, that would exempt providers with consistently clean approval histories from future prior-authorization requirements, an attempt to reward compliance rather than blanket every claim with review.

The Takeaway

WISeR is narrower than my original article may have implied: six states, 14 services, six named vendors, with a human clinician required to sign off on any denial. But the financial incentives baked into the model, and the fact that this is the first large-scale AI prior-auth experiment inside traditional Medicare, are exactly why it's worth watching closely as it expands.

Sources

CMS, “WISeR (Wasteful and Inappropriate Service Reduction) Model”: https://www.cms.gov/priorities/innovation/innovation-models/wiser

CMS, “WISeR Model Frequently Asked Questions”: https://www.cms.gov/priorities/innovation/files/document/wiser-model-frequently-asked-questions

CMS, “WISeR Fact Sheet”: https://www.cms.gov/files/document/wiser-fact-sheet.pdf

CMS, “WISeR Model Provider and Supplier Operational Guide”: https://www.cms.gov/priorities/innovation/files/wiser-provider-supplier-guide.pdf

Nick Hut, “The WISeR prior authorization model for Medicare is set to pose challenges for hospitals,” HFMA (Dec. 23, 2025): https://www.hfma.org/revenue-cycle/the-wiser-prior-authorization-model-for-medicare-is-set-to-pose-challenges-for-hospitals/

ASRA Update, “CMS Provides More Details on WISeR Prior Authorization Model” (Oct. 24, 2025): https://asra.com/news-publications/asra-update-item/asra-updates/2025/10/24/cms-provides-more-details-on-wiser-prior-authorization-model

Dinsmore & Shohl, “Is AI WISeR? CMS Models AI-based Prior Authorization Process in Six States” (Jan. 30, 2026): https://www.dinsmore.com/publications/is-ai-wiser-cms-models-ai-based-prior-authorization-process-in-six-states/

Moss Adams, “Medicare WISeR Model Requires Prior Authorizations in Six States” (Jan. 13, 2026): https://www.mossadams.com/articles/2026/01/medicare-wiser-model

DLA Piper, “CMS launches WISeR Model: What providers need to know” (2026): https://www.dlapiper.com/en/insights/publications/2026/01/cms-wiser-model

Center for Medicare Advocacy, “Early Reports on WISeR Model Are Troubling” (Mar. 26, 2026): https://medicareadvocacy.org/early-reports-on-wiser-model-are-troubling/

Congressional Research Service via Congress.gov, “Overview of the Medicare Wasteful and Inappropriate Service Reduction (WISeR) Model” (May 12, 2026): https://www.congress.gov/crs-product/IF13133

Medical Economics, “WISeR spending or unneeded delays in health care? Prior authorizations, AI in Medicare prompt concerns” (July 15, 2026): https://www.medicaleconomics.com/view/wiser-spending-or-unneeded-delays-in-health-care-prior-authorizations-ai-in-medicare-prompt-concerns