Showing posts with label ChatGPT. Show all posts
Showing posts with label ChatGPT. Show all posts

Wednesday, August 12, 2026

The Biggest Mistake Historians Make When Using ChatGPT

 

by Alvin Blackshear  |  Historian & Researcher

A historian asks ChatGPT a seemingly straightforward question: “Who was the first African American elected to public office in Delaware after Reconstruction?

Within seconds, the system produces a polished answer. It supplies a name, an election date, a description of the office, and perhaps even several citations. The prose is orderly, confident, and specific. Nothing about it appears careless. The historian copies the response into a research notebook and proceeds to the next question.

Only later does the answer begin to unravel. One citation does not exist. A quotation attributed to a nineteenth-century newspaper is actually a modern paraphrase. The date belongs to another officeholder. An appointment has been confused with an election. Details from two people have been combined into a single biography.

Nothing in the original response sounded implausible. That is precisely the danger.

The greatest mistake historians make when using ChatGPT is not merely trusting an occasional incorrect answer. It is treating AI-generated prose as though it were historical evidence. ChatGPT produces narratives. Historians require evidence. Those activities may sometimes overlap, but they are not the same.

The Abandonment of Historical Method

Historians already possess a sophisticated method for evaluating claims. When examining a letter, newspaper, memoir, photograph, government record, or oral history, they ask familiar questions. Who created the source? When was it created? For what purpose? Under what circumstances? What interests shaped its contents? Can its claims be corroborated? Has another scholar interpreted the evidence differently?

These questions constitute more than an academic ritual. They are the means by which historical knowledge is distinguished from rumor, memory, advocacy, folklore, and invention.

Yet something curious happens when many researchers begin using generative artificial intelligence. The normal habits of source criticism are suspended. A response is accepted because it sounds reasonable. A citation is trusted because it resembles an academic citation. A quotation is repeated because the language seems appropriate to the period.

The historian who would never accept an unsigned reminiscence without examining its provenance may accept an AI-generated paragraph without asking where any of its claims originated.

ChatGPT should therefore be treated as an intelligent research assistant, not as a documentary source. It can suggest search terms, identify possible lines of inquiry, compare concepts, organize notes, formulate questions, and help clarify difficult prose. It can also point a researcher toward archives, books, people, and events that merit investigation. Its usefulness, however, does not convert its output into evidence.

The distinction is fundamental. A research assistant may tell a historian that a document probably exists. The historian must still locate and examine it.

Why Fluency Is So Persuasive

Large language models are especially dangerous when they are almost correct. An obviously absurd answer invites skepticism. A plausible answer, containing real names, accurate background information, and one fabricated detail, may pass unnoticed.

ChatGPT’s authority is largely rhetorical. It writes in complete sentences, arranges events chronologically, supplies transitions, and often presents conclusions without visible hesitation. Readers naturally associate these qualities with knowledge. Confidence, specificity, and coherence become substitutes for provenance.

Recent research suggests that this problem is not simply the result of careless prompting. Kalai and his colleagues argue that common accuracy-based evaluations may reward models for guessing instead of acknowledging uncertainty. If a benchmark gives credit only for correct answers and imposes no meaningful cost for plausible errors, the model is encouraged to answer even when abstention would be more responsible. The researchers propose open scoring rubrics that disclose how errors and abstentions will be evaluated.

This observation has important consequences for historians. A model that always produces an answer may appear more useful than one that frequently says, “I do not have enough evidence.” In historical research, however, an admission of uncertainty may be the more accurate and intellectually responsible response.

The historian must therefore evaluate not only what the system says, but whether the system should have answered at all.

Five Forms of Historical Hallucination

Historical hallucinations frequently assume recognizable forms.

The first is the invented citation. ChatGPT may produce a realistic book title, journal article, archival collection, author, volume number, or page range. Every component may appear academically credible even though the source does not exist.

The second is the misquotation of a primary source. A model may modernize, compress, or reconstruct a statement and then present the resulting language inside quotation marks. The general sentiment may be accurate, but the words are not documentary evidence.

The third is the composite biography. Details belonging to several individuals may be joined into one apparently coherent life. This is especially dangerous when researching people who share names, occupations, institutions, military units, or geographic locations.

The fourth is incorrect chronology. A model may identify real events but arrange them in the wrong order, confuse the date of an appointment with the date of an election, or place a person at an institution before that person arrived there.

The fifth is false causation. ChatGPT may connect two events with phrases such as “therefore,” “as a result,” or “this led to,” even when the evidence establishes only sequence or correlation.

These errors are not equally visible. An invented person may be discovered quickly. A subtle chronological error or unsupported causal inference may survive several rounds of editing because it fits an expected narrative.

Why Better Prompts Are Not Enough

Prompt design can improve an AI response. A historian can request citations, ask the system to distinguish facts from interpretations, require expressions of uncertainty, or instruct it not to invent missing information. Such practices are worthwhile.

They are not verification.

A carefully written prompt may reduce the probability of an error, but it cannot transform generated language into historical evidence. Even a response that includes citations must be checked against the cited material. Even a quotation accompanied by a page number must be located on that page. Even a claim presented as certain must be corroborated.

Research on hallucinations in academic writing identifies fabricated citations, factual distortion, named-entity errors, logical inconsistency, and propagation errors as threats to scholarly integrity. The final responsibility remains with the human author. An historian cannot excuse a false statement by explaining that ChatGPT supplied it.

Better prompting is therefore a research technique, not an evidentiary standard.

Testing AI as Historians Test Evidence

Historians need a repeatable method for testing AI-generated research. The evaluation should move beyond impressions such as “the answer seemed good” or “the model performed well.”

A practical scoring rubric might assess eight dimensions:

1.      Citation existence: Does every cited work or archival collection exist?

2.      Citation accuracy: Does the cited source support the specific claim?

3.      Quotation fidelity: Are quoted words reproduced exactly and in context?

4.      Identity accuracy: Have people with similar names or careers been distinguished?

5.      Chronological accuracy: Are dates and sequences correct?

6.      Geographic and institutional accuracy: Are places, offices, organizations, and jurisdictions correctly identified?

7.      Provenance: Can each important assertion be traced to accessible evidence?

8.      Historiographical consistency: Does the answer acknowledge significant scholarly disagreement?

Each category could be scored from zero to four. A zero would indicate fabrication or complete failure. A score of one would reflect major errors. Two would indicate partial support or substantial ambiguity. Three would represent generally accurate work with minor deficiencies. Four would require full and independently verified support.

A weighted rubric would assign greater penalties to invented citations, false quotations, and merged identities than to minor stylistic or contextual omissions. This matters because not all errors have the same scholarly consequence.

Benchmark testing should also use a fixed set of questions rather than memorable anecdotes. Researchers could create a collection of historical questions at several difficulty levels: well-documented national events, obscure local events, disputed interpretations, biographical identity problems, chronological puzzles, and questions whose correct answer is genuinely unknown.

The same questions could then be submitted to several AI systems under controlled conditions. Evaluators would record accuracy, citation validity, unsupported assertions, appropriate abstentions, and changes across repeated runs. Such testing would reveal not only whether a model produces correct answers, but how it fails.

What RAG Can and Cannot Do

Retrieval-augmented generation, commonly known as RAG, offers one method of improving reliability. A RAG system retrieves documents from a designated collection before generating an answer. For historians, that collection might contain archival finding aids, digitized newspapers, oral-history transcripts, government records, or scholarly articles.

This approach can make AI responses more transparent and more closely connected to identifiable sources. It may reduce fabricated citations and permit researchers to inspect the evidence used in producing an answer.

Yet RAG does not eliminate the need for historical judgment.

Recent research distinguishes factuality from faithfulness. A statement may be faithful to a retrieved document while the document itself is inaccurate, biased, incomplete, or misinterpreted. Conversely, a statement may be factually correct but unsupported by the particular documents retrieved. A serious evaluation must test both dimensions.

For historians, retrieval is only the beginning. The fact that a system retrieved a newspaper article does not establish that the article was truthful. The fact that a claim appears in a memoir does not establish that memory was reliable. The fact that several sources repeat a statement does not prove they were independent.

RAG can retrieve evidence. It cannot perform the full intellectual work of source criticism.

The Historian’s Continuing Responsibility

ChatGPT does not threaten historical scholarship simply because it occasionally invents facts. Historical scholarship has always confronted error, forgery, propaganda, partial memory, and misleading testimony.

The greater danger arises when historians allow the fluency of artificial intelligence to replace the discipline’s oldest habit: skepticism.

For centuries, historians have questioned manuscripts, memoirs, newspapers, photographs, statistics, and archives. Artificial intelligence deserves no exemption from that tradition. Its output should be tested, scored, corroborated, and traced to evidence.

The historian’s task remains unchanged. It is not merely to discover information or compose a persuasive narrative. It is to determine what the surviving evidence permits us to say, what remains uncertain, and why the distinction matters.

References

  1. Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. “Evaluating Large Language Models for Accuracy Incentivizes Hallucinations.” Nature 653 (2026): 1047–1051. https://www.nature.com/articles/s41586-026-10549-w
  2. Rahman, Subhey Sadi, et al. “Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models.” Artificial Intelligence Review 59 (2026), article  70. https://link.springer.com/article/10.1007/s10462-025-11454-w
  3. Fadeeva, Ekaterina, et al. “Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation.” Findings of the Association for Computational Linguistics: ACL 2026 (2026): 6814–6836. https://aclanthology.org/2026.findings-acl.338
  4. Giray, Louie, Md Emdadul Islam, and John Arvin Glo. “AI Hallucinations in Academic Writing: Implications for Research Integrity.” Naunyn-Schmiedeberg’s Archives of Pharmacology (2026). https://pubmed.ncbi.nlm.nih.gov/42217043 

Sunday, August 02, 2026

AI for Older Adults


AI for Older Adults

by Alvin Blackshear | Historian & Researcher

Artificial intelligence tools like ChatGPT, Claude, Gemini, Perplexity, and Copilot have moved from novelty to necessity faster than almost any technology in memory. For adults over 60, that speed can feel less like an invitation and more like being left behind at the station. But research on how older adults actually use, and struggle with, these tools tells a more encouraging story: the gap isn't about ability, it's about approach. A 2026 study of older adults' technology support requests found that when queries were rewritten to add missing context and clearer language, the accuracy of AI-generated solutions jumped from 46% to 69%, and related Google search results improved from 35% to 69% accuracy. In other words, the tools already work well enough. The skill worth learning is how to talk to them.

This article walks through four practical steps: choosing the right AI for your needs, writing prompts that get better answers, understanding what AI "agents" are and when to trust them, and avoiding the mistakes that trip up new users of every age, but especially those who didn't grow up parsing tech jargon.

Choosing the Right AI

There is no single "best" AI assistant. Each is tuned for different strengths, and picking the right one for the task at hand saves time and frustration. A simple way to think about it:

·   Best for everyday questions: ChatGPT is the most general-purpose assistant, well suited to answering quick questions, drafting emails, or explaining a topic in conversation.

·   Best for writing and long documents: Claude tends to produce clear, well-organized long-form writing and is a strong choice for letters, essays, or reviewing lengthy documents.

·   Best for web research: Perplexity is built around search, pairing its answers with cited sources so you can click through and verify what it tells you, which is useful when accuracy matters most.

·   Best for Google users: Gemini integrates directly with Gmail, Google Docs, and Google Search, which makes it a natural fit if that's already your everyday ecosystem.

·   Best for Microsoft 365 users: Copilot is built into Word, Outlook, and Excel, so if your family or workplace runs on Microsoft software, it can work inside the programs you already know.

None of these choices is permanent or exclusive. Many people end up using two: one for quick daily questions and another for anything they want double-checked with sources. The point isn't to find the "correct" one, but to match the tool to the task, the same way you'd choose a phone call over a text message depending on what you're trying to accomplish.

Learning to Write Better Prompts

The single biggest lever an older adult has for getting better answers from AI isn't a special app or setting. It's the wording of the question itself. Research on communication between older adults and AI systems identified four recurring patterns that reduce answer quality: verbosity (explaining too much surrounding detail), incompleteness (leaving out a key fact the AI needs), over-specification (adding assumptions that aren't actually true), and under-specification (being too vague about what's wanted). The good news from that same research is that these patterns are fixable, not through retraining how you think, but through a few reliable habits when you type or speak your question.

Ask for plain English. AI models default to whatever register they think fits the question, and that can drift into jargon. Simply adding "explain this in plain English, no technical terms" or "explain it like you would to a smart friend who isn't a computer person" changes the entire tone of the reply.

Request step-by-step explanations. Instead of asking "how do I attach a photo to an email," try "walk me through this one step at a time, and wait for me to say 'next' before moving to the next step." This breaks a wall of instructions into a manageable back-and-forth, and it also gives you a natural point to ask a follow-up question without losing the thread.

Ask for examples. Abstract instructions are harder to follow than concrete ones. Adding "give me a real example" or "show me what that would actually look like on my screen" turns a vague description into something you can match against what you're actually seeing.

Ask the AI to critique its own answer. This is a lesser-known but powerful technique: after receiving an answer, ask "is there anything wrong with this answer, or a simpler way to do it?" AI models are often able to catch their own errors or oversimplifications when prompted to double-check themselves, and this habit builds in a layer of quality control you don't have to supply yourself.

The underlying finding from the research is worth repeating: rephrased, well-specified questions were understood far better by others (93.7% versus 65.8% comprehension in one comparison) and led to solutions people felt confident following. A better-written question isn't a nicety. It's the difference between a correct answer and a confusing one.

Using AI Agents

You may have started to hear the term "AI agent," distinct from a regular chatbot. It's worth understanding the difference, because agents come with real convenience and real risk.

What an agent is. A standard AI chat answers a question and stops. An agent is designed to take a goal and carry out a sequence of actions on its own (searching the web, filling out a form, booking something, or working through a multi-step task) before reporting back. Industry analysis on the state of AI agents in 2026 notes that the defining shift in newer systems is their ability to handle progressively longer and more complex tasks, maintaining context across many steps rather than answering a single quick question.

When to use one. Agents are useful for tasks that are repetitive, well-defined, and low-stakes if something goes slightly wrong, for example, summarizing a folder of documents, comparing prices across several websites, or drafting a set of similar emails. They save the most time exactly where a task has many small, mechanical steps that would otherwise take you an hour of clicking around.

Risks of autonomous agents. The tradeoff for that convenience is reduced oversight. Because an agent acts across many steps without checking in at each one, a small misunderstanding early on can compound into a larger error by the end, the AI equivalent of a wrong turn early in a road trip that takes you far off course before anyone notices. This matters most for anything involving money, personal information, medical decisions, or sending communications on your behalf. Never give an agent standing access to your bank accounts, passwords, or ability to make purchases without a manual confirmation step at the end.

Practical everyday examples. Reasonable, lower-risk uses include: asking an agent to research and summarize options for a specific purchase (with you making the final decision), having it organize digital photos or files, or using it to draft, but not send, a batch of similar messages, such as holiday cards or appointment reminders, for you to review individually before anything goes out.

Avoiding Common Mistakes

Even with the right tool and better prompts, a handful of habits cause most of the frustration people report with AI, regardless of age.

Trusting AI too quickly. AI models can sound confident even when they're wrong. They occasionally generate plausible-sounding but incorrect information, a phenomenon often called "hallucination." Treat a confident tone as no guarantee of accuracy, especially for anything involving health, legal, or financial specifics.

Asking vague questions. As the communication research above showed, vague or incomplete questions are the most common source of poor answers. If a response feels off-target, the fastest fix usually isn't switching tools. It's adding the missing context and asking again.

Accepting the first answer. The first response an AI gives is a starting point, not a verdict. Asking a follow-up like "are you sure?" or "what would make this answer wrong?" often surfaces caveats or corrections the first answer left out.

Failing to verify important facts. For anything that matters, such as a medication interaction, a legal deadline, or a financial figure, cross-check the AI's answer against a trusted human source, such as your doctor, pharmacist, financial advisor, or a reputable website. Tools like Perplexity, which show their sources directly, make this step easier by letting you click through and confirm the claim yourself rather than taking it on faith.

The Bigger Picture

None of this requires becoming a technical expert. The research is fairly consistent on this point: older adults who use AI report high confidence and ease once they receive well-matched, clearly explained answers. The 2026 diary study found 89.8% of older adult participants felt able to answer AI's follow-up questions, and 94.7% felt able to follow the solutions they were given. The barrier isn't capability; it's a mismatch in how questions get asked and answered. Broader research on aging and generative AI echoes the same conclusion, pointing to real promise in health and cognitive support as long as tools are built and used thoughtfully, with attention to privacy and genuine ease of use. Choose a tool that matches the task, ask questions the way outlined above, use agents cautiously and only for low-stakes chores, and verify anything that truly matters. That combination, not fluency in tech jargon, is what determines whether AI becomes a genuinely useful assistant in daily life.

___________________________ 

Sources:

Dhanuka, N. (2026, February 26). State of AI agents 2026: Autonomy is here. Prosus. https://www.prosus.com/news-insights/2026/state-of-ai-agents-2026-autonomy-is-here

Ernst & Young Global Limited. (2026, March). Understanding older generations' adoption of AI [Ripples Research Report]. https://www.ey.com/content/dam/ey-unified-site/ey-com/en-gl/about-us/corporate-responsibility/documents/ey-gl-how-older-generations-are-engaging-with-ai-03-2026.pdf

Liu, T. (2026, May 29). The role of multimodal generative AI in older adults' health management: Systematic scoping review. JMIR AI. https://ai.jmir.org/2026/1/e84695

Shomee, H. H. (2026, January 15). Empowering older adults in digital technology use with foundation models. arXiv. https://doi.org/10.48550/arXiv.2601.10018