Are LLMs Stuck in Time?
Documenting historical events accurately is no small feat. As time passes, human memory fades and books, letters, posters, and other artifacts become difficult to find. By some estimates, we know less than 1% of what happened 500 years ago.
So, as researchers and technologists seek to build better large language models (LLMs), they face a challenge. AI can reasonably reconstruct the past through data and language, but it lacks the context to handle nuance: how people spoke, what clothes they wore, and how they approached daily life.
These LLM blind spots and possibly biases show up in predictable ways: frontier models can reduce cultures to caricatures and stereotypes, misrepresent events, and sneak in objects, words, and ideas that didn’t exist in that era. “A history of the past isn’t the past,” said Matthew Wilkens, associate professor of Information Science at Cornell University.
As a result, researchers are exploring ways to address the gaps. This includes tools to measure historical accuracy, as well as methods—such as building time-locked models trained exclusively on content from a specific era—that paint a more realistic picture. The goal is to help LLMs reason from inside a period, rather than through a modern looking glass.
“Right now, when you prompt a model to behave like a Victorian, it often performs like a costume drama,” said Hamed Yaghoobian, assistant professor of Computer Science at Muhlenberg College. “It cannot unlearn the 20th century.”
It isn’t surprising that LLMs have a blurry view of the past. Missing words, symbols, expressions, inflections, idioms, and data cause systems to flatten and distort events. LLMs can also have blind spots on parts of the physical world: fashion, music, architecture, weapons, technology, and lifestyles.
The result is a mechanical yet confident regurgitation of past events. Wilkens likens it to logic bugs in software: “It looks like everything is working fine, but then you suddenly see things fall apart.” The problem is rooted in a basic issue: the LLM doesn’t reason within the historical period. It reconstructs events and applies learned modern-day assumptions, values, and knowledge.
Frontier models are especially prone to this problem, said David Bamman, associate professor at the University of California, Berkeley School of Information. While LLMs generally get key historical facts right, “We can’t be sure whether we’re measuring a model’s true reasoning ability or its ability to retrieve an answer someone else has already given,” he said.
“There’s a homogenization process that takes place during training, so these models are not able to speak authentically from a historical perspective,” added Ted Underwood, professor of Information Sciences and English at the University of Illinois, Urbana-Champaign. What results, in many cases, is what he described as a “middle-of-the-road, 21st century, Western point of view.”
For instance, an LLM might mimic a character from a 19th-century Irish novel—and the dialogue can sound entirely convincing. What’s really taking place could be deceptive, however. For instance, the LLM might spin up a caricature of an Irish lord or peasant that speaks in a stereotypical brogue. Underwood described this as “leprechaun talk.”
There’s no simple way to fix this problem. Feeding LLMs more media—letters, posters, record liner notes, old newspaper stories, and handwritten ledgers—can help fill gaps. Yet digitizing physical documents is incredibly time-consuming and labor-intensive, and that alone can’t address censorship, distortions, and biases that affect human historians as well. “Data matters more than architecture,” said Kaspar Beelen, a technical lead at the University of London’s School of Advanced Study.
Data contamination is another problem. Even time-locked models—trained only on content specifically between set dates—can’t completely wall off the present. For instance, when researchers built a closed model, Talkie, with a cutoff date of 1930, it still somehow learned about World War II and the New Deal. No one is sure whether this came from library metadata, footnotes that snuck into content, or contact with modern tokens.
Regardless, convincing yet deceptive results often follow. Making matters worse, as an LLM bends to modern logic, current perspectives, or stereotypes of the past, “You wind up with content that doesn’t represent what you think it represents,” Underwood said.
While there’s no simple fix for historical myopia, researchers believe they can rein in unruly LLMs. Developing a way to measure accuracy is a starting point. For example, Underwood is part of a group developing a tool that checks output against 714 questions drawn from period sources such as books, newspapers, diaries, religious texts, and archival records.
These questions span knowledge and inference across what the authors call “parallax” categories—where the right answer depends on the vantage point. This makes it possible for the LLM to alter its response based on differences such as, say, a British colonial administrator in India, a National Congress member, a common citizen, and a British socialist from the early 1900s.
Wilkens is exploring another approach: running AI-based simulations of past events. His idea is to replay moments and watch a range of scenarios unfold. He then can compare the model’s performance against the actual historical record. Other researchers are exploring ways for models to “unlearn” with an approach that strips out anachronistic facts so that the model doesn’t have access to knowledge that didn’t exist in that period.
Meanwhile, Bamman is focused on “measuring cultural phenomena in a multimodal setting, including movies and TV shows.” His research group is digitizing films dating as far back as 1922 to trace the rise of various concepts—including the development of the close-up and camera movement techniques. The goal is to “understand how methods from computer vision, natural language processing, and AI can shed light on the history of film,” he said. This method could help gain greater insight into historical cultural nuances, at least as far back as the last century.
Researchers also are developing closed models. One of the most ambitious projects is talkie, a 13-billion parameter open-weight model trained on 260 billion tokens of pre-1931 English—books, periodicals, court records, and reference works pulled from public-domain archives. The development team hopes to scale it to a GPT-3 level by the end of summer 2026.
A similar project, TimeCapsuleLLM, encompasses roughly 90 gigabytes of books, periodicals, and legal documents published in London between 1800 and 1875. The model attempts to understand links between words in that era rather than the links that exist today. It was developed by Hayk Grigorian, a former graduate student at Muhlenberg College who developed the project with Yaghoobian. “If I train from scratch, the language model won’t pretend to be old. It just will be,” Grigorian noted on his GitHub page.
The end goal, Wilkens said, isn’t to simulate a conversation with George Washington or Genghis Khan; it’s to create historical grounding that contributes to a better understanding of the past—along with an ability to apply the lessons to the present. “There are basic questions in the humanities that remain unsolved. LLMs offer the prospect of accelerating research,” he said.
Researchers are optimistic that a combination of fine-tuning, retrieval augmented generation, machine unlearning, multimodal approaches, closed models, and other methods will help them gain better visibility into the past. “A lot has been lost and there are limits to what we can reconstruct,” Underwood said. “But we are making progress.”
Samuel Greengard is an author and journalist based in West Linn, OR, USA.
Related Stories
AI News
Hebbia helped popularize AI for Wall Street. Now it's trying to stay ahead of the companies chasing the massive opportunity. The startup, one of the earliest companies to use AI to analyze large collections of financial documents, is overhauling its...
49 minutes ago
AI News
Nearly 700 AI agents coordinated Hugging Face attack, says report
49 minutes ago
AI News
Marvell Technology Sees AI Connectivity as Next Big Infrastructure Bottleneck
49 minutes ago
AI News
Jobs Artificial Intelligence Will Not Be Able to Replace for Now
49 minutes ago
AI News
Z.ai shares surge 8% after releasing new AI model running only on Chinese chips
50 minutes ago
AI News
The Regulatory Vacuum in AI
50 minutes ago
AI News
World's first patient to undergo live AI
2 hours ago
AI News
Nvidia discussed buying AI startup Hugging Face: report
2 hours ago