Cracking The Code: Can AI Help Decode Ancient Languages?

A self-taught AI engineer's claim about Linear A shows both the promise and the pitfalls of using AI for ancient language decipherment

Every ancient language deciphered so far needed a reference point, usually a bilingual text like the Rosetta Stone, or a known related language for comparison. Linear A, the writing system of the Bronze Age Minoan civilization on Crete, has neither. It’s described as an “isolated language” because it has no confirmed connection to any known language, living or dead, making it a major challenge for linguists over the past century.

Etruscan, a language once spoken in Italy before the rise of the Roman Empire, has fared only slightly better. Researchers have compiled a partial vocabulary from short funerary inscriptions, but its grammar and deeper meaning still elude linguists. Both languages represent exactly the kind of puzzle AI has recently gotten involved in, not by replacing linguists but by sharpening their instincts. So does AI actually stand a chance of finishing what a century of human effort couldn’t?

A simple hypothesis

A recent case highlights both the appeal and the problem. In June 2026, a self-taught AI engineer and amateur linguist claimed to have made a major breakthrough on Linear A. It started with a simple hypothesis that an unknown word in a prayer inscription came from a Semitic root meaning “to dwell” or “to live.” He then used AI-generated scripts to compare this phonetic pattern against the rest of the collected Linear A corpus. He reportedly managed to assign values to 40 symbols and compile a 408-word dictionary, arguing this shows Linear A is an extinct member of the Semitic language family, which includes Hebrew and Aramaic.

The need for creativity

It’s worth being precise about what AI did and didn’t do here. It didn’t come up with the idea, the engineer did. AI served as the research assistant, quickly testing a human hypothesis against thousands of characters, a job that would have taken a person months to do by hand. That division of labor is the pattern worth watching here, more than any single claim still under scrutiny.

“Cross-lingual transfer”

Writing in The Conversation, researcher Jane Adkins asks where AI actually helps and where it hits limits. She notes that it excels at large-scale pattern analysis, checking a hypothesis about a symbol or word against an entire corpus in minutes rather than years, and can catch repeating sequences a human eye would miss while reconstructing damaged or fragmentary inscriptions by predicting likely missing characters. AI also performs a distinct process called “cross-lingual transfer,” where a model trained on a known language can sometimes infer patterns in an unfamiliar but closely related one, similar to how knowing Spanish helps someone guess at Portuguese.

Researchers have already used this approach successfully on scripts like Ugaritic, an extinct Semitic language spoken in the late Bronze Age (roughly 1300-1190 BCE) in the ancient coastal city of Ugarit, in what is now Syria, since its language family is known. But there’s a clear limit: statistical pattern-matching can’t generate meaning from nothing. It needs a reference point, a known language family or a bilingual text, to tell which patterns matter and which are coincidence. Linear A and Etruscan are difficult precisely because that reference point is missing or incomplete, and no amount of computing power can create one from scratch.

Limited evidence

Given a large enough dataset, an AI model could in theory become fluent enough to hold a conversation with a Minoan or Etruscan speaker, on its own statistical terms. That means it could recognize recurring word patterns and other structural units and reuse those terms in a different context, something that could approximate a limited conversation, at least on the surface.

Only a time machine could confirm it

Still, an AI system couldn’t yet hand a person an actual translation, because fluency and meaning aren’t the same thing here. A model can learn which symbols follow each other and which words cluster together without ever knowing what any of them actually refers to. That’s exactly why verifying any AI-based claim about these languages is so hard, and it’s a bigger problem than any single case. Normally you’d check a proposed translation against native speakers, other texts, or decades of expert consensus. None of that exists for a language that hasn’t really been deciphered. Short of a time machine, there’s no way to confirm it. The small surviving text corpora make the problem worse: the entire surviving body of Linear A texts adds up to roughly 7,500 characters, small enough to fit on a single screen, and with so little data, almost any hypothesis can find scattered matches to support it.

A real accelerator

That’s why claims in this field lean so heavily on independent expert review and peer evaluation rather than statistical confidence scores, and why “AI spotted a pattern” and “AI found the right meaning” are very different claims that are easy to mix up. AI is a genuine accelerator, capable of compressing years of manual cross-checking into minutes and letting far more people work on these problems than institutional resources ever allowed. But it doesn’t eliminate the two things decipherment has always required: a genuine comparative reference point, and rigorous human judgment to tell a real breakthrough from an attractive coincidence. Until one of those shows up for Linear A or Etruscan, AI’s role in language decipherment will stay what it is today: a very fast assistant to a very old, very human puzzle.

Follow tovima.com on Google News to keep up with the latest stories
Exit mobile version