From Romani To Tsakonian, AI Learns Greece’s “Small” Languages

Of the roughly 7,000 languages in the world, very few count in the training of language models. In Greece, researchers are recording Romani, Tsakonian, and dialects before they fall silent and bringing them to Large Language Models through the Lingua program

What language do Large Language Models speak? Before we say “all of them,” let’s think about how often we interact with our favorite model in English, and the reasons why we do.

Before we say “all of them,” let’s think about languages that don’t have the institutional privileges of an official language: languages that are only spoken and not written, languages spoken in a very small or remote area by a small number of speakers, languages whose use is systematically stigmatized or even criminalized.

Before we say “all of them,” let’s think about how it would feel if, before we started a conversation, a language model asked us “ti ‘a ftiásoumi gia séna” (“what shall we do for you”), as a speaker of the Grevena dialect would say.

The question is not theoretical. It was raised in the most practical way a few days ago at the Museum of Cycladic Art, where Microsoft presented the results of the first Greek round of Lingua, the program that supports the digital recording of languages and dialects with a small digital footprint. Five projects, from Greek Romani to Tsakonian, were selected and are receiving funding or technological support.

Linguistic Inequality and Artificial Intelligence

The development of Artificial Intelligence was bound to reach the critical issue of the digital representation of languages. Large Language Models unlock human knowledge, but the way they are developed, drawing information from digitally available texts among other sources, contributes to widening linguistic inequality, since spoken languages are not equally represented in the digital world.

“The dominance of English in Large Language Models is undeniable, but there is a very large performance gap between less widely spoken languages and dominant ones,” said Wassim Hamidouche, a researcher at Microsoft’s AI for Good Lab, while presenting the program.

Wassim Hamidouche (AI for Good Lab) and Gianna Andronopoulou (Microsoft, right) at the presentation of the Greek part of Lingua.

According to UNESCO data, of the roughly 7,000 spoken languages, about 1,000 have a digital presence. “Of these, fewer than 100 have the kind of presence in the digital environment that meaningfully contributes to training large language models,” he noted.

Bridging this gap necessarily involves the companies that develop the models themselves. Lingua is one of the main pillars of Microsoft’s AI for Good Lab and has so far provided funding and technological resources to organizations, institutions, and researchers around the world who work on recording, documenting, and preserving less widely spoken languages and varieties that are at risk because of their small number of speakers.

Hamidouche was part of the research team that last January developed the “Bring Your Own Language into the LLMs” toolkit, which provides a complete digital toolkit for recording, mapping, and classifying dialects.

The Greek Lingua

The Ministry of Culture and the Ministry of Digital Governance are working together on the Greek part of Lingua. Funding of 50,000 euros went to ARSIS’s program for recording and preserving Greek Romani, the spoken language used by Roma communities in Greece. Four other projects received support through Azure services: the University of Crete’s Computational Dialectal Atlas (CoDAG), the Athena Research Center’s LALIES, the Tsakonia Archive’s T.S.A.K.O.N.I.A., and Istorima, a digital archive of oral histories that records the stories of dialect speakers (Tsakonian, Cappadocian, Arvanitika, and Pontic) in Greece.

Gianna Andronopoulou, General Manager of Microsoft Greece, Cyprus, and Malta, noted that the program’s first round “marks the starting point for building an ecosystem that will bring researchers and the state together and provide the motivation to continue these initiatives through collaboration.”

A Dialect With an Army, a Navy, and Data

The question of which languages will fit into the models is not new. It is the latest version of a very old question. The saying is well known even outside academic linguistics: a language is a dialect with an army and a navy. Unlike other pseudoscientific and pseudo-linguistic sayings, its accuracy is not disputed, even if its origin remains unclear.

It was made famous by Max Weinreich, who attributed it to an anonymous member of the audience at one of his lectures, to stress that the criteria that make a language variety standard, and later dominant, are social and historical. Today he might also add Large Language Models.

In Greece, Standard Greek, the one we learn at school, write, and use on every official occasion, the one documented in school grammars and dictionaries, has flattened dialectal diversity. It has exiled the non-standard varieties spoken in the country to the realm of folklore, destined to survive (not for long) within small communities whose populations keep shrinking.

We often think of natural languages as living organisms: they are born, they develop, and they die, following a pattern of natural decline. This metaphor is misleading. The death of a language or dialect is not a natural process but a political and historical one. As harsh as it may sound, it is more accurate to say that languages are killed. In the age of Artificial Intelligence, this mechanism takes on a new dimension: whatever is not in the data does not exist for the models either.

Digital Footprint and Living Speakers

“If a language has no digital footprint, it is in danger of disappearing,” Ms. Andronopoulou stressed. “Lingua helps stop the rapid path toward extinction faced by languages without a large digital footprint, given the fast pace of Artificial Intelligence development.” As she said, the conversation about Artificial Intelligence has moved from what it can do to how we choose to use it.

Those working in the field completed the picture from their side. “Digital tools are necessary, but languages survive as long as there are speakers, and that must be the priority,” stressed Stella Markantonatou, emeritus researcher at the Athena Research Center.

Stella Markantonatou of the ILSP (LALIES project) stressed the importance of field linguistics for recording and spreading dialectal varieties.

In the same spirit, Eleni Doundoulaki, Secretary General of Contemporary Culture, noted that “a language is saved when it is spoken,” stressing the importance of research work both in the digital environment and within communities. She said the people and communities who help collect the material should have a say in how it is preserved and used, and she also referred to copyright protection as part of the Ministry of Culture’s related initiatives.

Ms. Andronopoulou, for her part, spoke of “the democratization of technology, since not only are access and representation given to communities and groups that are not dominant, but these tools also return to the communities.” In fact, digital recording can also work the other way around: it can give communities themselves the tools to study and use their native language or the language of their heritage.

The Greek Projects

Ms. Markantonatou described the work of the Institute for Language and Speech Processing in recording nine Modern Greek dialects, from Pontic and Cypriot to the Apeiranthos dialect of Naxos, as well as fully documenting Pomak, stressing the importance of field linguistics for collecting material.

Stergios Chatzikyriakidis, until recently Professor of Computational Linguistics at the University of Crete, discussed the digital dialect corpora of the Computational Dialectal Atlas and the models developed for analyzing them further.

Linguist Stergios Chatzikyriakidis presents the Computational Dialectal Atlas (CoDAG) developed by the University of Crete.

The team has already digitized about 40,000 pages of dialect material, while the corpora total around 4.3 million words, drawn from sources freely available online. As he said, “the infrastructure provided by Lingua makes it possible to develop additional applications or even platforms.”

One of these is Svarna, “an open-source online corpus that brings together material from different databases.” It contains more than 500 million words, covering various levels of style in Modern Greek as well as many stages of its historical development and dialectal variation.

Georgia Kalpazidou, a PhD candidate in linguistics at the Democritus University of Thrace, represented ARSIS and spoke about the digital recording of oral histories from native speakers of the varieties of Greek Romani, as they are used in various Roma communities in Greece. As she notes, “there is significant variation from region to region.”

Georgia Kalpazidou, a PhD candidate at the Democritus University of Thrace, stressed the importance of recording folk stories in Romani.

Part of the material consists of folk tales, which begin with a specific phrase: “There was an old man and an old woman.” Romani is her native language, but she comments that “folk tales and folk stories are passed down less and less from generation to generation. Many Roma don’t know them, and their limited spread also affects the language itself, since it is an oral dialect that has not been documented.” For her, an additional goal of the program is to remove the stigma attached to using Romani and to break down stereotypes and negative attitudes toward its speakers. Here too, the challenge is not only technological: a stigmatized language is not saved just because it is archived.

Angeliki Politi (center) and Sofia Kamvysi (left) of the Tsakonia Archive presented the T.S.A.K.O.N.I.A. project on using the Tsakonian dialect in education.

Sofia Kamvysi and Angeliki Politi highlighted the educational side of recording one of the best-known and most studied dialects, Tsakonian, through the T.S.A.K.O.N.I.A. project (Tsakonian Speech and Knowledge Open Network for Inclusive AI). Tsakonian has the status of an endangered language, and as Ms. Politi noted, based on the number and age of its native speakers, it has an estimated life span of ten years. Ms. Politi, a scientific associate in charge of educational design for digital projects on Tsakonian, added that T.S.A.K.O.N.I.A. “will contribute significantly to supporting the Tsakonian dialect in the digital space, by developing tools for recording, studying, spreading, and learning it.”

The Digital Readiness of Greek

Relatively speaking, Greek is also under significant pressure from other, more powerful languages with many more native speakers. Vassilis Katsouros, Director of the ILSP, pointed out that the goal is for “Greek to be competitive and not fall behind other, more powerful languages,” with tools that are not just modern or impressive but relevant to the needs of each language community.

For Giannis Mastrogeorgiou, Special Secretary for Long-Term Planning, the rapid development of Artificial Intelligence makes it urgent to highlight the intangible cultural heritage carried by the Greek language and to explore our linguistic wealth more deeply. “As Artificial Intelligence moves into a phase of greater autonomy with agents, posing real, but not existential, risks to humanity, it is important to look at how it will work for the benefit of society,” he noted, also stressing the importance of cooperation between the state, the academic and research community, and the private sector, as happened with Lingua.

The Safeguards

The marginalization of dialects and their symbolic weakening are often accompanied by the stigmatization of their speakers, something all the researchers emphasized. During the Lingua presentation, everyone made it clear that introducing the country’s linguistic wealth through formal education, preserving it and giving it a digital presence through innovative technologies, and empowering small language communities are necessary safeguards against cultural flattening and the homogenization of monolingualism. Technology can keep a language in the archive; only its speakers can keep it alive.

Follow tovima.com on Google News to keep up with the latest stories
Exit mobile version