A teacher stands with his back to the camera, instructing a boardroom filled with attentive students. The teacher gestures emphatically, while referencing a laptop screen that says “OpenAI”.
Paul Azunre, founder of Khaya.ai, presents his work to faculty and students at the School of Engineering Sciences, University of Ghana, Legon. Credit: Peter Ayorigo

The Distributed AI Research Institute (DAIR) protects and promotes underserved communities and language groups’ work to design, develop, maintain, and govern AI systems.

  

Artificial intelligence (AI) tools are prolific in English, cropping up in everyday activities and services.

Yet for the roughly 7 million people whose native language is Ethiopia’s Tigrinya, the functionality, use, and application of AI is much more complicated.

One of 70 or so languages spoken in Ethiopia, Tigrinya is considered a “low-resource language” in the AI world. A low-resource language is one that lacks sufficient digital data and linguistic resources available to train AI systems, making it more difficult for AI tools to understand and perform in that language. People speaking it and other low-resource languages are at a distinct disadvantage with technology compared to those speaking English or other predominant world languages.

It plays out in a fundamental way: AI tools translating these languages, which also include Vietnamese, Navajo, and Haitian Creole, are dysfunctional. Something as mundane as sending a dictated text or reading a locally relevant Wikipedia entry might require a translator or simply cannot be done. A more daunting result is that AI tools in low-resource languages are much more vulnerable to malicious, harmful content.

Those shortcomings underscore how AI exacerbates the marginalization and isolation of already under-resourced societies.

That is the challenging environment where the Distributed Artificial Intelligence Research Institute (DAIR) works to achieve a more equitable alternative.

 

“One of our core missions is to build and execute on the alternative tech future that we want.”

DAIR, an AI research organization, protects the rights and interests of underserved communities and language groups by enabling their participation in designing, developing, maintaining, and governing AI systems in their regions.

“One of our core missions is to build and execute on the alternative tech future that we want,” said Timnit Gebru, a Stanford-educated engineer who was born and raised in Ethiopia. She founded DAIR in 2021.

“There is this assumption,” Gebru added, “that someone has to have a certain amount of literacy or that they have to speak a certain language in order for technology like AI to work very well for them.”

DAIR flips that assumption.

Tech should adapt to fit and meet humans where they are, not the other way around, said Gebru, who co-founded Black in AI, a nonprofit strengthening the role of Black people in the field.

The more equitable alternative calls on technology companies to consider the full impact of emerging AI-related tools, such as voice transcription or language translation.

John Palfrey talks to Timnit Gebru, Founder and Executive Director of the Distributed AI Research Institute, about the evolution and ethics of AI, and the ways new technology can positively impact communities around the world.

Beyond advocating for existing tools and platforms to be designed and developed in the languages of the populations they serve, DAIR seeks to develop local talent as data modelers, who organize and structure data so AI can use it effectively, and system designers, who create the overall AI infrastructure for how all the components fit together—moving away from modifying a Silicon Valley-developed tool for local language.

DAIR also works to build new tools and platforms that use local, responsibly sourced data—instead of files scraped from the web, a practice that fails to screen out racist, inaccurate, or otherwise damaging materials from incorporation into AI models.

In addition, tech companies must understand and take responsibility for how their product could be used and ensure systems are in place to prevent labor theft or abuse, Gebru said.

It is an ambitious vision that expands beyond Tigrinya.

DAIR has teamed up with other African partners to form the Huniki Federation. They are developing text translation, voice-to-voice translation, and speech recognition platforms for several languages, including Ghana-based Twi, Ewe, and Dagbani, Ethiopia-based Amharic, and South Africa-based isiZulu.

Gebru’s model for DAIR and the inspiration for the Huniki Federation emerged from witnessing the effect that poor oversight of U.S.-developed social media tools had on her communities during the 2020-21 Tigray War in Ethiopia. More than 600,000 civilians were killed in the conflict.

Many popular platforms in Ethiopia—Facebook and TikTok, as examples—were not designed to be monitored for inflammatory or violent commentary in low-demand languages like Tigrinya, according to Amnesty International reports. The shortcoming exacerbated an already catastrophic conflict and underscores the importance of locally developed and linguistically knowledgeable AI tools.

Nuredin Ali, a research intern at the Distributed AI Research Institute (DAIR), studies human-AI Interaction with a focus on designing and implementing machine learning approaches that prioritize the human element and account for diverse societal and cultural contexts.

“I've seen that an Ethiopian company focused on Ethiopian languages means that they understand the context of those languages,” Gebru said. “They understand the sources of the data and what they mean.”

Huniki Federation members are ideally suited for those roles.

Side-by-side speech bubbles demonstrating text in English being translated into Twi by an AI model. English text:

An automated translation of English into Akuapem Twi, generated by Khaya.ai.

Individual research groups are developing models for their own country’s language to be used by local speakers and businesses. Huniki is supporting the individual organizations and plans to push for a broader system that can handle many of Africa’s languages.

Sharing resources across the federation also enhances the global visibility and credibility of the African language processing developers, a critical distinction when others are attempting to dominate the market by building products that compare poorly to those developed locally.

“We have to focus on data quality, and we have to treat it carefully in a way that Big Tech does not,” said Paul Azunre, founder of Khaya.ai, a Ghana-focused AI service offering translation, transcription and generation tools in over 30 languages. Born in Ghana, Azunre studied computer programming at Swarthmore College and MIT, using the opportunity to work on one of the first speech recognition tools for several African languages.

“They’ll (Big Tech) come up with all these automated ways of generating data, which they wouldn't do for their own languages,” he said, “because it doesn't work as well as actually creating the data and doing it properly.”

A young woman in a colorful blouse and red headscarf sits at a table, resting her chin on her hand while holding a smartphone, with other people in the background.

Attendee Saudiyatu Sulleyman listens to Paul Azunre's presentation at the Bolga Technical University, Bolga, Ghana. Credit: Victor Awmakoba

High-Performing Translation

Huniki requires its members to source data ethically, a practice that has made Tigrinya- and Amharic-focused Lesan.ai a much stronger product, said Asmelash Teka Hadgu, a Huniki partner and co-founder of Lesan.ai.

“If we relied on the old, simple Silicon Valley mantra of scraping the web, we would miss the full range of populations in Tigray,” Hadgu said.

Instead, Lesan has reached out to the Tigray community, paying for expertise in various dialects, and treating those working on the model as partners.

Side-by-side speech bubbles demonstrating text in English being translated into Tigrinya by an AI model. English text:

A sample of English into the Enderta Tigrinya dialect, translated by the team at Lesan AI.

“We've gathered some unique data sets, things you cannot find online,” Hadgu said, including colloquial speech and different dialects. “We're able to resource that from our community members, because as soon as you tell them we want your dialect to be part of this technology, they are very, very happy.”

As a result, Lesan’s Amharic and Tigrinya machine translation system has been demonstrated to outperform translation models created by Google Translate and Microsoft Translator. Local Tigray residents say the tool has enabled them to tell the stories of people who endured the war.

And while DAIR, the Huniki Federation, and its individual language developers are relatively new to the language technology space, their early progress has attracted attention from prospective members and investors. As part of an aggressive outreach plan to establish linguistic representation throughout Africa, the federation is currently finalizing membership requirements.

“It's more than profit, or the number of users—it's really about preserving culture,” Azunre said. “We also think that leads to better models.”

Since 2023, MacArthur has provided $2.25 million in general operating and flexible support to DAIR.