An essay for the Ada Lovelace Institute arguing that mainstream large language models marginalise linguistic communities and risk amplifying real injustice for people who don't speak a "high-resource" language.
Motivation
Machine translation and LLMs are still built for a few of the world's 7,000+ languages. When systems that shape immigration decisions, public services, or access to information rely on those tools, the languages left out become a gateway to exclusion, an example of what philosophers call epistemic injustice, and what I describe here as a kind of computational silencing.
Approach
The piece traces how this gap shows up in practice: from asylum interviews to everyday digital life. I argue that decolonising Language AI means centring minoritised language communities in how models and datasets get built, not treating their languages as an afterthought.
Lexical datasets as a starting point
One concrete direction the piece points to: lexical datasets built with and for minoritised language communities, which can support the development of models for languages that mainstream LLMs currently ignore.
Accurate language technology isn't a matter of comfort, for many people, it's a gateway to their rights.
Outcomes & impact
Published as a guest article on the Ada Lovelace Institute's blog, this piece connects my PhD research on low-resource LLMs to a wider public conversation about AI ethics and linguistic justice.
Collaborators
- Solo-authored; published by the Ada Lovelace Institute
Credits & disclosure
Header image: From the blog page of the Ada Lovelace Institute.
None of the text above was written or generated by AI.
