An open-source pipeline for mobile-assisted language learning and data collection, built for endangered language communities with a focus on primarily oral languages. The app teaches users to read and write their own language and lets communities collect and own their own data. This project is currently being evaluated with Dzardzongke (Nepal) and Puno Quechua (Peru) speakers.
Motivation
Language technology for endangered languages has to centre the communities it serves. Communities are increasingly asking for more representation online and a role in collecting data for their own use, but most tooling isn't built with that in mind. This project builds an open pipeline any community can use, not just the ones with in-house engineering support.
Approach
The app combines vocabulary learning, literacy support through a newly standardised romanised orthography, cultural heritage content, a predictive-text keyboard, and a built-in data collection tool. We're now evaluating it with Dzardzongke and Puno Quechua speakers spanning teenagers to elders, using task completion, an adapted System Usability Scale, and semi-structured interviews and focus groups.
Five evaluation questions
The evaluation asks: how usable is the app across different levels of digital literacy? Does it improve reading and writing in the standardised orthography? How does each community respond to the cultural heritage content? How accessible is the pipeline itself? And where do needs diverge between language communities?
Even the existence of a digital tool has proven symbolically significant for oral and endangered language communities, but symbolic significance isn't enough. We need empirical evidence of real educational and cultural impact.
Outcomes & impact
Early feedback shows strong engagement with the cultural heritage content, while the orthography-based activities need more scaffolding for participants without prior Latin-script literacy — feedback we're feeding straight back into the next iteration. The goal is a methodology and pipeline that other endangered-language communities can adopt without needing specialist app-development expertise.
Collaborators
- Anna Korhonen (PI), University of Cambridge
- Elwin Huaman, University of Cambridge
- Centre for Human-Inspired AI
- Partner institutions in Nepal, and Peru
Credits & disclosure
Header image: "Textiles and Tech 2" by Hanna Barakat, part of the Better Images of AI project. Hanna Barakat & Archival Images of AI + AIxDESIGN / Textiles and Tech 2 / Licenced by CC-BY 4.0
None of the text above was written or generated by AI.
