First published: September 21, 2026. Last update: September 22, 2026.
On August 31, 2026, after about one year and one month since starting, my initially 6-month contract with Mozilla Data Collective (MDC) has come to an end. It is a great opportunity to wrap up what's been done, take a break and reflect on what to do next.
Mozilla Data Collective is a platform where language data owners can share their data, under their control and terms. It has more than a thousand datasets now, both free and paid, and is still growing! Go check it out!
My primary responsibility at MDC was to land, or help others to land, new datasets on the platform. Initially, the languages I was responsible for were from the regions Central Asia and the Caucasus, but later this was extended to also include languages of Europe and Russia.
The majority of the datasets I was involved with came from other people and organizations, and are hosted under their accounts on MDC. There are too many to list them all here. I wish to express my sincere gratitude to all data providers who decided to host their datasets on MDC. It's been a pleasure to work with you!
There are also a few datasets which I have rights to distribute under Taruen's account on MDC. (Taruen is a custom language technology and data studio that I run.) The vast majority of these datasets are in the public domain. The full list can be found under Taruen's account on MDC.
I wish to thank the Mozilla Foundation and the Mozilla Data Collective for the opportunity provided! I'm very happy to have taken part in this important initiative at its very beginnings. Special thanks to Francis Tyers, Robert Pugh, Kostis Zarkias, Christine Kim, Michelle Ulea-Vang, and fellow regional language researchers Emmanuel Ngue Um, Yacub Fahmilda, Meesum Alam, and others.
So, what's next? If in short, the goal is to do more of machine learning work (especially now with so many new datasets on Mozilla Data Collective!). There is a project that has been sitting in the back of my mind for some time now, and finally I am able to give it some attention. I'm talking about Taruen Edge, a small set of language tools that run entirely on your own device and can work offline — no audio or text ever leaves your browser. It already does speech-to-text with the stock Whisper-small. Next, I want to fine-tune that Whisper-small model on Common Voice Tatar and ship it quantized to Taruen Edge, so that Tatar speakers get fast, private, offline transcription for free.
I'll write a separate post about this, so stay tuned — or subscribe to the newsletter below.
Home | Work with me | Résumé | Projects | Publications | Talks | Now | Email | Reading log | Movies log