English
|
Deutsch |
Русский |
Татарча |
Türkçe
Résumé
Ilnar Salimzianov — NLP / ML Engineer
Computational linguist by training — a decade building NLP and
speech-to-text systems, especially for under-resourced languages, and taking
them from data pipeline to production and on-device deployment.
Open to full-time roles — employment or long-term contract — remote,
or on-site/hybrid in the EU or the Balkans (open to relocation).
Hire me →
Websites
ifs.name, taruen.com
Contact
Email: ilnar@selimcan.org,
Telegram,
LinkedIn, GitLab, GitHub
Please also see the Projects page.
Programming skills
- Languages
-
Python, Racket, Clojure, GNU Bash
- Experience with
-
Data & Backend Infra: Elasticsearch, FastAPI, Pandas, Scrapy, Docker | MLOps & NLP: Azure
ML, ONNX Runtime, HF Transformers, Scikit-learn, Spacy, NLTK |
Low-Resource NLP: Apertium, HFST, VISL CG-3 | Testing: Pytest, pytest-bdd, Locust
Education
- University of Stuttgart (2014-2017)
- M.Sc. degree in Computational Linguistics
- Kazan State University (2006-2011)
- Specialist's degree in German Philology, focus Linguistics
Natural Languages
Tatar (native), Russian, German (C1, TestDaF 5,5,5,5), English (C1, TOEFL
iBT 112), Kazakh, Turkish
Experience
- 07/2025-08/2026 Regional Language Researcher (Independent
Contractor)
- Mozilla Data Collective (Remotely)
-
- Sourced and helped external dataset owners publish ML-ready
corpora on the Mozilla Data Collective — including Armenian speech
corpora (question–answer dialogues; refugee testimonials), a Kyrgyz
human-annotated NER dataset, a Kyrgyz LLM-evaluation benchmark, and a
Georgian logical-reasoning LLM benchmark.
- Onboarded 8 archival speech corpora from the INEL project at the
University of Hamburg (Institute for Finno-Ugric/Uralic Studies) and
25 news-text corpora from a major international broadcaster.
- Created, scraped, and published 12 ML-ready datasets under the
Taruen account: folklore and literature corpora (Tatar, Kyrgyz,
Dolgan, Finnish, Polish), proverb collections (Kazakh, Uzbek, Kumyk,
Chechen, Bulgarian), a Chuvash TTS speech set, and a structured World
Factbook JSON dataset.
- Wrote datasheets and standardized packaging so each corpus is
directly ingestible for downstream LLM/NLP training.
- 01/2024-12/2025 Computational Linguist / NLP Developer (Independent
Contractor)
- US LegalTech Startup (Remotely)
-
- Architected and deployed an end-to-end data ingestion pipeline,
scraping and processing over 20 million public trademark records from
the USPTO API.
- Designed and optimized a highly complex Elasticsearch schema
capable of executing multi-layered similarity queries, including
phonetic matching, orthographic homoglyphs, and cross-lingual semantic
translations.
- Built a RESTful API utilizing FastAPI to serve real-time trademark
clearance search results to downstream frontend applications.
- Drove development test-first: wrote business-readable acceptance
tests with pytest-bdd (Given/When/Then scenarios with example tables),
so non-technical stakeholders could review and sign off on
behavior.
- Implemented automated data clearance pipelines to isolate and
verify so-called distinctive components within registered
trademarks.
- 08/2021-10/2023 Computational linguist / NLP developer
-
Orpheus Technology Ltd (prowritingaid.com)
-
- Helped launch ProWritingAid's premium-tier tone detection feature
by training and fine-tuning models using Scikit-learn and Huggingface
Transformers.
- Deployed models on Microsoft Azure ML and monitored performance
over time to ensure zero downtime.
- Converted PyTorch models into ONNX format, optimizing and
quantizing them for specific Azure ML hardware using the Huggingface
Optimum library.
- Load-tested Azure ML endpoints using Locust to right-size instance
types; the ONNX-quantized models let a workload the team had
originally scoped for 8 GPU machines run on 1 (with a second added for
redundancy), at negligible accuracy loss.
- Used Large Language Models (ChatGPT, Claude) for data generation
and labeling, and fine-tuned models via the OpenAI API.
- Prototyped GUI apps in Racket (gui-easy) and Clojure (cljfx) to
explore combining ProWritingAid with LLMs.
- 01/2018-11/2021 Remote research assistant/computational
linguist
- Nazarbayev University, Nur-Sultan, Kazakhstan
-
- Trained a new speech-to-text system for Kazakh using the Coqui STT
framework (details: https://arxiv.org/abs/2107.10637).
- Wrote a web interface to an existing ESPnet-based Kazakh
speech-to-text system and packaged the app into a Docker image.
- Gathered data and led the launching of https://commonvoice.mozilla.org/kk and https://commonvoice.mozilla.org/tt.
- Extended the Kazakh morphological transducer apertium-kaz with new
stems, affixes, and a Constraint Grammar including dependency parsing
capabilities in the Universal Dependencies framework.
- 07/2017-03/2018 Remote software contractor
- Central Eurasian Studies Department, Indiana University, Bloomington
(United States)
- Developed a closed domain, rule-based Tatar-to-English machine
translator to translate Tatar population records (1828-1918).
- 04/2015-03/2017 Research assistant
- Institute for Natural Language Processing, University of
Stuttgart
- Developed a dependency-based sentence simplifier/compressor for
German.
- 05/2014-08/2014 Student developer
- Apertium project as part of the Google Summer of Code 2014
programme
- Developed a prototype Tatar-Russian machine translator.
- 05/2012-08/2012 Student developer
- Apertium project, Google Summer of Code 2012 programme
- Developed a rule-based Kazakh-Tatar machine translator.
Independent Coursework
- MITx 6.00.1x: Introduction to Computer Science and Programming Using
Python = 93% (on edx.org)
- MITx 6.00.2x Introduction to Computational Thinking and Data Science =
88% (on edx.org)
- UTAustinX UT.5.02x Linear Algebra - Foundations to Frontiers = 62% (on
edx.org)
- SPD1x: Systematic Program Design - Part 1
- Andrew Ng's "Machine Learning" course
Home | Hire me | Résumé | Projects | Publications | Talks | Now | Email | Reading
log | Movies log