Ilnar Salimzianov's Personal Site
English
|
Deutsch |
Русский |
Татарча |
Türkçe
Work with me
NLP / ML engineer specializing in speech-to-text and low-resource
languages. I take language and speech models from notebook to production
— fine-tuning, ONNX optimization, APIs, deployment, and load-testing — and I
do it for languages the big cloud providers handle poorly or not at all.
Computational linguist by training (M.Sc., University of Stuttgart,
2017), with a decade building NLP, speech-to-text, and MLOps systems across
industry and research.
Availability
About 20 hours per week, across 3 days — a deliberate, stable
commitment, so the rest of my week can go to my own language-technology work
at taruen.com and selimcan.org. You get a focused
senior contractor who is in this for the long run, not someone who will
churn the moment a full-time role appears.
Remote, worldwide. My working hours overlap European business
hours and US-East mornings. I'm based in Istanbul and relocating to Serbia —
putting me squarely in the European working day.
How I can help
- Speech-to-text for your language or domain
- Fine-tune, evaluate, and deploy automatic speech recognition (Whisper
and others), including languages not supported out of the box. Honest
WER/CER benchmarking, not vendor hand-waving.
- ML models from notebook to production
- Take a working model and make it deployable and affordable: ONNX
conversion and quantization, a FastAPI service, Docker, load-testing, and
cost optimization. (At ProWritingAid, ONNX conversion and quantization let
a workload the team had scoped for 8 GPU machines run on 1 — plus a warm
standby for redundancy — at negligible accuracy loss.)
- NLP data pipelines at scale
- Scrape, clean, and structure large text corpora for NLP/LLM training
and search — for example, an end-to-end ingestion pipeline over 20M+ USPTO
trademark records feeding a multi-layered Elasticsearch similarity
engine.
- Low-resource & multilingual NLP
- Morphological analysers, machine translation, data collection, and
evaluation for Turkic, Caucasian, and other under-resourced languages.
Eight publications and several shipped open-source tools in this
area.
Selected results
- Shipped ProWritingAid's premium-tier tone-detection feature —
training, fine-tuning, Azure ML deployment, and ONNX optimization.
- Built an end-to-end pipeline over 20M+ USPTO trademark records with
phonetic, orthographic, and cross-lingual similarity search.
- Fine-tuned Coqui STT for Kazakh (arXiv:2107.10637)
and Whisper for Laz, and launched
Common Voice for Kazakh and Tatar.
Get in touch
The fastest way to start is a short message describing what you're
building and the outcome you need.
Home | Work with me | Résumé | Projects | Publications | Talks | Now | Email | Reading
log | Movies log