Ilnar Salimzianov's Personal Site

English | Deutsch | Русский | Татарча | Türkçe

Automatic Speech Recognition for Tatar

First published: October 6, 2026. Last update: October 11, 2026.

Summary

I fine-tuned OpenAI's Whisper-large-v3 to transcribe spoken Tatar. On speakers and sentences it never heard in training, it gets 4.3% of characters wrong, down from 32% for the stock model. That makes it the most accurate Whisper model for Tatar among those I could evaluate, and the one with the most readable output of all models tested: it writes capital letters and punctuation. Both training runs together cost about $24 of rented GPU time.

Between my two runs, the recipe stayed the same and only the data changed. The default Common Voice splits leave just 2 speakers for training; building my own splits raised that to 9 speakers in run 1 and 165 in run 2, and run 2 makes 16% fewer errors than run 1.

Try it in the demo, or download the weights: run 2 (recommended) and run 1.

Why Tatar ASR

At the end of a previous post, I announced an automatic speech recognition (ASR) model for Tatar: a fine-tuned Whisper. This post describes how I made it, and how it compares to what already exists.

The goal is captions for Tatar YouTube videos. Tatar is a Turkic language with several million speakers, and a small but steady number of YouTubers make videos in it. YouTube's automatic captions run on Google Cloud Speech-to-Text, which does not support Tatar, although it does support Azerbaijani, Kazakh, Kyrgyz, Turkish and Uzbek. Every Tatar YouTuber is bilingual, and each faces the same question: keep making videos in Tatar, or switch to a bigger language and get more views? Machine translation from Tatar is already decent, whether with Google Translate, DeepL or an LLM. What is missing is the tedious first step: typing the Tatar captions. With those in place, translated subtitles become easy, and a Tatar video can reach a wider audience.

I also have a personal link to the data. Years ago, I helped launch the Tatar version of Mozilla Common Voice, the website where volunteers recorded the dataset I used for fine-tuning.

Finally, friends at the Yasalma community pointed me to the BuzzASR paper, which fine-tunes Whisper-large-v3 for many languages, but not for Tatar. This post fills that gap.

Mine is not the first Tatar ASR model, and that points to a wider problem. Big providers skip languages like Tatar because the market is small. Enthusiasts train models for the languages they care about, but rarely keep them running for people who can't run Python. The most accurate models need a GPU, which costs from about $0.30 an hour. There are cheaper ways to serve them: free tiers such as Google Colab and Kaggle (meant for developers, not end users), Hugging Face's ZeroGPU Spaces (included in its $9-a-month Pro plan), serverless GPU providers that bill only for the seconds used, and small models that run on the user's own device, even in the browser. That last option is what my side project https://edge.taruen.com is about, and it will get a post of its own.

These are the Tatar ASR models I found:

  1. https://huggingface.co/AigizK/wav2vec2-large-mms-1b-tatar-v2
  2. https://huggingface.co/yasalma/whisper-finetuned-tt-asr
  3. https://huggingface.co/AigizK/wav2vec2-large-mms-1b-tatar
  4. https://huggingface.co/yasalma/whisper-finetuned-tatartts-asr
  5. https://huggingface.co/anton-l/wav2vec2-large-xlsr-53-tatar
  6. https://huggingface.co/crang/wav2vec2-large-xlsr-53-tatar
  7. https://huggingface.co/emre/wav2vec2-large-xlsr-53-W2V2-TATAR-SMALL
  8. https://huggingface.co/infinitejoy/wav2vec2-large-xls-r-300m-tatar
  9. https://huggingface.co/kingabzpro/wav2vec2-large-xls-r-300m-Tatar
  10. https://huggingface.co/sammy786/wav2vec2-xlsr-tatar
  11. https://huggingface.co/ctaguchi/wav2vec2-xls-r-300m-tatar
  12. https://github.com/IS2AI/Soyle
  13. https://github.com/IS2AI/TurkicASR
  14. https://github.com/facebookresearch/omnilingual-asr

For a good introduction to speech recognition, see chapters 15 and 16 of the upcoming 3rd edition of Jurafsky and Martin's Speech and Language Processing, which the authors share online.

For a long time, ASR was dominated by Hidden Markov Models (HMMs). Chapter 16 describes the two modern approaches:

The models in the list above follow one of these two approaches, or, like Omnilingual's LLM-ASR variant, combine them.

For each model, I looked at three things:

  1. the base model and its licence;
  2. for fine-tuned models, the datasets they were fine-tuned on, and the licences of those datasets;
  3. accuracy: word error rate (WER) and character error rate (CER), including CER on the raw output, with capitalization and punctuation kept.

These are the Tatar speech datasets I found:

Objectives

  1. Evaluate the 14 ASR models listed above on Common Voice Scripted Speech 27.0 Tatar.
  2. Fine-tune Whisper-large-v3 on the same dataset.

For this first iteration, I trained only on Common Voice Scripted Speech (CC0). ISSAI's TatSC and TatarTTS also allow commercial use, but together they are more than ten times its size, and training time grows with them; adding them is the next step. The Spontaneous Speech dataset is CC0 too, but at 14 clips it is too small to matter yet.

The data

Common Voice sorts its clips into buckets. The validated bucket holds every clip that passed community review: for Tatar, 30,555 clips from 270 speakers. The official train, dev and test splits are subsets of it, chosen so that each sentence is used only once and no speaker appears in more than one split.

For Tatar, this gives a lopsided result. The official test split holds 258 of the 270 speakers. The official training split has 8,249 clips from just 2 speakers. And 12,558 validated clips are in no split at all, mostly further readings of sentences that are already used. On top of that, contributions are very uneven: one speaker has recorded about 12,000 clips, while most have recorded a few dozen.

Training on the official training split would mean learning Tatar from two voices. So for both of my training runs, I built my own training and validation sets, always keeping them separate from the test data in two ways: no speaker and no sentence may appear on both sides.

Run 1 kept the official test split unchanged, so that its results are comparable with anyone else's. For training and validation, it used every validated clip that shares neither a speaker nor a sentence with the test split. That left just 11 speakers. Two of them, with few clips each, became the validation set (628 clips, one per sentence). The other 9 became the training set: 19,605 clips, more than twice the official training split, but still with one speaker accounting for 58% of it.

Run 2 re-split the data from scratch, to train on many more voices. The test and validation sets were drawn only from clips in the official test split, which run 1 never trained on, so both runs can be compared on them fairly. The test set takes up to 40 clips from each of 80 randomly chosen speakers (1,122 clips), and the validation set the same from 25 other speakers (418 clips), with one clip per sentence in both. Every other validated clip that shares neither a speaker nor a sentence with them became training data: 26,855 clips from 165 speakers. One speaker still accounts for 52% of it.

Evaluation setup

All models were evaluated on two test sets:

The main metric is the character error rate (CER). Tatar is agglutinative, so words are long, and a single wrong suffix makes a whole word count as wrong. CER shows more fairly how much correcting a transcript takes. I also report the word error rate (WER), and CER in two versions: normalized, after lowercasing and removing punctuation, and raw, comparing the output exactly as produced. Raw scores matter for readable subtitles; normalized scores are fair to models that were never trained to produce capitals and punctuation.

All Whisper models were decoded with the same settings, following BuzzASR, so that differences come from the models, not from decoding.

Results

Tatar ASR models: base model, fine-tuning data, and the licences of each.
Model Base model (licence) Fine-tuned on (licence) Licence
AigizK/wav2vec2-large-mms-1b-tatar, with its language model facebook/mms-1b-all (CC BY-NC 4.0) not documented CC BY-NC 4.0
sammy786/wav2vec2-xlsr-tatar XLS-R, probably the 1B version (Apache-2.0) Common Voice 8.0 (CC0) Apache-2.0
My run 2 (165 training speakers) openai/whisper-large-v3 (MIT) Common Voice Scripted Speech 27.0 (CC0) MIT
My run 1 (9 training speakers) openai/whisper-large-v3 (MIT) Common Voice Scripted Speech 27.0 (CC0) MIT
infinitejoy/wav2vec2-large-xls-r-300m-tatar facebook/wav2vec2-xls-r-300m (Apache-2.0) Common Voice 7.0 (CC0) Apache-2.0
anton-l/wav2vec2-large-xlsr-53-tatar facebook/wav2vec2-large-xlsr-53 (Apache-2.0) Common Voice 6.1, train + dev (CC0) Apache-2.0
yasalma/whisper-finetuned-tt-asr openai/whisper-small (MIT) TatSC (CC BY 4.0), tat_hackathon_asr (unclear), Common Voice (CC0) Apache-2.0
crang/wav2vec2-large-xlsr-53-tatar facebook/wav2vec2-large-xlsr-53 (Apache-2.0) Common Voice, train + dev, version not stated (CC0) Apache-2.0
kingabzpro/wav2vec2-large-xls-r-300m-Tatar facebook/wav2vec2-xls-r-300m (Apache-2.0) Common Voice 8.0 (CC0) Apache-2.0
emre/wav2vec2-large-xlsr-53-W2V2-TATAR-SMALL facebook/wav2vec2-large-xlsr-53 (Apache-2.0) Common Voice, version not stated (CC0) Apache-2.0
AigizK/wav2vec2-large-mms-1b-tatar-v2 (no language model) facebook/mms-1b-all (CC BY-NC 4.0) not documented CC BY-NC 4.0
ctaguchi/wav2vec2-xls-r-300m-tatar XLS-R, size unclear (Apache-2.0) probably Common Voice 16.1 (CC0) not stated
yasalma/whisper-finetuned-tatartts-asr Whisper, size not stated (MIT) not documented not stated
openai/whisper-large-v3, without fine-tuning — — MIT
IS2AI/Soyle Whisper, size not stated (MIT) Common Voice 13.0 (CC0), TatSC (CC BY 4.0) CC BY 4.0 (project page)
IS2AI/TurkicASR none: trained from scratch with ESPnet Common Voice 10.0 (CC0), for Tatar not stated
facebookresearch/omnilingual-asr none: Meta's own encoder large multilingual mix, incl. Meta's own corpus (CC BY 4.0) Apache-2.0
Error rates in percent; lower is better. CER and WER are computed after lowercasing and removing punctuation; raw CER keeps capitalization and punctuation. Run 2 test: 1,122 clips from 80 speakers. Official test: the official test split, 4,983 clips from 258 speakers.
Run 2 test Official test
Model CER WER raw CER CER WER raw CER
AigizK/wav2vec2-large-mms-1b-tatar, with its language model 3.2 19.3 10.6 3.1 18.9 10.2
sammy786/wav2vec2-xlsr-tatar 3.9 17.6 11.1 3.9 17.5 10.9
My run 2 (165 training speakers) 4.3 18.2 6.0 — — —
My run 1 (9 training speakers) 5.2 22.8 6.7 5.5 23.6 7.0
infinitejoy/wav2vec2-large-xls-r-300m-tatar 5.3 24.3 12.4 5.4 25.2 12.3
Omnilingual omniASR_LLM_1B_v2 5.5 22.9 12.5 5.1 21.6 11.9
Omnilingual omniASR_LLM_7B_v2 6.5 27.2 13.5 6.5 27.2 13.2
anton-l/wav2vec2-large-xlsr-53-tatar 7.3 28.8 14.2 7.3 29.1 14.0
yasalma/whisper-finetuned-tt-asr 7.9 32.3 9.7 8.2 33.2 10.0
crang/wav2vec2-large-xlsr-53-tatar 8.1 32.3 14.9 8.2 33.0 14.8
kingabzpro/wav2vec2-large-xls-r-300m-Tatar 9.0 38.3 15.9 9.0 39.2 15.6
emre/wav2vec2-large-xlsr-53-W2V2-TATAR-SMALL 9.0 36.3 15.9 9.4 37.2 15.9
AigizK/wav2vec2-large-mms-1b-tatar-v2 (no language model) 9.1 34.8 16.6 9.2 35.3 16.3
ctaguchi/wav2vec2-xls-r-300m-tatar 12.3 49.2 18.9 12.5 50.0 18.8
yasalma/whisper-finetuned-tatartts-asr 13.1 50.4 14.9 13.7 51.2 15.3
Omnilingual omniASR_CTC_1B_v2 21.2 53.6 26.8 20.5 52.7 25.9
openai/whisper-large-v3, without fine-tuning 32.1 89.3 33.5 33.0 90.2 34.3
IS2AI/Soyle not evaluated: the weights linked from its README (dhcppc0/soyle_onnx) were unavailable in October 2026
IS2AI/TurkicASR not evaluated: it needs a 2022 ESPnet setup

Notes on Tables 1 and 2:

Stock Whisper-large-v3, with no Tatar-specific training, gets about 32% CER. That is the baseline any fine-tuning has to beat.

Among the existing models, sammy786's wav2vec2 model is the strongest acoustic model, at 3.9% CER. AigizK's two repositories hold the same acoustic model, and one of them adds an n-gram language model. That language model alone cuts CER from 9.1% to 3.2%, and to 2.2% once a bug that drops every й is accounted for (see the notes above).

Two models beat mine on normalized CER: AigizK's with its language model, and sammy786's. On raw output, mine win clearly: 6.0% raw CER for run 2 and 6.7% for run 1, against 9.7% for the next best model and 10.6% or more for all the others. The wav2vec2 and Omnilingual models write mostly or entirely in lowercase without punctuation, so their output needs editing before it can be used as subtitles.

Licences matter for anyone building a product. AigizK's models build on MMS, which is licensed for non-commercial use only. sammy786's model is Apache-2.0, and mine are MIT.

Checking for contamination

Common Voice keeps growing, and models trained on older releases may have seen clips that are now in the current test set. That would make them look better than they are.

To check, I grouped the test clips by when they were recorded, and used my run 1 model, which certainly never saw them, as a control. Newer clips turned out to be harder for every model, including the control, and sammy786's lead stayed constant across clip ages. Had the model memorized older clips, its advantage would shrink on newer ones. It doesn't, so I see no sign of contamination. Only the file lists of its original training data (Common Voice 8.0) could rule it out completely.

Fine-tuning

I followed the BuzzASR recipe: full fine-tuning of Whisper-large-v3, with a lower learning rate for the encoder than for the decoder, a cosine learning rate schedule, and early stopping on the validation set. The scripts are in the appendices.

Unlike my earlier Laz model, I kept capital letters and punctuation in the training labels. Subtitles have to be readable, and lowercase text without punctuation is tiring to read.

Training ran on a rented A100 on RunPod, for about $24 in total. Three lessons from the process:

Run 1 against run 2

Run 1, trained on 9 speakers, reached 5.2% CER on run 2's test set. Run 2, trained on 165 speakers with the same recipe, reached 4.3%.

Only the data changed between the runs: run 2 had 18 times as many speakers and 37% more clips. Both probably helped, and I didn't separate the two effects. Either way, fixing the splits gave a clear gain without touching a single hyperparameter.

Smaller models for the browser

Whisper-large-v3 is far too big to run in a browser, and even Whisper-small proved too slow there. So the next step is to fine-tune Whisper-tiny and Whisper-base the same way and run them on Taruen Edge. I will update this post with how much accuracy the smaller models give up.

Limitations

Next steps

Try it

The demo transcribes uploaded or recorded audio with either model and gives you an SRT subtitle file. The weights are on Hugging Face under the MIT licence: run 2 and run 1.

I'm open to ML, NLP and speech roles: remote, hybrid or on-site.

Hire me →

Appendix A: finetune.py for run 1

This is the script exactly as it ran on the GPU pod for run 1 (official test split, 9 training speakers), copied from the pod before it was deleted.

Download: finetune-run1.py (661 lines)

Show the code
"""
Fine-tune Whisper on Common Voice Scripted Speech 27.0 Tatar.

The recipe follows the simple fine-tuning (SFT) setup of BuzzASR
(arxiv:2609.09554): full fine-tuning of Whisper-large-v3, an encoder learning
rate of 0.3x the decoder's, a cosine schedule, early stopping on dev, and greedy
decoding with repetition penalties.

Data, all drawn from the validated bucket:
- test:  the official CV test split, untouched, so results compare with eval.py
- dev:   about 500 clips from randomly chosen light speakers, one per sentence
         (kept small: the test split holds most speakers, so few are left)
- train: every other validated clip, sharing no speaker and no sentence with
         dev or test

Speakers per split in CV Scripted Speech 27.0 Tatar (clips from the TSV files):

    split        clips  speakers
    dev           4765         9
    invalidated    580       149
    other          264        19
    test          4983       258
    train         8249         2
    validated    30555       270

Test holds 258 of the 270 validated speakers, which leaves only a dozen
for train and dev together. Reproduce with:

    dataset.groupby("split").agg(
        clips=("audio_path", "size"), speakers=("speaker_id", "nunique")
    )

Labels keep their original casing and punctuation, so the model learns to write
readable subtitles. Metrics are computed on normalized text, as in eval.py.

Usage:
    python3 languages/tat/finetune.py --dry-run # show the splits, train nothing
    python3 languages/tat/finetune.py           # defaults suit an A100 80 GB
    python3 languages/tat/finetune.py --batch-size 2 --grad-accum 16 --optim-8bit
                                                # A100 40 GB
    python3 languages/tat/finetune.py --resume  # continue from the last check-
                                                # point

To evaluate the result on the test split:
    python3 languages/tat/eval.py --models languages/tat/whisper-large-v3-tt/final

Requires the datacollective fork [1] with the TSV quoting fix [2], the script
refuses to run if load_dataset returns fewer rows than the TSV files contain.

[1] https://github.com/IlnarSelimcan/datacollective-python/
[2] https://github.com/IlnarSelimcan/datacollective-python/commit/3e9312ee2d9df9d2f4438667b66cd6b0bed0a854
"""

import os
import random
from collections import defaultdict
from collections.abc import Callable, Sequence
from dataclasses import dataclass
from itertools import combinations
from pathlib import Path
from typing import Annotated, TypedDict

import librosa
import numpy as np
import numpy.typing as npt
import pandas as pd
import torch
import typer
from datacollective import load_dataset
from transformers import (
    EarlyStoppingCallback,
    EvalPrediction,
    Seq2SeqTrainer,
    Seq2SeqTrainingArguments,
    WhisperForConditionalGeneration,
    WhisperProcessor,
    WhisperTokenizer,
    set_seed,
)

from eval import (
    COMMON_VOICE_SCRIPTED_SPEECH_27_0_TATAR,
    SAMPLING_RATE,
    WHISPER_GENERATE_KWARGS,
    Audio,
    load_audio,
    score,
)

HERE = Path(__file__).resolve().parent

MODEL_ID = "openai/whisper-large-v3"
OUTPUT_DIR = HERE / "whisper-large-v3-tt"

LABEL_PAD = -100
"""Label value ignored by the loss (PyTorch's cross-entropy default)."""

SPEED_RATES = (0.9, 1.0, 1.1)
"""Speed perturbation factors, used only with --speed-perturb."""


# -- Types --------------------------------------------------------------------


class Example(TypedDict):
    """One training example: log-mel features and label token ids."""

    input_features: npt.NDArray[np.float32]
    labels: list[int]


class Batch(TypedDict):
    """A padded batch, as the model's forward pass expects it."""

    input_features: torch.Tensor
    labels: torch.Tensor


@dataclass(frozen=True)
class Splits:
    """Train, dev and test clips, one row per clip."""

    train: pd.DataFrame
    dev: pd.DataFrame
    test: pd.DataFrame

    def items(self) -> list[tuple[str, pd.DataFrame]]:
        """Return (name, frame) pairs in a fixed order."""
        return [("train", self.train), ("dev", self.dev), ("test", self.test)]


@dataclass(frozen=True)
class Config:
    """Everything that defines one training run."""

    model_id: str
    output_dir: Path
    learning_rate: float
    encoder_lr_ratio: float
    batch_size: int
    grad_accum: int
    epochs: int
    warmup_steps: int
    weight_decay: float
    eval_steps: int
    patience: int
    dev_clips: int
    max_dev_speaker_clips: int
    max_clips_per_speaker: int
    speed_perturb: bool
    optim_8bit: bool
    workers: int
    seed: int
    resume: bool
    dry_run: bool


# -- Data ---------------------------------------------------------------------


def load_common_voice() -> pd.DataFrame:
    """
    Return every CV Scripted Speech 27.0 Tatar row, after checking that none
    were dropped.
    """
    if not os.getenv("MDC_API_KEY"):
        raise RuntimeError(
            "Set the MDC_API_KEY environment variable to download the dataset."
        )
    dataset = load_dataset(COMMON_VOICE_SCRIPTED_SPEECH_27_0_TATAR)
    check_complete(dataset)
    return dataset


def tsv_row_count(path: Path) -> int:
    """
    Return the number of data rows in the TSV file at `path`
    (assuming it has a header line).
    """
    with path.open(encoding="utf-8") as f:
        return sum(1 for _ in f) - 1


def check_complete(dataset: pd.DataFrame) -> None:
    """
    Raise if `dataset` has fewer rows per split than its TSV file, which is
    what an unpatched datacollective returns.
    """
    cv_dir = Path(dataset.audio_path.iloc[0]).parent.parent
    for split, loaded in dataset.split.value_counts().items():
        expected = tsv_row_count(cv_dir / f"{split}.tsv")
        if loaded != expected:
            raise RuntimeError(
                f"load_dataset returned {loaded} '{split}' rows, but "
                f"{split}.tsv has {expected}. Install the patched datacollective."
            )


def exclude(frame: pd.DataFrame, held_out: pd.DataFrame) -> pd.DataFrame:
    """
    Return the rows of `frame` sharing no speaker and no sentence with
    `held_out`.
    """
    shared = frame.speaker_id.isin(
        held_out.speaker_id
    ) | frame.sentence_id.isin(held_out.sentence_id)
    return frame[~shared]


def pick_dev_speakers(
    pool: pd.DataFrame, target_clips: int, max_clips: int, seed: int
) -> set[str]:
    """
    Return random "light" speakers from `pool` until together they have
    `target_clips` clips.

    Contributions are very uneven: a couple of speakers recorded thousands of
    clips each, most others a few dozen. Only speakers with at most `max_clips`
    clips qualify ("light"), because every dev speaker's clips are removed from
    train to keep the two splits speaker-disjoint. A heavy speaker in dev would
    take thousands of clips out of an already small training set; a light one
    costs only a few dozen.
    """
    counts = pool.speaker_id.value_counts()
    light = counts[counts <= max_clips].sample(frac=1, random_state=seed)
    clips_before = light.cumsum().shift(fill_value=0)
    return set(light[clips_before < target_clips].index)


def cap_per_speaker(frame: pd.DataFrame, cap: int, seed: int) -> pd.DataFrame:
    """Return at most `cap` random clips per speaker from `frame`."""
    shuffled = frame.sample(frac=1, random_state=seed)
    return shuffled.groupby("speaker_id").head(cap)


def check_disjoint(splits: Splits) -> None:
    """Raise if any two splits share a speaker or a sentence."""
    for (name_a, a), (name_b, b) in combinations(splits.items(), 2):
        for column in ("speaker_id", "sentence_id"):
            shared = set(a[column]) & set(b[column])
            if shared:
                raise ValueError(
                    f"{name_a} and {name_b} share {len(shared)} values of {column}"
                )


def make_splits(
    dataset: pd.DataFrame,
    dev_clips: int,
    max_dev_speaker_clips: int,
    max_clips_per_speaker: int,
    seed: int,
) -> Splits:
    """
    Return train/dev/test splits: the official test split, a small dev set of
    light speakers, and every other validated clip as train.
    `max_clips_per_speaker` of 0 means no cap.
    """
    test = dataset[dataset.split == "test"]
    pool = exclude(dataset[dataset.split == "validated"], test)

    speakers = pick_dev_speakers(pool, dev_clips, max_dev_speaker_clips, seed)
    if not speakers:
        raise ValueError(
            f"No speaker outside test has at most {max_dev_speaker_clips} clips, "
            "so dev would be empty; raise --max-dev-speaker-clips."
        )
    dev = pool[pool.speaker_id.isin(speakers)].drop_duplicates("sentence_id")

    train = exclude(pool, dev)
    if max_clips_per_speaker:
        train = cap_per_speaker(train, max_clips_per_speaker, seed)

    splits = Splits(
        train=train.sample(frac=1, random_state=seed).reset_index(drop=True),
        dev=dev.reset_index(drop=True),
        test=test.reset_index(drop=True),
    )
    check_disjoint(splits)
    return splits


def describe(splits: Splits) -> pd.DataFrame:
    """Return clip, speaker and sentence counts per split."""
    return pd.DataFrame.from_dict(
        {
            name: {
                "clips": len(frame),
                "speakers": frame.speaker_id.nunique(),
                "sentences": frame.sentence_id.nunique(),
                "top_speaker_share": round(
                    frame.speaker_id.value_counts(normalize=True).iloc[0], 3
                ),
            }
            for name, frame in splits.items()
        },
        orient="index",
    )


def save_splits(splits: Splits, directory: Path) -> None:
    """
    Write each split's clips to `directory`/<split>.tsv, for
    reproducibility.
    """
    directory.mkdir(parents=True, exist_ok=True)
    columns = ["audio_path", "transcription", "speaker_id", "sentence_id"]
    for name, frame in splits.items():
        frame[columns].to_csv(directory / f"{name}.tsv", sep="\t", index=False)


# -- Features -----------------------------------------------------------------


def change_speed(audio: Audio, rate: float) -> Audio:
    """
    Return `audio` played `rate` times faster, pitch included, as in Kaldi-style
    speed perturbation.
    """
    if rate == 1.0:
        return audio
    return librosa.resample(
        audio, orig_sr=round(SAMPLING_RATE * rate), target_sr=SAMPLING_RATE
    )


class WhisperDataset:
    """Map-style dataset that decodes audio and tokenizes text on access."""

    def __init__(
        self,
        frame: pd.DataFrame,
        processor: WhisperProcessor,
        speed_rates: Sequence[float] = (),
    ) -> None:
        self.paths = frame.audio_path.tolist()
        self.texts = frame.transcription.tolist()
        self.processor = processor
        self.speed_rates = speed_rates

    def __len__(self) -> int:
        return len(self.paths)

    def __getitem__(self, i: int) -> Example:
        audio = load_audio(self.paths[i])
        if self.speed_rates:
            # `random` is reseeded per DataLoader worker, unlike numpy's global
            # RNG
            audio = change_speed(audio, random.choice(self.speed_rates))
        features = self.processor.feature_extractor(
            audio, sampling_rate=SAMPLING_RATE
        ).input_features[0]
        labels = self.processor.tokenizer(self.texts[i]).input_ids
        return {"input_features": features, "labels": labels}


@dataclass(frozen=True)
class Collator:
    """
    Pads examples into a batch; padded label positions are ignored by the
    loss.
    """

    processor: WhisperProcessor
    decoder_start_token_id: int

    def __call__(self, examples: list[Example]) -> Batch:
        features = self.processor.feature_extractor.pad(
            [{"input_features": e["input_features"]} for e in examples],
            return_tensors="pt",
        )
        labels = self.processor.tokenizer.pad(
            [{"input_ids": e["labels"]} for e in examples], return_tensors="pt"
        )
        ids = labels["input_ids"].masked_fill(
            labels.attention_mask.ne(1), LABEL_PAD
        )
        # the model prepends the start token itself during training
        if (ids[:, 0] == self.decoder_start_token_id).all():
            ids = ids[:, 1:]
        return {"input_features": features["input_features"], "labels": ids}


# -- Model --------------------------------------------------------------------


def load_model(
    model_id: str,
) -> tuple[WhisperForConditionalGeneration, WhisperProcessor]:
    """Return the model and processor, set up to transcribe Tatar."""
    processor = WhisperProcessor.from_pretrained(
        model_id, language="tatar", task="transcribe"
    )
    model = WhisperForConditionalGeneration.from_pretrained(model_id, dtype=torch.float32)
    model.config.use_cache = False  # incompatible with gradient checkpointing
    configure_generation(model)
    return model, processor


def configure_generation(model: WhisperForConditionalGeneration) -> None:
    """
    Decode Tatar with the same greedy settings as eval.py, both for dev
    evaluation and in the saved model.
    """
    for key, value in WHISPER_GENERATE_KWARGS.items():
        setattr(model.generation_config, key, value)
    model.generation_config.forced_decoder_ids = None
    model.config.forced_decoder_ids = None


def parameter_groups(
    model: torch.nn.Module,
    lr: float,
    encoder_lr_ratio: float,
    weight_decay: float,
) -> list[dict[str, object]]:
    """
    Return AdamW parameter groups: the encoder gets `encoder_lr_ratio` times
    the decoder's learning rate; biases and norms get no weight decay.
    """
    grouped: dict[tuple[bool, bool], list[torch.nn.Parameter]] = defaultdict(
        list
    )
    for name, param in model.named_parameters():
        if param.requires_grad:
            is_encoder = name.startswith("model.encoder.")
            decays = param.ndim >= 2
            grouped[(is_encoder, decays)].append(param)
    return [
        {
            "params": params,
            "lr": lr * encoder_lr_ratio if is_encoder else lr,
            "weight_decay": weight_decay if decays else 0.0,
        }
        for (is_encoder, decays), params in grouped.items()
    ]


def build_optimizer(
    model: torch.nn.Module, cfg: Config
) -> torch.optim.Optimizer:
    """Return AdamW, or its 8-bit version, over the model's parameter groups."""
    groups = parameter_groups(
        model, cfg.learning_rate, cfg.encoder_lr_ratio, cfg.weight_decay
    )
    if cfg.optim_8bit:
        import bitsandbytes as bnb

        return bnb.optim.AdamW8bit(groups)
    return torch.optim.AdamW(groups)


# -- Training ------------------------------------------------------------------


def make_compute_metrics(
    tokenizer: WhisperTokenizer,
) -> Callable[[EvalPrediction], dict[str, float]]:
    """Return a function scoring generated predictions with eval.score."""

    def decode(ids: npt.NDArray[np.int64]) -> list[str]:
        ids = np.where(ids == LABEL_PAD, tokenizer.pad_token_id, ids)
        return tokenizer.batch_decode(ids, skip_special_tokens=True)

    def compute_metrics(pred: EvalPrediction) -> dict[str, float]:
        texts = pd.DataFrame(
            {
                "reference": decode(pred.label_ids),
                "hypothesis": decode(pred.predictions),
            }
        )
        return score(texts)

    return compute_metrics


def use_bf16() -> bool:
    """
    Return whether the GPU supports bfloat16, which is safer than fp16 for
    training.
    """
    return torch.cuda.is_available() and torch.cuda.is_bf16_supported()


def training_args(cfg: Config) -> Seq2SeqTrainingArguments:
    """Return the Trainer settings for `cfg`."""
    return Seq2SeqTrainingArguments(
        output_dir=str(cfg.output_dir),
        per_device_train_batch_size=cfg.batch_size,
        per_device_eval_batch_size=cfg.batch_size * 4,
        gradient_accumulation_steps=cfg.grad_accum,
        learning_rate=cfg.learning_rate,  # scheduler scales each group's own rate
        lr_scheduler_type="cosine",
        warmup_steps=cfg.warmup_steps,
        num_train_epochs=cfg.epochs,
        bf16=use_bf16(),
        gradient_checkpointing=True,
        gradient_checkpointing_kwargs={"use_reentrant": False},
        eval_strategy="steps",
        eval_steps=cfg.eval_steps,
        save_strategy="steps",
        save_steps=cfg.eval_steps,
        save_total_limit=2,
        load_best_model_at_end=True,
        metric_for_best_model="cer",
        greater_is_better=False,
        predict_with_generate=True,
        generation_max_length=225,
        logging_steps=25,
        report_to=["tensorboard"],
        dataloader_num_workers=cfg.workers,
        remove_unused_columns=False,
        label_names=["labels"],
        seed=cfg.seed,
    )


def has_checkpoint(directory: Path) -> bool:
    """Return whether `directory` contains a Trainer checkpoint."""
    return directory.is_dir() and any(directory.glob("checkpoint-*"))


def build_trainer(
    cfg: Config, splits: Splits
) -> tuple[Seq2SeqTrainer, WhisperProcessor]:
    """Return a Trainer for `cfg` on `splits`, and the processor it uses."""
    model, processor = load_model(cfg.model_id)
    speed_rates = SPEED_RATES if cfg.speed_perturb else ()
    trainer = Seq2SeqTrainer(
        model=model,
        args=training_args(cfg),
        train_dataset=WhisperDataset(splits.train, processor, speed_rates),
        eval_dataset=WhisperDataset(splits.dev, processor),
        data_collator=Collator(processor, model.config.decoder_start_token_id),
        compute_metrics=make_compute_metrics(processor.tokenizer),
        processing_class=processor,
        optimizers=(build_optimizer(model, cfg), None),
        callbacks=[
            EarlyStoppingCallback(early_stopping_patience=cfg.patience)
        ],
    )
    return trainer, processor


def run(cfg: Config) -> Path | None:
    """Train according to `cfg`; return the final model's directory, or None on a dry run."""
    set_seed(cfg.seed)
    splits = make_splits(
        load_common_voice(),
        cfg.dev_clips,
        cfg.max_dev_speaker_clips,
        cfg.max_clips_per_speaker,
        cfg.seed,
    )
    print(describe(splits).to_string())
    if cfg.dry_run:
        return None

    save_splits(splits, cfg.output_dir / "splits")
    trainer, processor = build_trainer(cfg, splits)
    trainer.train(
        resume_from_checkpoint=cfg.resume and has_checkpoint(cfg.output_dir)
    )

    final = cfg.output_dir / "final"
    trainer.save_model(
        str(final)
    )  # the best checkpoint, thanks to load_best_model_at_end
    processor.save_pretrained(str(final))
    print(f"Best model saved to {final}")
    return final


# -- Main ---------------------------------------------------------------------


def main(
    model_id: Annotated[
        str, typer.Option(help="Model to fine-tune.")
    ] = MODEL_ID,
    output_dir: Annotated[
        Path, typer.Option(help="Checkpoints, splits and final model.")
    ] = OUTPUT_DIR,
    learning_rate: Annotated[
        float,
        typer.Option(
            help="Decoder learning rate (BuzzASR: 3.24e-6, 6.48e-6, 2e-5)."
        ),
    ] = 6.48e-6,
    encoder_lr_ratio: Annotated[
        float,
        typer.Option(
            help="Encoder learning rate as a fraction of the decoder's."
        ),
    ] = 0.3,
    batch_size: Annotated[int, typer.Option(help="Clips per GPU step.")] = 4,
    grad_accum: Annotated[
        int,
        typer.Option(
            help="Steps per optimizer update; effective batch = batch size x this."
        ),
    ] = 8,
    epochs: Annotated[int, typer.Option(help="Maximum epochs.")] = 6,
    warmup_steps: Annotated[
        int, typer.Option(help="Learning-rate warmup steps.")
    ] = 150,
    weight_decay: Annotated[
        float, typer.Option(help="AdamW weight decay.")
    ] = 4e-5,
    eval_steps: Annotated[
        int, typer.Option(help="Evaluate and save every N updates.")
    ] = 100,
    patience: Annotated[
        int,
        typer.Option(
            help="Stop after N evaluations without dev CER improvement."
        ),
    ] = 3,
    dev_clips: Annotated[
        int, typer.Option(help="Approximate dev set size, in clips.")
    ] = 500,
    max_dev_speaker_clips: Annotated[
        int,
        typer.Option(
            help="Only speakers with at most this many clips go to dev."
        ),
    ] = 200,
    max_clips_per_speaker: Annotated[
        int, typer.Option(help="Cap on train clips per speaker (0 = no cap).")
    ] = 0,
    speed_perturb: Annotated[
        bool,
        typer.Option(help="Randomly change speed by 0.9x/1.1x in training."),
    ] = False,
    optim_8bit: Annotated[
        bool,
        typer.Option(
            help="Use 8-bit AdamW (bitsandbytes) to save GPU memory."
        ),
    ] = False,
    workers: Annotated[
        int, typer.Option(help="DataLoader worker processes.")
    ] = 4,
    seed: Annotated[
        int, typer.Option(help="Random seed for splits and training.")
    ] = 42,
    resume: Annotated[
        bool, typer.Option(help="Resume from the last checkpoint.")
    ] = False,
    dry_run: Annotated[
        bool, typer.Option(help="Only build and describe the splits.")
    ] = False,
) -> Path | None:
    """Fine-tune Whisper on Common Voice 27.0 Tatar."""
    return run(
        Config(**locals())
    )  # the parameters are exactly Config's fields


if __name__ == "__main__":
    typer.run(main)

Appendix B: finetune.py for run 2

The script for run 2, as committed after the run. It adds the --resplit option, which builds the speaker-balanced splits described in The data; without that option it behaves like the run 1 version.

Download: finetune-run2.py (760 lines)

Show the code
"""
Fine-tune Whisper on Common Voice Scripted Speech 27.0 Tatar.

The recipe follows the simple fine-tuning (SFT) setup of BuzzASR
(arxiv:2609.09554): full fine-tuning of Whisper-large-v3, an encoder learning
rate of 0.3x the decoder's, a cosine schedule, early stopping on dev, and greedy
decoding with repetition penalties.

Data, all drawn from the validated bucket, in one of two ways.

Official split (default):
- test:  the official CV test split, untouched, so results compare with eval.py
- dev:   about 500 clips from randomly chosen light speakers, one per sentence
         (kept small: the test split holds most speakers, so few are left)
- train: every other validated clip, sharing no speaker and no sentence with
         dev or test

Speakers per split in CV Scripted Speech 27.0 Tatar (clips from the TSV files):

    split        clips  speakers
    dev           4765         9
    invalidated    580       149
    other          264        19
    test          4983       258
    train         8249         2
    validated    30555       270

Test holds 258 of the 270 validated speakers, which leaves only a dozen
for train and dev together. Reproduce with:

    dataset.groupby("split").agg(
        clips=("audio_path", "size"), speakers=("speaker_id", "nunique")
    )

Re-split (--resplit), to train on far more voices:
- test:  up to 40 clips each from 80 random speakers of the official test split
- dev:   the same from 25 other official test speakers
- train: every other validated clip, sharing no speaker and no sentence with
         dev or test (about 165 speakers instead of 9)
Test and dev come only from official test clips, which a model trained on the
official split has never seen, so both kinds of model can be compared on the
new test set (eval.py --test-tsv <output-dir>/splits/test.tsv).

Labels keep their original casing and punctuation, so the model learns to write
readable subtitles. Metrics are computed on normalized text, as in eval.py.

Usage:
    python3 languages/tat/finetune.py --dry-run # show the splits, train nothing
    python3 languages/tat/finetune.py --resplit --dry-run
    python3 languages/tat/finetune.py           # defaults suit an A100 80 GB
    python3 languages/tat/finetune.py --batch-size 2 --grad-accum 16 --optim-8bit
                                                # A100 40 GB
    python3 languages/tat/finetune.py --resume  # continue from the last check-
                                                # point

To evaluate the result on the test split:
    python3 languages/tat/eval.py --models languages/tat/whisper-large-v3-tt/final

Requires the datacollective fork [1] with the TSV quoting fix [2], the script
refuses to run if load_dataset returns fewer rows than the TSV files contain.

[1] https://github.com/IlnarSelimcan/datacollective-python/
[2] https://github.com/IlnarSelimcan/datacollective-python/commit/3e9312ee2d9df9d2f4438667b66cd6b0bed0a854
"""

import os
import random
from collections import defaultdict
from collections.abc import Callable, Sequence
from dataclasses import dataclass
from itertools import combinations
from pathlib import Path
from typing import Annotated, TypedDict

import librosa
import numpy as np
import numpy.typing as npt
import pandas as pd
import torch
import typer
from datacollective import load_dataset
from transformers import (
    EarlyStoppingCallback,
    EvalPrediction,
    Seq2SeqTrainer,
    Seq2SeqTrainingArguments,
    WhisperForConditionalGeneration,
    WhisperProcessor,
    WhisperTokenizer,
    set_seed,
)

from eval import (
    COMMON_VOICE_SCRIPTED_SPEECH_27_0_TATAR,
    SAMPLING_RATE,
    WHISPER_GENERATE_KWARGS,
    Audio,
    load_audio,
    score,
)

HERE = Path(__file__).resolve().parent

MODEL_ID = "openai/whisper-large-v3"
OUTPUT_DIR = HERE / "whisper-large-v3-tt"

LABEL_PAD = -100
"""Label value ignored by the loss (PyTorch's cross-entropy default)."""

SPEED_RATES = (0.9, 1.0, 1.1)
"""Speed perturbation factors, used only with --speed-perturb."""


# -- Types --------------------------------------------------------------------


class Example(TypedDict):
    """One training example: log-mel features and label token ids."""

    input_features: npt.NDArray[np.float32]
    labels: list[int]


class Batch(TypedDict):
    """A padded batch, as the model's forward pass expects it."""

    input_features: torch.Tensor
    labels: torch.Tensor


@dataclass(frozen=True)
class Splits:
    """Train, dev and test clips, one row per clip."""

    train: pd.DataFrame
    dev: pd.DataFrame
    test: pd.DataFrame

    def items(self) -> list[tuple[str, pd.DataFrame]]:
        """Return (name, frame) pairs in a fixed order."""
        return [("train", self.train), ("dev", self.dev), ("test", self.test)]


@dataclass(frozen=True)
class Config:
    """Everything that defines one training run."""

    model_id: str
    output_dir: Path
    learning_rate: float
    encoder_lr_ratio: float
    batch_size: int
    grad_accum: int
    epochs: int
    warmup_steps: int
    weight_decay: float
    eval_steps: int
    patience: int
    dev_clips: int
    max_dev_speaker_clips: int
    max_clips_per_speaker: int
    resplit: bool
    test_speakers: int
    dev_speakers: int
    max_held_out_speaker_clips: int
    speed_perturb: bool
    optim_8bit: bool
    workers: int
    seed: int
    resume: bool
    dry_run: bool


# -- Data ---------------------------------------------------------------------


def load_common_voice() -> pd.DataFrame:
    """
    Return every CV Scripted Speech 27.0 Tatar row, after checking that none
    were dropped.
    """
    if not os.getenv("MDC_API_KEY"):
        raise RuntimeError(
            "Set the MDC_API_KEY environment variable to download the dataset."
        )
    dataset = load_dataset(COMMON_VOICE_SCRIPTED_SPEECH_27_0_TATAR)
    check_complete(dataset)
    return dataset


def tsv_row_count(path: Path) -> int:
    """
    Return the number of data rows in the TSV file at `path`
    (assuming it has a header line).
    """
    with path.open(encoding="utf-8") as f:
        return sum(1 for _ in f) - 1


def check_complete(dataset: pd.DataFrame) -> None:
    """
    Raise if `dataset` has fewer rows per split than its TSV file, which is
    what an unpatched datacollective returns.
    """
    cv_dir = Path(dataset.audio_path.iloc[0]).parent.parent
    for split, loaded in dataset.split.value_counts().items():
        expected = tsv_row_count(cv_dir / f"{split}.tsv")
        if loaded != expected:
            raise RuntimeError(
                f"load_dataset returned {loaded} '{split}' rows, but "
                f"{split}.tsv has {expected}. Install the patched datacollective."
            )


def exclude(frame: pd.DataFrame, held_out: pd.DataFrame) -> pd.DataFrame:
    """
    Return the rows of `frame` sharing no speaker and no sentence with
    `held_out`.
    """
    shared = frame.speaker_id.isin(
        held_out.speaker_id
    ) | frame.sentence_id.isin(held_out.sentence_id)
    return frame[~shared]


def pick_dev_speakers(
    pool: pd.DataFrame, target_clips: int, max_clips: int, seed: int
) -> set[str]:
    """
    Return random "light" speakers from `pool` until together they have
    `target_clips` clips.

    Contributions are very uneven: a couple of speakers recorded thousands of
    clips each, most others a few dozen. Only speakers with at most `max_clips`
    clips qualify ("light"), because every dev speaker's clips are removed from
    train to keep the two splits speaker-disjoint. A heavy speaker in dev would
    take thousands of clips out of an already small training set; a light one
    costs only a few dozen.
    """
    counts = pool.speaker_id.value_counts()
    light = counts[counts <= max_clips].sample(frac=1, random_state=seed)
    clips_before = light.cumsum().shift(fill_value=0)
    return set(light[clips_before < target_clips].index)


def cap_per_speaker(frame: pd.DataFrame, cap: int, seed: int) -> pd.DataFrame:
    """Return at most `cap` random clips per speaker from `frame`."""
    shuffled = frame.sample(frac=1, random_state=seed)
    return shuffled.groupby("speaker_id").head(cap)


def check_disjoint(splits: Splits) -> None:
    """Raise if any two splits share a speaker or a sentence."""
    for (name_a, a), (name_b, b) in combinations(splits.items(), 2):
        for column in ("speaker_id", "sentence_id"):
            shared = set(a[column]) & set(b[column])
            if shared:
                raise ValueError(
                    f"{name_a} and {name_b} share {len(shared)} values of {column}"
                )


def make_splits(
    dataset: pd.DataFrame,
    dev_clips: int,
    max_dev_speaker_clips: int,
    max_clips_per_speaker: int,
    seed: int,
) -> Splits:
    """
    Return train/dev/test splits: the official test split, a small dev set of
    light speakers, and every other validated clip as train.
    `max_clips_per_speaker` of 0 means no cap.
    """
    test = dataset[dataset.split == "test"]
    pool = exclude(dataset[dataset.split == "validated"], test)

    speakers = pick_dev_speakers(pool, dev_clips, max_dev_speaker_clips, seed)
    if not speakers:
        raise ValueError(
            f"No speaker outside test has at most {max_dev_speaker_clips} clips, "
            "so dev would be empty; raise --max-dev-speaker-clips."
        )
    dev = pool[pool.speaker_id.isin(speakers)].drop_duplicates("sentence_id")

    train = exclude(pool, dev)
    if max_clips_per_speaker:
        train = cap_per_speaker(train, max_clips_per_speaker, seed)

    splits = Splits(
        train=train.sample(frac=1, random_state=seed).reset_index(drop=True),
        dev=dev.reset_index(drop=True),
        test=test.reset_index(drop=True),
    )
    check_disjoint(splits)
    return splits


def sample_speakers(
    frame: pd.DataFrame, n: int, cap: int, seed: int
) -> pd.DataFrame:
    """
    Return up to `cap` clips from each of `n` random speakers in `frame`,
    one clip per sentence.
    """
    speakers = frame.speaker_id.drop_duplicates()
    chosen = speakers.sample(n=min(n, len(speakers)), random_state=seed)
    clips = frame[frame.speaker_id.isin(chosen)].drop_duplicates("sentence_id")
    return cap_per_speaker(clips, cap, seed)


def make_resplits(
    dataset: pd.DataFrame,
    test_speakers: int,
    dev_speakers: int,
    max_held_out_speaker_clips: int,
    max_clips_per_speaker: int,
    seed: int,
) -> Splits:
    """
    Return train/dev/test splits drawn afresh: test and dev from speakers of
    the official test split, train from every other validated clip.
    `max_clips_per_speaker` of 0 means no cap on train.
    """
    official_test = dataset[dataset.split == "test"]
    test = sample_speakers(
        official_test, test_speakers, max_held_out_speaker_clips, seed
    )
    dev = sample_speakers(
        exclude(official_test, test),
        dev_speakers,
        max_held_out_speaker_clips,
        seed,
    )

    train = exclude(exclude(dataset[dataset.split == "validated"], test), dev)
    if max_clips_per_speaker:
        train = cap_per_speaker(train, max_clips_per_speaker, seed)

    splits = Splits(
        train=train.sample(frac=1, random_state=seed).reset_index(drop=True),
        dev=dev.reset_index(drop=True),
        test=test.reset_index(drop=True),
    )
    check_disjoint(splits)
    return splits


def build_splits(dataset: pd.DataFrame, cfg: Config) -> Splits:
    """Return the splits `cfg` asks for: official (default) or re-split."""
    if cfg.resplit:
        return make_resplits(
            dataset,
            cfg.test_speakers,
            cfg.dev_speakers,
            cfg.max_held_out_speaker_clips,
            cfg.max_clips_per_speaker,
            cfg.seed,
        )
    return make_splits(
        dataset,
        cfg.dev_clips,
        cfg.max_dev_speaker_clips,
        cfg.max_clips_per_speaker,
        cfg.seed,
    )


def describe(splits: Splits) -> pd.DataFrame:
    """Return clip, speaker and sentence counts per split."""
    return pd.DataFrame.from_dict(
        {
            name: {
                "clips": len(frame),
                "speakers": frame.speaker_id.nunique(),
                "sentences": frame.sentence_id.nunique(),
                "top_speaker_share": round(
                    frame.speaker_id.value_counts(normalize=True).iloc[0], 3
                ),
            }
            for name, frame in splits.items()
        },
        orient="index",
    )


def save_splits(splits: Splits, directory: Path) -> None:
    """
    Write each split's clips to `directory`/<split>.tsv, for
    reproducibility.
    """
    directory.mkdir(parents=True, exist_ok=True)
    columns = ["audio_path", "transcription", "speaker_id", "sentence_id"]
    for name, frame in splits.items():
        frame[columns].to_csv(directory / f"{name}.tsv", sep="\t", index=False)


# -- Features -----------------------------------------------------------------


def change_speed(audio: Audio, rate: float) -> Audio:
    """
    Return `audio` played `rate` times faster, pitch included, as in Kaldi-style
    speed perturbation.
    """
    if rate == 1.0:
        return audio
    return librosa.resample(
        audio, orig_sr=round(SAMPLING_RATE * rate), target_sr=SAMPLING_RATE
    )


class WhisperDataset:
    """Map-style dataset that decodes audio and tokenizes text on access."""

    def __init__(
        self,
        frame: pd.DataFrame,
        processor: WhisperProcessor,
        speed_rates: Sequence[float] = (),
    ) -> None:
        self.paths = frame.audio_path.tolist()
        self.texts = frame.transcription.tolist()
        self.processor = processor
        self.speed_rates = speed_rates

    def __len__(self) -> int:
        return len(self.paths)

    def __getitem__(self, i: int) -> Example:
        audio = load_audio(self.paths[i])
        if self.speed_rates:
            # `random` is reseeded per DataLoader worker, unlike numpy's global
            # RNG
            audio = change_speed(audio, random.choice(self.speed_rates))
        features = self.processor.feature_extractor(
            audio, sampling_rate=SAMPLING_RATE
        ).input_features[0]
        labels = self.processor.tokenizer(self.texts[i]).input_ids
        return {"input_features": features, "labels": labels}


@dataclass(frozen=True)
class Collator:
    """
    Pads examples into a batch; padded label positions are ignored by the
    loss.
    """

    processor: WhisperProcessor
    decoder_start_token_id: int

    def __call__(self, examples: list[Example]) -> Batch:
        features = self.processor.feature_extractor.pad(
            [{"input_features": e["input_features"]} for e in examples],
            return_tensors="pt",
        )
        labels = self.processor.tokenizer.pad(
            [{"input_ids": e["labels"]} for e in examples], return_tensors="pt"
        )
        ids = labels["input_ids"].masked_fill(
            labels.attention_mask.ne(1), LABEL_PAD
        )
        # the model prepends the start token itself during training
        if (ids[:, 0] == self.decoder_start_token_id).all():
            ids = ids[:, 1:]
        return {"input_features": features["input_features"], "labels": ids}


# -- Model --------------------------------------------------------------------


def load_model(
    model_id: str,
) -> tuple[WhisperForConditionalGeneration, WhisperProcessor]:
    """Return the model and processor, set up to transcribe Tatar."""
    processor = WhisperProcessor.from_pretrained(
        model_id, language="tatar", task="transcribe"
    )
    model = WhisperForConditionalGeneration.from_pretrained(
        model_id,
        dtype=torch.float32,  # transformers 5 defaults to float16
    )
    model.config.use_cache = False  # incompatible with gradient checkpointing
    configure_generation(model)
    return model, processor


def configure_generation(model: WhisperForConditionalGeneration) -> None:
    """
    Decode Tatar with the same greedy settings as eval.py, both for dev
    evaluation and in the saved model.
    """
    for key, value in WHISPER_GENERATE_KWARGS.items():
        setattr(model.generation_config, key, value)
    model.generation_config.forced_decoder_ids = None
    model.config.forced_decoder_ids = None


def parameter_groups(
    model: torch.nn.Module,
    lr: float,
    encoder_lr_ratio: float,
    weight_decay: float,
) -> list[dict[str, object]]:
    """
    Return AdamW parameter groups: the encoder gets `encoder_lr_ratio` times
    the decoder's learning rate; biases and norms get no weight decay.
    """
    grouped: dict[tuple[bool, bool], list[torch.nn.Parameter]] = defaultdict(
        list
    )
    for name, param in model.named_parameters():
        if param.requires_grad:
            is_encoder = name.startswith("model.encoder.")
            decays = param.ndim >= 2
            grouped[(is_encoder, decays)].append(param)
    return [
        {
            "params": params,
            "lr": lr * encoder_lr_ratio if is_encoder else lr,
            "weight_decay": weight_decay if decays else 0.0,
        }
        for (is_encoder, decays), params in grouped.items()
    ]


def build_optimizer(
    model: torch.nn.Module, cfg: Config
) -> torch.optim.Optimizer:
    """Return AdamW, or its 8-bit version, over the model's parameter groups."""
    groups = parameter_groups(
        model, cfg.learning_rate, cfg.encoder_lr_ratio, cfg.weight_decay
    )
    if cfg.optim_8bit:
        import bitsandbytes as bnb

        return bnb.optim.AdamW8bit(groups)
    return torch.optim.AdamW(groups)


# -- Training ------------------------------------------------------------------


def make_compute_metrics(
    tokenizer: WhisperTokenizer,
) -> Callable[[EvalPrediction], dict[str, float]]:
    """Return a function scoring generated predictions with eval.score."""

    def decode(ids: npt.NDArray[np.int64]) -> list[str]:
        ids = np.where(ids == LABEL_PAD, tokenizer.pad_token_id, ids)
        return tokenizer.batch_decode(ids, skip_special_tokens=True)

    def compute_metrics(pred: EvalPrediction) -> dict[str, float]:
        texts = pd.DataFrame(
            {
                "reference": decode(pred.label_ids),
                "hypothesis": decode(pred.predictions),
            }
        )
        return score(texts)

    return compute_metrics


def use_bf16() -> bool:
    """
    Return whether the GPU supports bfloat16, which is safer than fp16 for
    training.
    """
    return torch.cuda.is_available() and torch.cuda.is_bf16_supported()


def training_args(cfg: Config) -> Seq2SeqTrainingArguments:
    """Return the Trainer settings for `cfg`."""
    return Seq2SeqTrainingArguments(
        output_dir=str(cfg.output_dir),
        per_device_train_batch_size=cfg.batch_size,
        per_device_eval_batch_size=cfg.batch_size * 4,
        gradient_accumulation_steps=cfg.grad_accum,
        learning_rate=cfg.learning_rate,  # scheduler scales each group's own rate
        lr_scheduler_type="cosine",
        warmup_steps=cfg.warmup_steps,
        num_train_epochs=cfg.epochs,
        bf16=use_bf16(),
        gradient_checkpointing=True,
        gradient_checkpointing_kwargs={"use_reentrant": False},
        eval_strategy="steps",
        eval_steps=cfg.eval_steps,
        save_strategy="steps",
        save_steps=cfg.eval_steps,
        save_total_limit=2,
        load_best_model_at_end=True,
        metric_for_best_model="cer",
        greater_is_better=False,
        predict_with_generate=True,
        generation_max_length=225,
        logging_steps=25,
        report_to=["tensorboard"],
        dataloader_num_workers=cfg.workers,
        remove_unused_columns=False,
        label_names=["labels"],
        seed=cfg.seed,
    )


def has_checkpoint(directory: Path) -> bool:
    """Return whether `directory` contains a Trainer checkpoint."""
    return directory.is_dir() and any(directory.glob("checkpoint-*"))


def build_trainer(
    cfg: Config, splits: Splits
) -> tuple[Seq2SeqTrainer, WhisperProcessor]:
    """Return a Trainer for `cfg` on `splits`, and the processor it uses."""
    model, processor = load_model(cfg.model_id)
    speed_rates = SPEED_RATES if cfg.speed_perturb else ()
    trainer = Seq2SeqTrainer(
        model=model,
        args=training_args(cfg),
        train_dataset=WhisperDataset(splits.train, processor, speed_rates),
        eval_dataset=WhisperDataset(splits.dev, processor),
        data_collator=Collator(processor, model.config.decoder_start_token_id),
        compute_metrics=make_compute_metrics(processor.tokenizer),
        processing_class=processor,
        optimizers=(build_optimizer(model, cfg), None),
        callbacks=[
            EarlyStoppingCallback(early_stopping_patience=cfg.patience)
        ],
    )
    return trainer, processor


def run(cfg: Config) -> Path | None:
    """Train according to `cfg`; return the final model's directory, or None on a dry run."""
    set_seed(cfg.seed)
    splits = build_splits(load_common_voice(), cfg)
    print(describe(splits).to_string())
    if cfg.dry_run:
        return None

    save_splits(splits, cfg.output_dir / "splits")
    trainer, processor = build_trainer(cfg, splits)
    trainer.train(
        resume_from_checkpoint=cfg.resume and has_checkpoint(cfg.output_dir)
    )

    final = cfg.output_dir / "final"
    trainer.save_model(
        str(final)
    )  # the best checkpoint, thanks to load_best_model_at_end
    processor.save_pretrained(str(final))
    print(f"Best model saved to {final}")
    return final


# -- Main ---------------------------------------------------------------------


def main(
    model_id: Annotated[
        str, typer.Option(help="Model to fine-tune.")
    ] = MODEL_ID,
    output_dir: Annotated[
        Path, typer.Option(help="Checkpoints, splits and final model.")
    ] = OUTPUT_DIR,
    learning_rate: Annotated[
        float,
        typer.Option(
            help="Decoder learning rate (BuzzASR: 3.24e-6, 6.48e-6, 2e-5)."
        ),
    ] = 6.48e-6,
    encoder_lr_ratio: Annotated[
        float,
        typer.Option(
            help="Encoder learning rate as a fraction of the decoder's."
        ),
    ] = 0.3,
    batch_size: Annotated[int, typer.Option(help="Clips per GPU step.")] = 4,
    grad_accum: Annotated[
        int,
        typer.Option(
            help="Steps per optimizer update; effective batch = batch size x this."
        ),
    ] = 8,
    epochs: Annotated[int, typer.Option(help="Maximum epochs.")] = 6,
    warmup_steps: Annotated[
        int, typer.Option(help="Learning-rate warmup steps.")
    ] = 150,
    weight_decay: Annotated[
        float, typer.Option(help="AdamW weight decay.")
    ] = 4e-5,
    eval_steps: Annotated[
        int, typer.Option(help="Evaluate and save every N updates.")
    ] = 100,
    patience: Annotated[
        int,
        typer.Option(
            help="Stop after N evaluations without dev CER improvement."
        ),
    ] = 3,
    dev_clips: Annotated[
        int, typer.Option(help="Approximate dev set size, in clips.")
    ] = 500,
    max_dev_speaker_clips: Annotated[
        int,
        typer.Option(
            help="Only speakers with at most this many clips go to dev."
        ),
    ] = 200,
    max_clips_per_speaker: Annotated[
        int, typer.Option(help="Cap on train clips per speaker (0 = no cap).")
    ] = 0,
    resplit: Annotated[
        bool,
        typer.Option(help="Draw new test/dev/train splits (see docstring)."),
    ] = False,
    test_speakers: Annotated[
        int, typer.Option(help="With --resplit: speakers in the test set.")
    ] = 80,
    dev_speakers: Annotated[
        int, typer.Option(help="With --resplit: speakers in the dev set.")
    ] = 25,
    max_held_out_speaker_clips: Annotated[
        int,
        typer.Option(
            help="With --resplit: clips per test/dev speaker, at most."
        ),
    ] = 40,
    speed_perturb: Annotated[
        bool,
        typer.Option(help="Randomly change speed by 0.9x/1.1x in training."),
    ] = False,
    optim_8bit: Annotated[
        bool,
        typer.Option(
            help="Use 8-bit AdamW (bitsandbytes) to save GPU memory."
        ),
    ] = False,
    workers: Annotated[
        int, typer.Option(help="DataLoader worker processes.")
    ] = 4,
    seed: Annotated[
        int, typer.Option(help="Random seed for splits and training.")
    ] = 42,
    resume: Annotated[
        bool, typer.Option(help="Resume from the last checkpoint.")
    ] = False,
    dry_run: Annotated[
        bool, typer.Option(help="Only build and describe the splits.")
    ] = False,
) -> Path | None:
    """Fine-tune Whisper on Common Voice 27.0 Tatar."""
    return run(
        Config(**locals())
    )  # the parameters are exactly Config's fields


if __name__ == "__main__":
    typer.run(main)

Appendix C: eval.py

The evaluation script used for all models in Table 2. Given --test-tsv, it evaluates on a saved test set, such as run 2's, instead of the official test split.

Download: eval.py (418 lines)

Show the code
"""
Evaluate existing Tatar ASR models on the test split of Common Voice
Scripted Speech 27.0 Tatar.

Usage:
    python3 languages/tat/eval.py             # all models, full test split
    python3 languages/tat/eval.py --limit 50  # quick smoke test
    python3 languages/tat/eval.py --models yasalma/whisper-finetuned-tt-asr
    python3 languages/tat/eval.py --models omniASR_LLM_1B_v2
    python3 languages/tat/eval.py --test-tsv <finetune output>/splits/test.tsv

Everything expensive is cached on disk, so reruns are cheap:
    - the test split         -> cv27_tat_test.parquet
    - each model's output    -> results/<model>[__n<limit>].tsv

With --test-tsv, results go to a results/ folder next to that file instead,
so they never mix with results on the official split.

Delete a model's TSV to re-run that model.

Model ids starting with "omniASR_" are Meta's Omnilingual ASR models and need
`pip install omnilingual-asr` (imported lazily, so the rest works without it).
Everything else is loaded with the Hugging Face transformers pipeline.

Set OMP_NUM_THREADS=1 on cloud machines: NumPy and PyTorch otherwise start a
thread per host core, far more than the cores a pod actually gets, and the
oversubscription slows feature extraction to a crawl.
"""

import os
import re
import unicodedata
from pathlib import Path
from typing import Annotated, TypedDict

import jiwer
import librosa
import numpy as np
import numpy.typing as npt
import pandas as pd
import torch
import typer
from datacollective import download_dataset, load_dataset
from torch.utils.data import Dataset
from tqdm import tqdm
from transformers import pipeline

COMMON_VOICE_SCRIPTED_SPEECH_27_0_TATAR = "cmu62fysg00noo1077ljif8fv"

HERE = Path(__file__).resolve().parent
TEST_CACHE = HERE / "cv27_tat_test.parquet"
RESULTS_DIR = HERE / "results"

SAMPLING_RATE = 16_000

OMNI_LANG = "tat_Cyrl"

WORKERS = 1
"""
DataLoader processes preparing audio while the GPU transcribes. The ASR
pipeline is a ChunkPipeline, which transformers limits to one worker.
"""

MODELS = [
    # Baselines (see the BuzzASR paper, arXiv:2609.09554)
    "openai/whisper-large-v3",  # zero-shot reference for our own fine-tune
    "omniASR_CTC_1B_v2",
    "omniASR_LLM_1B_v2",
    "omniASR_LLM_7B_v2",  # needs ~20 GB GPU memory
    # Tatar fine-tunes on Hugging Face
    "AigizK/wav2vec2-large-mms-1b-tatar-v2",
    "yasalma/whisper-finetuned-tt-asr",
    "AigizK/wav2vec2-large-mms-1b-tatar",
    "yasalma/whisper-finetuned-tatartts-asr",
    "anton-l/wav2vec2-large-xlsr-53-tatar",
    "crang/wav2vec2-large-xlsr-53-tatar",
    "emre/wav2vec2-large-xlsr-53-W2V2-TATAR-SMALL",
    "infinitejoy/wav2vec2-large-xls-r-300m-tatar",
    "kingabzpro/wav2vec2-large-xls-r-300m-Tatar",
    "sammy786/wav2vec2-xlsr-tatar",
    # Not tested here
    # https://github.com/IS2AI/Soyle
    # https://github.com/IS2AI/TurkicASR
]

# Same greedy settings as the BuzzASR paper, so Whisper numbers are comparable.
WHISPER_GENERATE_KWARGS = {
    "language": "tatar",
    "task": "transcribe",
    "num_beams": 1,
    "no_repeat_ngram_size": 3,
    "repetition_penalty": 1.2,
}


# -- Types --------------------------------------------------------------------


Audio = npt.NDArray[np.float32]
"""Mono waveform at SAMPLING_RATE."""


class HFAudioInput(TypedDict):
    """One input item for the Hugging Face ASR pipeline."""

    raw: Audio
    sampling_rate: int


class OmniAudioInput(TypedDict):
    """One input item for Omnilingual's ASRInferencePipeline."""

    waveform: Audio
    sample_rate: int


Row = dict[str, str | float]
"""One line of the summary table: model id plus scores, or an error message."""


# -- Data ---------------------------------------------------------------------


def load_test_split() -> pd.DataFrame:
    """
    Return the CV 27.0 Tatar test split, downloading and caching it on first
    use.
    """
    if TEST_CACHE.exists():
        return pd.read_parquet(TEST_CACHE)
    if not os.getenv("MDC_API_KEY"):
        raise RuntimeError(
            "Set the MDC_API_KEY environment variable to download the dataset."
        )
    download_dataset(COMMON_VOICE_SCRIPTED_SPEECH_27_0_TATAR)
    dataset = load_dataset(COMMON_VOICE_SCRIPTED_SPEECH_27_0_TATAR)
    test = dataset[dataset.split == "test"].reset_index(drop=True)
    test.to_parquet(TEST_CACHE)
    return test


def load_audio(path: str) -> Audio:
    """
    Decode the audio file at `path` to a mono float array at SAMPLING_RATE.
    """
    audio, _ = librosa.load(path, sr=SAMPLING_RATE, mono=True)
    return audio


class AudioDataset(Dataset):
    """
    Clips decoded on access, in the format the HF pipeline expects. Being a
    Dataset (not a generator) lets the pipeline prepare upcoming clips in a
    worker process while the GPU transcribes the current batch.
    """

    def __init__(self, paths: list[str]) -> None:
        self.paths = paths

    def __len__(self) -> int:
        return len(self.paths)

    def __getitem__(self, i: int) -> HFAudioInput:
        return {
            "raw": load_audio(self.paths[i]),
            "sampling_rate": SAMPLING_RATE,
        }


# -- Text normalization -------------------------------------------------------


def normalize(text: str) -> str:
    """
    Lowercase, drop punctuation, collapse whitespace.
    \\w is Unicode-aware, so ә ө ү җ ң һ are kept.
    """
    text = unicodedata.normalize("NFC", text).lower()
    text = re.sub(r"[^\w\s]|_", " ", text)
    return " ".join(text.split())


# -- Inference ---------------------------------------------------------------


def free_gpu() -> None:
    """Release cached GPU memory so the next model has room to load."""
    if torch.cuda.is_available():
        torch.cuda.empty_cache()


def transcribe_hf(
    model_id: str, paths: list[str], batch_size: int
) -> list[str]:
    """Transcribe `paths` with a Hugging Face model, one hypothesis per path."""
    use_cuda = torch.cuda.is_available()
    asr = pipeline(
        "automatic-speech-recognition",
        model=model_id,
        device="cuda:0" if use_cuda else "cpu",
        dtype=torch.float16 if use_cuda else torch.float32,
    )
    inputs = AudioDataset(paths)
    if "whisper" in model_id.lower():
        outputs = asr(
            inputs,
            batch_size=batch_size,
            num_workers=WORKERS,
            generate_kwargs=WHISPER_GENERATE_KWARGS,
        )
    else:
        outputs = asr(inputs, batch_size=batch_size, num_workers=WORKERS)
    hyps = [
        out["text"] for out in tqdm(outputs, total=len(paths), desc=model_id)
    ]
    del asr
    free_gpu()
    return hyps


def transcribe_omnilingual(
    model_card: str, paths: list[str], batch_size: int
) -> list[str]:
    """
    Transcribe `paths` with a Meta Omnilingual ASR model, one hypothesis per
    path.
    """
    from omnilingual_asr.models.inference.pipeline import ASRInferencePipeline

    try:
        from omnilingual_asr.models.wav2vec2_llama.lang_ids import (
            supported_langs,
        )

        assert OMNI_LANG in supported_langs, (
            f"{OMNI_LANG} not supported by Omnilingual"
        )
    except ImportError:
        pass

    asr = ASRInferencePipeline(model_card=model_card)
    use_lang = (
        "_LLM_" in model_card
    )  # CTC models have no language conditioning?

    hyps: list[str] = []
    for i in tqdm(range(0, len(paths), batch_size), desc=model_card):
        chunk = paths[i : i + batch_size]
        audio: list[OmniAudioInput] = [
            {"waveform": load_audio(p), "sample_rate": SAMPLING_RATE}
            for p in chunk
        ]
        kwargs = {"lang": [OMNI_LANG] * len(chunk)} if use_lang else {}
        hyps.extend(asr.transcribe(audio, batch_size=len(chunk), **kwargs))

    del asr
    free_gpu()
    return hyps


def transcribe(model_id: str, paths: list[str], batch_size: int) -> list[str]:
    """Transcribe `paths` with `model_id`, dispatching to the right backend."""
    if model_id.startswith("omniASR"):
        return transcribe_omnilingual(model_id, paths, batch_size)
    return transcribe_hf(model_id, paths, batch_size)


def load_test_tsv(path: Path) -> pd.DataFrame:
    """Return a test set saved by finetune.py (its splits/test.tsv)."""
    return pd.read_csv(path, sep="\t", keep_default_na=False)


def results_dir_for(test_tsv: Path | None) -> Path:
    """
    Return where results go: RESULTS_DIR for the official split, else a
    results/ folder next to `test_tsv`.
    """
    return RESULTS_DIR if test_tsv is None else test_tsv.parent / "results"


def predictions_path(
    model_id: str, limit: int, results_dir: Path | None = None
) -> Path:
    """
    Return the cache file for `model_id`'s predictions on the first `limit`
    clips, in `results_dir` (default RESULTS_DIR).
    """
    name = model_id.replace("/", "__")
    suffix = f"__n{limit}" if limit else ""
    return (results_dir or RESULTS_DIR) / f"{name}{suffix}.tsv"


def get_predictions(
    model_id: str,
    test: pd.DataFrame,
    limit: int,
    batch_size: int,
    results_dir: Path | None = None,
) -> pd.DataFrame:
    """
    Return reference/hypothesis pairs for `model_id` on `test`, from cache if
    present.
    """
    path = predictions_path(model_id, limit, results_dir)
    if path.exists():
        return pd.read_csv(path, sep="\t", keep_default_na=False)
    hyps = transcribe(model_id, list(test.audio_path), batch_size)
    df = pd.DataFrame(
        {
            "audio_path": test.audio_path,
            "reference": test.transcription,
            "hypothesis": hyps,
        }
    )
    path.parent.mkdir(parents=True, exist_ok=True)
    df.to_csv(path, sep="\t", index=False)
    return df


# -- Scoring ------------------------------------------------------------------


def score(df: pd.DataFrame) -> dict[str, float]:
    """
    Compute WER and CER of `df.hypothesis` against `df.reference`, normalized
    and raw.
    """
    refs, hyps = list(df.reference), list(df.hypothesis)
    refs_n, hyps_n = [normalize(t) for t in refs], [normalize(t) for t in hyps]
    return {
        "wer": jiwer.wer(refs_n, hyps_n),
        "cer": jiwer.cer(refs_n, hyps_n),
        "wer_raw": jiwer.wer(refs, hyps),
        "cer_raw": jiwer.cer(refs, hyps),
    }


def evaluate(
    model_id: str,
    test: pd.DataFrame,
    limit: int,
    batch_size: int,
    results_dir: Path | None = None,
) -> tuple[pd.DataFrame | None, Row]:
    """
    One model -> (predictions, summary row). Errors become a row, not a crash.
    """
    try:
        preds = get_predictions(model_id, test, limit, batch_size, results_dir)
        return preds, {"model": model_id, **score(preds)}
    except Exception as e:
        print(f"!! {model_id} failed: {type(e).__name__}: {e}")
        return None, {"model": model_id, "error": f"{type(e).__name__}: {e}"}


def summarize(
    rows: list[Row], limit: int, results_dir: Path | None = None
) -> pd.DataFrame:
    """
    Combine per-model rows into a table sorted by CER, and save it as TSV in
    `results_dir` (default RESULTS_DIR).
    """
    summary = pd.DataFrame(rows)
    if "cer" in summary:
        summary = summary.sort_values("cer", na_position="last")
    directory = results_dir or RESULTS_DIR
    directory.mkdir(parents=True, exist_ok=True)
    name = f"summary__n{limit}.tsv" if limit else "summary.tsv"
    summary.to_csv(directory / name, sep="\t", index=False)
    return summary


# ── Main ──────────────────────────────────────────────────────────────────


def main(
    models: Annotated[
        list[str],
        typer.Option(help="Model ids to evaluate (repeat the flag)."),
    ] = MODELS,
    limit: Annotated[
        int,
        typer.Option(
            help="Evaluate on the first N test clips only (0 = all)."
        ),
    ] = 0,
    batch_size: Annotated[
        int, typer.Option(help="Inference batch size.")
    ] = 16,
    test_tsv: Annotated[
        Path | None,
        typer.Option(help="Test set saved by finetune.py (splits/test.tsv)."),
    ] = None,
) -> tuple[pd.DataFrame, dict[str, pd.DataFrame], pd.DataFrame]:
    """Evaluate Tatar ASR models on the CV 27.0 Tatar test split."""
    test = load_test_split() if test_tsv is None else load_test_tsv(test_tsv)
    results_dir = results_dir_for(test_tsv)
    if limit:
        test = test.head(limit)
    print(f"Test clips: {len(test)}, speakers: {test.speaker_id.nunique()}")

    results = {
        m: evaluate(m, test, limit, batch_size, results_dir) for m in models
    }
    preds = {m: p for m, (p, _) in results.items() if p is not None}
    summary = summarize(
        [row for _, row in results.values()], limit, results_dir
    )

    with pd.option_context(
        "display.max_colwidth", 60, "display.float_format", "{:.3f}".format
    ):
        print(summary.to_string(index=False))
    return test, preds, summary


if __name__ == "__main__":
    typer.run(main)

Appendix D: software environment

Both runs used a single NVIDIA A100 80GB PCIe on RunPod, with the runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404 image. I did not run pip freeze before deleting the pod, so the package list below is reconstructed. The versions I know for certain are these:

Known software and hardware versions on the training pod.
Component Version Source
GPU driver 580.159.04 (run 1), 595.91.07 (run 2) nvidia-smi
Python 3.12.3 the pod
PyTorch 2.8.0+cu128 the pod image
Transformers 5.18.0 both runs' TensorBoard logs
NumPy, SciPy, librosa 2.1.2, 1.18.1, 1.0.0 the pod
datacollective my fork at commit 3e9312ee, with a fix for TSV quoting; release 0.6.3 includes the same fix the pod

Everything else was resolved with uv pip compile for Python 3.12 on Linux, with the known versions pinned and every package released after the morning of run 1 (October 1, 2026) excluded. The header of the file explains how to regenerate it.

Download: requirements-reconstructed.txt (285 lines)

Show the code
# Reconstructed after the fact: the RunPod pod was deleted before anyone ran
# `pip freeze`. Versions marked "known" below were read from the pod or from the
# training logs; everything else is what pip would most likely have resolved on
# the morning of run 1 (October 1, 2026). Regenerate with:
#
#   uv pip compile requirements.in --python-version 3.12 \
#     --python-platform x86_64-manylinux_2_28 \
#     --exclude-newer 2026-10-01T06:00:00Z -o requirements-reconstructed.txt
#
# where requirements.in pins the known versions:
#   torch==2.8.0           (known: preinstalled in the RunPod image, +cu128 build)
#   transformers==5.18.0   (known: recorded in both runs' TensorBoard logs)
#   numpy==2.1.2           (known)
#   scipy==1.18.1          (known)
#   librosa==1.0.0         (known)
# and lists the rest unpinned: accelerate, tensorboard, jiwer, pandas, pyarrow,
# tqdm, typer, datacollective.
#
# Installed later, only to evaluate AigizK/wav2vec2-large-mms-1b-tatar with its
# language model (known versions, not needed for training):
#   kenlm==0.3.0
#   pyctcdecode==0.5.0
#
absl-py==2.5.0
    # via tensorboard
accelerate==1.15.0
    # via -r requirements.in
annotated-doc==0.0.5
    # via typer
annotated-types==0.8.0
    # via pydantic
anyio==4.15.1
    # via httpx
certifi==2026.7.22
    # via
    #   httpcore
    #   httpx
    #   requests
cffi==2.1.1
    # via soundfile
charset-normalizer==3.5.2
    # via requests
click==8.5.0
    # via
    #   huggingface-hub
    #   jiwer
cloudpickle==3.1.2
    # via joblib
datacollective==0.6.3
    # via -r requirements.in
decorator==5.3.1
    # via librosa
filelock==4.0.7
    # via
    #   huggingface-hub
    #   torch
fox-progress-bar==0.1.3
    # via datacollective
fsspec==2026.9.0
    # via
    #   huggingface-hub
    #   torch
grpcio==1.84.0
    # via tensorboard
h11==0.16.0
    # via httpcore
hf-xet==1.6.0
    # via huggingface-hub
httpcore==1.0.9
    # via httpx
httpx==0.28.1
    # via huggingface-hub
huggingface-hub==1.33.0
    # via
    #   accelerate
    #   tokenizers
    #   transformers
idna==3.20
    # via
    #   anyio
    #   httpx
    #   requests
jinja2==3.1.6
    # via torch
jiwer==4.0.0
    # via -r requirements.in
joblib==1.6.0
    # via
    #   librosa
    #   scikit-learn
lazy-loader==0.6
    # via librosa
librosa==1.0.0
    # via -r requirements.in
llvmlite==0.50.0
    # via numba
markdown==3.11
    # via tensorboard
markdown-it-py==4.2.0
    # via rich
markupsafe==3.0.3
    # via
    #   jinja2
    #   werkzeug
mdurl==0.1.2
    # via markdown-it-py
mpmath==1.3.0
    # via sympy
msgpack==1.2.3
    # via librosa
narwhals==2.26.0
    # via scikit-learn
networkx==3.7
    # via torch
numba==0.68.0
    # via librosa
numpy==2.1.2
    # via
    #   -r requirements.in
    #   accelerate
    #   librosa
    #   numba
    #   pandas
    #   scikit-learn
    #   scipy
    #   soundfile
    #   soxr
    #   tensorboard
    #   transformers
nvidia-cublas-cu12==12.8.4.1
    # via
    #   nvidia-cudnn-cu12
    #   nvidia-cusolver-cu12
    #   torch
nvidia-cuda-cupti-cu12==12.8.90
    # via torch
nvidia-cuda-nvrtc-cu12==12.8.93
    # via torch
nvidia-cuda-runtime-cu12==12.8.90
    # via torch
nvidia-cudnn-cu12==9.10.2.21
    # via torch
nvidia-cufft-cu12==11.3.3.83
    # via torch
nvidia-cufile-cu12==1.13.1.3
    # via torch
nvidia-curand-cu12==10.3.9.90
    # via torch
nvidia-cusolver-cu12==11.7.3.90
    # via torch
nvidia-cusparse-cu12==12.5.8.93
    # via
    #   nvidia-cusolver-cu12
    #   torch
nvidia-cusparselt-cu12==0.7.1
    # via torch
nvidia-nccl-cu12==2.27.3
    # via torch
nvidia-nvjitlink-cu12==12.8.93
    # via
    #   nvidia-cufft-cu12
    #   nvidia-cusolver-cu12
    #   nvidia-cusparse-cu12
    #   torch
nvidia-nvtx-cu12==12.8.90
    # via torch
packaging==26.3
    # via
    #   accelerate
    #   huggingface-hub
    #   lazy-loader
    #   pooch
    #   tensorboard
    #   transformers
pandas==3.0.6
    # via
    #   -r requirements.in
    #   datacollective
pillow==12.3.0
    # via tensorboard
platformdirs==4.12.2
    # via pooch
pooch==1.9.0
    # via librosa
protobuf==7.36.2
    # via tensorboard
psutil==7.2.2
    # via accelerate
pyarrow==25.0.1
    # via -r requirements.in
pycparser==3.0
    # via cffi
pydantic==2.13.5
    # via datacollective
pydantic-core==2.46.5
    # via pydantic
pygments==2.21.0
    # via rich
python-dateutil==2.9.0.post0
    # via pandas
python-dotenv==1.2.4
    # via datacollective
pyyaml==6.0.3
    # via
    #   accelerate
    #   datacollective
    #   huggingface-hub
    #   transformers
rapidfuzz==3.14.6
    # via jiwer
regex==2026.9.29
    # via transformers
requests==2.34.2
    # via
    #   datacollective
    #   pooch
rich==15.0.0
    # via typer
safetensors==0.8.0
    # via
    #   accelerate
    #   transformers
scikit-learn==1.9.1
    # via librosa
scipy==1.18.1
    # via
    #   -r requirements.in
    #   librosa
    #   scikit-learn
setuptools==84.0.0
    # via
    #   tensorboard
    #   torch
    #   triton
shellingham==1.5.4
    # via typer
six==1.17.0
    # via python-dateutil
soundfile==0.14.0
    # via librosa
soxr==1.1.0
    # via librosa
sympy==1.14.0
    # via torch
tensorboard==2.21.0
    # via -r requirements.in
tensorboard-data-server==0.7.2
    # via tensorboard
threadpoolctl==3.7.0
    # via scikit-learn
tokenizers==0.23.2
    # via transformers
torch==2.8.0
    # via
    #   -r requirements.in
    #   accelerate
tqdm==4.70.1
    # via
    #   -r requirements.in
    #   huggingface-hub
    #   transformers
transformers==5.18.0
    # via -r requirements.in
triton==3.4.0
    # via torch
typer==0.27.2
    # via
    #   -r requirements.in
    #   transformers
typing-extensions==4.16.0
    # via
    #   anyio
    #   grpcio
    #   huggingface-hub
    #   pydantic
    #   pydantic-core
    #   soundfile
    #   torch
    #   typing-inspection
typing-inspection==0.4.4
    # via pydantic
urllib3==2.8.0
    # via requests
werkzeug==3.1.9
    # via tensorboard

  1. Or paired not with CTC but with a Transformer decoder, as in the LLM-ASR variant of Meta's Omnilingual. ↩

Home | Hire me | Résumé | Projects | Publications | Talks | Now | Email | Reading log | Movies log

Powered by Buttondown.

Hire me