whisper-heb-v7-multi

A full fine-tune of ivrit-ai/whisper-large-v3-turbo (809M, 4 decoder layers) for Hebrew, English, Arabic, Russian and Yiddish, with language identification trained in: the model predicts the <|lang|> token itself, so you do not have to tell it which language a clip is in. Built for short telephony segments (8 kHz call-center audio) at Voicenter; Hebrew is the primary language, and the other four are what Israeli callers switch to.

Trained 2026-10-01/02 on one B200, 14,691 steps (one epoch over 3.76M clips, ~10.5k hours), warm-started from our Hebrew-only whisper-heb-v6, with token-level knowledge distillation from one teacher per language. Best checkpoint (step 14500) selected on the macro WER over five held-out dev slices decoded with auto language detection.

Usage

from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="VoicenterTeam/whisper-heb-v7-multi",
               torch_dtype="bfloat16", device="cuda")
asr("call.wav")                                              # language auto-detected
asr("call.wav", generate_kwargs={"language": "he"})          # or force one of he/en/ar/ru/yi

The shipped generation_config has language=None, so plain generate() runs Whisper's language detection with the fine-tuned weights. Forcing the language gives slightly better WER on short Arabic/Russian clips (see below) and is what to do when the language is known. The model is a stock Whisper checkpoint: it converts to CTranslate2 / faster-whisper with ct2-transformers-converter unchanged.

Results

WER on normalized text (nikud/tashkeel/punctuation stripped, case folded, ё→е, Arabic alef/ta marbuta/ya folded), greedy decoding, bf16. auto = language detected by the model. Compared with our previous Hebrew-only model (v6) and openai/whisper-large-v3-turbo. Full tables incl. 8 kHz simulated-telephony condition, forced-language mode, CER and per-set flags: results/whisper-heb-v7-multi/SUMMARY.md in the training repo.

Per language (macro over the language's test sets, clean audio, auto LID)

lang sets v7 v6 (Hebrew-only) openai turbo v7 LID acc
he ILSpeech, 10k telephony 0.128 0.134 0.240 1.00
en FLEURS+CV+AMI, AppTek call-center 0.138 1.067 0.133 0.99
ar FLEURS, CV17, Casablanca jo/ps 0.376 1.101 0.292 0.96
ru FLEURS+CV+Golos, ru.test 0.114 1.038 0.122 0.99
yi Omnilingual ydd (read) 0.347 0.983 1.196 1.00
macro 0.221 0.865 0.397 0.99

Under a simulated 8 kHz μ-law phone channel the macro is 0.239 (v6 0.867, openai 0.430).

Selected sets

set refs v7 v6 openai turbo
he telephony, 10k (1000 sample) soniox 0.155 0.166 0.342
he telephony, soniox+witness agree ≤0.2 WER (5,915) soniox 0.079 0.088 0.221
he ILSpeech speaker 1 (579) human 0.126 0.125 0.148
he whatsapp (ivrit.ai, long-form) human 0.071 0.068 0.152
he Hebrew suite macro (d1, whatsapp, fleurs, CV, kan, saspeech; forced he) human 0.102 0.094 0.196
en eval mix (FLEURS/CV/AMI) human 0.099 1.024 0.088
en AppTek call-center (3000) human 0.181 1.093 0.183
ru eval mix human 0.095 1.030 0.091
ru test (CV/SOVA/RuLS/FLEURS) human 0.133 1.047 0.153
ru real phone calls (Open STT, eval-only) human 0.388 1.154 0.316
ar eval mix human 0.376 1.101 0.292
yi Omnilingual (182) human 0.347 0.983 1.196
yi prompt-disjoint held-out (1,104) human 0.377 0.995 1.155

What this means

  • Hebrew telephony improved over the Hebrew-only v6 (10k: 0.166 → 0.155; agreement subset 0.088 → 0.079); long-form whatsapp and ILSpeech are unchanged within noise. Hebrew read speech regressed a little (FLEURS 0.180 → 0.204, Common Voice 0.134 → 0.150, forced-he).
  • English and Russian are at openai-turbo level (within ±0.01), from WER ≈ 1.0 in v6, which had lost them entirely during Hebrew fine-tuning.
  • Yiddish (Hebrew script, Hasidic orthography) works: 0.35 WER where openai turbo writes Latin-script German-ish text (1.2).
  • Arabic is the weak spot: 0.376 vs 0.292 for openai turbo. Arabic was 5% of the training mix with a Levantine-biased teacher; the test set mixes MSA and several dialects.
  • Russian telephony also lags openai turbo (0.388 vs 0.316 on real calls): the commercial Russian training data has no telephone speech.
  • Language ID: 0.99 overall; Hebrew is never misdetected on telephony. On very short clips (<2 s) LID drops (ar 0.58, yi 0.83, ru 0.89), and on Hebrew read speech (FLEURS, CV) ~5% of clips are detected as Arabic — force language="he" when the language is known. He↔Yi confusion is 4 of 1,104 Yiddish clips and 0 of 7,875 Hebrew clips.

Known limitations

  • Hallucinates short phrases on speech-free audio (line noise, music, hold tones): 98% of real line-noise clips get a transcript, 27% of them "תודה רבה". A follow-up (v8) trains explicit speech-free negatives; until then gate the model with a VAD.
  • No conversational Yiddish test set exists; the Yiddish numbers are read speech.
  • Hebrew telephony references are soniox transcripts (agreement, not ground truth); the human-referenced Hebrew sets are the ivrit.ai suite above.
  • No code-switching (he+en in one utterance) evaluation.

Training

  • Data: 3.76M clips. Hebrew 53% (Voicenter 8 kHz telephony pseudo-labelled by soniox, ivrit.ai audio-v2 podcasts, Knesset, crowd recordings, synthetic TTS), English 19%, Russian 19%, Arabic 5%, Yiddish 3% (3× oversampled). Only commercially-licensed sources for en/ru/ar/yi. Labels keep case and punctuation where the source had it; Hebrew/Yiddish nikud stripped.
  • Label format: <|startoftranscript|><|lang|><|transcribe|><|notimestamps|> text <|endoftext|> with the row's own language token, so the language token is the first predicted token.
  • Knowledge distillation: loss = 0.2·CE + 0.8·T²·KL(teacher‖student), T=2, one frozen teacher per language sharing the Whisper large-v3 tokenizer — he: whisper-heb-v5 (ours), en: TheStageAI/thewhisper-large-v3-turbo, ru: openai/whisper-large-v3, ar: VoicenterTeam/whisper-levantine-hf, yi: ivrit-ai/yi-whisper-large-v3. The language token position is excluded from the KL.
  • Optimisation: lr 3e-6 (encoder 0.3×), cosine, 500 warmup, effective batch 256, bf16, label smoothing 0.1, SpecAugment, speed perturbation 0.9–1.1×, telephone band-pass augmentation.
  • Checkpoint selection on multi_wer: macro WER over 150-clip dev slices per language (sentence-disjoint from training) + ILSpeech speaker 1, auto language detection.

Code: the whisper-finetune repo (scripts/launch_multilingual_training.sh, src/kd.py, scripts/bench_v7.sh).

Data attribution

Hebrew: ivrit.ai (crowd-recital, crowd-transcribe, audio-v2; ivrit.ai license v2 — not for voice cloning), VoxKnesset, Voicenter telephony. English: Common Voice 17 (CC0), People's Speech (CC-BY), Granary/YODAS (CC-BY), VoxPopuli (CC0), AMI & ICSI (CC-BY-4.0), LibriSpeech (CC-BY-4.0), EdAcc (CC-BY-SA), NOTSOFAR-1 (CC-BY-4.0), DiPCo (CDLA-P), SLUE-HVB/VoxCeleb (CC-BY-4.0), Vystadial (CC-BY-SA-3.0), MLS (CC-BY-4.0), speechocean762 (CC-BY-4.0), MInDS-14 (CC-BY-4.0). Russian: Golos (SberDevices license), YODAS/Granary (CC-BY), Russian LibriSpeech (CC-BY-4.0), SOVA RuDevices (CC-BY-4.0), FLEURS (CC-BY-4.0), Common Voice (CC0), ToneWebinars (MIT), Dialogs (OpenRAIL). Arabic: Common Voice 17 (CC0), Omnilingual ASR corpus (CC-BY-4.0), Lahgtna Levantine TTS (CC-BY-4.0, synthetic), SADA, ClArTTS, ArVoice, FLEURS, SVQ, YODAS. Yiddish: ivrit.ai crowd-recital-yi / crowd-whatsapp-yi, Omnilingual ASR corpus ydd (CC-BY-4.0), REYD (CC-BY-4.0), Common Voice 22 (CC0). Teachers: TheStage AI (CC-BY-4.0), OpenAI Whisper (MIT), ivrit.ai (Apache-2.0).

Evaluation data notes

ILSpeech speaker 2 turned out to be Making History podcast audio that is also in the ivrit.ai audio-v2 training corpus; all ILSpeech numbers here use speaker 1 only (the full-set figure is contaminated for v5/v6/v7). Casablanca and Common Voice test sentences appear in some public Arabic fine-tunes' training data; none of ours.

Downloads last month
14
Safetensors
Model size
0.8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VoicenterTeam/whisper-heb-v7-multi

Finetuned
(13)
this model
Finetunes
1 model