Instructions to use VoicenterTeam/whisper-heb-v7-multi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VoicenterTeam/whisper-heb-v7-multi with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="VoicenterTeam/whisper-heb-v7-multi")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("VoicenterTeam/whisper-heb-v7-multi") model = AutoModelForSpeechSeq2Seq.from_pretrained("VoicenterTeam/whisper-heb-v7-multi", device_map="auto") - Notebooks
- Google Colab
- Kaggle
whisper-heb-v7-multi
A full fine-tune of ivrit-ai/whisper-large-v3-turbo (809M, 4 decoder layers) for Hebrew,
English, Arabic, Russian and Yiddish, with language identification trained in: the model
predicts the <|lang|> token itself, so you do not have to tell it which language a clip is in.
Built for short telephony segments (8 kHz call-center audio) at Voicenter; Hebrew is the primary
language, and the other four are what Israeli callers switch to.
Trained 2026-10-01/02 on one B200, 14,691 steps (one epoch over 3.76M clips, ~10.5k hours),
warm-started from our Hebrew-only whisper-heb-v6, with token-level knowledge distillation from
one teacher per language. Best checkpoint (step 14500) selected on the macro WER over five
held-out dev slices decoded with auto language detection.
Usage
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="VoicenterTeam/whisper-heb-v7-multi",
torch_dtype="bfloat16", device="cuda")
asr("call.wav") # language auto-detected
asr("call.wav", generate_kwargs={"language": "he"}) # or force one of he/en/ar/ru/yi
The shipped generation_config has language=None, so plain generate() runs Whisper's language
detection with the fine-tuned weights. Forcing the language gives slightly better WER on short
Arabic/Russian clips (see below) and is what to do when the language is known. The model is a
stock Whisper checkpoint: it converts to CTranslate2 / faster-whisper with ct2-transformers-converter
unchanged.
Results
WER on normalized text (nikud/tashkeel/punctuation stripped, case folded, ё→е, Arabic alef/ta
marbuta/ya folded), greedy decoding, bf16. auto = language detected by the model. Compared
with our previous Hebrew-only model (v6) and openai/whisper-large-v3-turbo.
Full tables incl. 8 kHz simulated-telephony condition, forced-language mode, CER and per-set
flags: results/whisper-heb-v7-multi/SUMMARY.md in the training repo.
Per language (macro over the language's test sets, clean audio, auto LID)
| lang | sets | v7 | v6 (Hebrew-only) | openai turbo | v7 LID acc |
|---|---|---|---|---|---|
| he | ILSpeech, 10k telephony | 0.128 | 0.134 | 0.240 | 1.00 |
| en | FLEURS+CV+AMI, AppTek call-center | 0.138 | 1.067 | 0.133 | 0.99 |
| ar | FLEURS, CV17, Casablanca jo/ps | 0.376 | 1.101 | 0.292 | 0.96 |
| ru | FLEURS+CV+Golos, ru.test | 0.114 | 1.038 | 0.122 | 0.99 |
| yi | Omnilingual ydd (read) | 0.347 | 0.983 | 1.196 | 1.00 |
| macro | 0.221 | 0.865 | 0.397 | 0.99 |
Under a simulated 8 kHz μ-law phone channel the macro is 0.239 (v6 0.867, openai 0.430).
Selected sets
| set | refs | v7 | v6 | openai turbo |
|---|---|---|---|---|
| he telephony, 10k (1000 sample) | soniox | 0.155 | 0.166 | 0.342 |
| he telephony, soniox+witness agree ≤0.2 WER (5,915) | soniox | 0.079 | 0.088 | 0.221 |
| he ILSpeech speaker 1 (579) | human | 0.126 | 0.125 | 0.148 |
| he whatsapp (ivrit.ai, long-form) | human | 0.071 | 0.068 | 0.152 |
| he Hebrew suite macro (d1, whatsapp, fleurs, CV, kan, saspeech; forced he) | human | 0.102 | 0.094 | 0.196 |
| en eval mix (FLEURS/CV/AMI) | human | 0.099 | 1.024 | 0.088 |
| en AppTek call-center (3000) | human | 0.181 | 1.093 | 0.183 |
| ru eval mix | human | 0.095 | 1.030 | 0.091 |
| ru test (CV/SOVA/RuLS/FLEURS) | human | 0.133 | 1.047 | 0.153 |
| ru real phone calls (Open STT, eval-only) | human | 0.388 | 1.154 | 0.316 |
| ar eval mix | human | 0.376 | 1.101 | 0.292 |
| yi Omnilingual (182) | human | 0.347 | 0.983 | 1.196 |
| yi prompt-disjoint held-out (1,104) | human | 0.377 | 0.995 | 1.155 |
What this means
- Hebrew telephony improved over the Hebrew-only v6 (10k: 0.166 → 0.155; agreement subset 0.088 → 0.079); long-form whatsapp and ILSpeech are unchanged within noise. Hebrew read speech regressed a little (FLEURS 0.180 → 0.204, Common Voice 0.134 → 0.150, forced-he).
- English and Russian are at openai-turbo level (within ±0.01), from WER ≈ 1.0 in v6, which had lost them entirely during Hebrew fine-tuning.
- Yiddish (Hebrew script, Hasidic orthography) works: 0.35 WER where openai turbo writes Latin-script German-ish text (1.2).
- Arabic is the weak spot: 0.376 vs 0.292 for openai turbo. Arabic was 5% of the training mix with a Levantine-biased teacher; the test set mixes MSA and several dialects.
- Russian telephony also lags openai turbo (0.388 vs 0.316 on real calls): the commercial Russian training data has no telephone speech.
- Language ID: 0.99 overall; Hebrew is never misdetected on telephony. On very short clips
(<2 s) LID drops (ar 0.58, yi 0.83, ru 0.89), and on Hebrew read speech (FLEURS, CV) ~5% of
clips are detected as Arabic — force
language="he"when the language is known. He↔Yi confusion is 4 of 1,104 Yiddish clips and 0 of 7,875 Hebrew clips.
Known limitations
- Hallucinates short phrases on speech-free audio (line noise, music, hold tones): 98% of real line-noise clips get a transcript, 27% of them "תודה רבה". A follow-up (v8) trains explicit speech-free negatives; until then gate the model with a VAD.
- No conversational Yiddish test set exists; the Yiddish numbers are read speech.
- Hebrew telephony references are soniox transcripts (agreement, not ground truth); the human-referenced Hebrew sets are the ivrit.ai suite above.
- No code-switching (he+en in one utterance) evaluation.
Training
- Data: 3.76M clips. Hebrew 53% (Voicenter 8 kHz telephony pseudo-labelled by soniox, ivrit.ai audio-v2 podcasts, Knesset, crowd recordings, synthetic TTS), English 19%, Russian 19%, Arabic 5%, Yiddish 3% (3× oversampled). Only commercially-licensed sources for en/ru/ar/yi. Labels keep case and punctuation where the source had it; Hebrew/Yiddish nikud stripped.
- Label format:
<|startoftranscript|><|lang|><|transcribe|><|notimestamps|> text <|endoftext|>with the row's own language token, so the language token is the first predicted token. - Knowledge distillation:
loss = 0.2·CE + 0.8·T²·KL(teacher‖student), T=2, one frozen teacher per language sharing the Whisper large-v3 tokenizer — he:whisper-heb-v5(ours), en:TheStageAI/thewhisper-large-v3-turbo, ru:openai/whisper-large-v3, ar:VoicenterTeam/whisper-levantine-hf, yi:ivrit-ai/yi-whisper-large-v3. The language token position is excluded from the KL. - Optimisation: lr 3e-6 (encoder 0.3×), cosine, 500 warmup, effective batch 256, bf16, label smoothing 0.1, SpecAugment, speed perturbation 0.9–1.1×, telephone band-pass augmentation.
- Checkpoint selection on
multi_wer: macro WER over 150-clip dev slices per language (sentence-disjoint from training) + ILSpeech speaker 1, auto language detection.
Code: the whisper-finetune repo (scripts/launch_multilingual_training.sh, src/kd.py,
scripts/bench_v7.sh).
Data attribution
Hebrew: ivrit.ai (crowd-recital, crowd-transcribe, audio-v2; ivrit.ai license v2 — not for voice cloning), VoxKnesset, Voicenter telephony. English: Common Voice 17 (CC0), People's Speech (CC-BY), Granary/YODAS (CC-BY), VoxPopuli (CC0), AMI & ICSI (CC-BY-4.0), LibriSpeech (CC-BY-4.0), EdAcc (CC-BY-SA), NOTSOFAR-1 (CC-BY-4.0), DiPCo (CDLA-P), SLUE-HVB/VoxCeleb (CC-BY-4.0), Vystadial (CC-BY-SA-3.0), MLS (CC-BY-4.0), speechocean762 (CC-BY-4.0), MInDS-14 (CC-BY-4.0). Russian: Golos (SberDevices license), YODAS/Granary (CC-BY), Russian LibriSpeech (CC-BY-4.0), SOVA RuDevices (CC-BY-4.0), FLEURS (CC-BY-4.0), Common Voice (CC0), ToneWebinars (MIT), Dialogs (OpenRAIL). Arabic: Common Voice 17 (CC0), Omnilingual ASR corpus (CC-BY-4.0), Lahgtna Levantine TTS (CC-BY-4.0, synthetic), SADA, ClArTTS, ArVoice, FLEURS, SVQ, YODAS. Yiddish: ivrit.ai crowd-recital-yi / crowd-whatsapp-yi, Omnilingual ASR corpus ydd (CC-BY-4.0), REYD (CC-BY-4.0), Common Voice 22 (CC0). Teachers: TheStage AI (CC-BY-4.0), OpenAI Whisper (MIT), ivrit.ai (Apache-2.0).
Evaluation data notes
ILSpeech speaker 2 turned out to be Making History podcast audio that is also in the ivrit.ai audio-v2 training corpus; all ILSpeech numbers here use speaker 1 only (the full-set figure is contaminated for v5/v6/v7). Casablanca and Common Voice test sentences appear in some public Arabic fine-tunes' training data; none of ours.
- Downloads last month
- 14
Model tree for VoicenterTeam/whisper-heb-v7-multi
Base model
openai/whisper-large-v3