Mozilla's crowdsourced speech corpus — 31K hours, 120+ languages, CC0.
Common Voice 17.0 is Mozilla's crowdsourced speech corpus — the largest openly licensed voice dataset in the world. Volunteers read sentences aloud, other contributors validate the clips, and everything is released under CC0: no attribution, no usage restrictions.
Over 31,000 hours of recorded MP3 clips (roughly 20,000 hours validated) in 120+ languages, each with a sentence-level transcription and speaker metadata: age bracket, sex, accent. Every language ships with reviewed train/dev/test splits. Coverage runs from English with millions of clips down to dozens of low-resource languages; Russian alone offers over 150 hours of validated speech.
The go-to corpus for training and benchmarking ASR neuronets such as Whisper or wav2vec2. CC0 licensing means models fine-tuned on it can be deployed commercially without legal friction, and the fixed per-language test splits give miners and validators a reproducible accuracy yardstick across languages — including Russian. Note: the HF repo is gated (speaker-privacy terms must be accepted), but the license itself is unrestricted.