Tool:Wikispeech/SpeechDataCollector
| Website | https://wikispeech-sdc.toolforge.org/ |
| Keywords | phabricator, python, admin |
| Author(s) | Sebastian Berlin, Viktoria Hillerud |
| Maintainer(s) | Sebastian Berlin, Viktoria Hillerud (View all) |
| License | No license specified |
| Issues | Open tasks · Report a bug |
The Wikispeech Speech Data Collector is a Toolforge tool for collecting speech data for Wikispeech. It lets users record themselves reading prompts, producing speech recordings that can be used to train and improve text-to-speech (and potentially speech-to-text) models for Wikimedia projects.
The tool is built on top of the Lingua Libre code base and is developed in close collaboration with the Lingua Libre team. It is developed and maintained by Wikimedia Sverige. It was previously developed as a MediaWiki extension (Extension:WikispeechSpeechDataCollector, now archived); development has since moved to Toolforge.
This page contains technical documentation aimed at maintainers and developers. For general information about the speech data collection project, see Wikispeech/Speech data collection on Meta.
Complete list of configuration options
| Config file | Default value | Documentation |
|---|---|---|
| config.ini & .env.localhost | ALLOW_ANONYMOUS=false
VITE_ALLOW_ANONYMOUS=false
|
To test the tool without having an account. |