Jump to content

Tool:Wikispeech/SpeechDataCollector

From Wikitech
Toolforge tools
Website https://wikispeech-sdc.toolforge.org/
Keywords phabricator, python, admin
Author(s) Sebastian Berlin, Viktoria Hillerud
Maintainer(s) Sebastian Berlin, Viktoria Hillerud (View all)
License No license specified
Issues Open tasks · Report a bug

The Wikispeech Speech Data Collector is a Toolforge tool for collecting speech data for Wikispeech. It lets users record themselves reading prompts, producing speech recordings that can be used to train and improve text-to-speech (and potentially speech-to-text) models for Wikimedia projects.

The tool is built on top of the Lingua Libre code base and is developed in close collaboration with the Lingua Libre team. It is developed and maintained by Wikimedia Sverige. It was previously developed as a MediaWiki extension (Extension:WikispeechSpeechDataCollector, now archived); development has since moved to Toolforge.

This page contains technical documentation aimed at maintainers and developers. For general information about the speech data collection project, see Wikispeech/Speech data collection on Meta.

Complete list of configuration options

Config file Default value Documentation
config.ini & .env.localhost
ALLOW_ANONYMOUS=false
&
VITE_ALLOW_ANONYMOUS=false
To test the tool without having an account.