labari-voice/fr-sn-speech-pilot
fr_sn, Senegalese French read-speech pilot Read speech in Senegalese French, in the FLEURS format. The sentences were written in Senegal, about local realities, and read by Senegalese speakers in their own French. For evaluation, not for training. See DATASHEET.md for provenance and intended use. Segments 210 Sentences 105, all covered Speakers 5 (3 female, 2 male) Duration 19.3 min Words 2832 Audio WAV PCM 16-bit, 16 kHz, mono Split single test split… See the full description on the dataset page: https://huggingface.co/datasets/labari-voice/fr-sn-speech-pilot.
fr_sn, Senegalese French read-speech pilot
Read speech in Senegalese French, in the FLEURS format. The sentences were written in Senegal, about local realities, and read by Senegalese speakers in their own French. For evaluation, not for training.
See DATASHEET.md for provenance and intended use.
Files
audio/test/ segments + metadata.csv
test.tsv FLEURS metadata
speakers.tsv speaker table
manifest.jsonl per-segment measurements
checksums.sha256 audio integrity
DATASHEET.md datasheettest.tsv follows the FLEURS schema: no header, seven tab-separated columns.
An id appears once per speaker who read it. File names encode utterance and speaker as NNNSS.wav. Transcriptions come from the reading script, not from automatic speech recognition.
Speakers
Age follows the Common Voice vocabulary. Identifiers are pseudonymised.
Measurements
Per-segment values are in manifest.jsonl and allow filtering before use.
Limitations
- Volume. Twenty minutes, against roughly twelve hours per language in FLEURS. Not sized for training.
- Statistical power. With 2832 words, an error rate carries a margin of about ± 1.5 point at a 20 % error rate.
- Speaker imbalance, from 10 to 99 segments per voice. Weight accordingly.
- Two single-speaker domains, Telecommunications and Administration.
- Read speech, not spontaneous speech.
- Partition on
idrather than on segments, so that every voice reading a given text stays on the same side.
Licence
CC BY 4.0: use, redistribution and derivative works, including commercial use, with attribution.
Recordings were made with narrators under contract in Labari Voice studios, with informed consent for public release. Speaker identifiers are pseudonymised and carry no directly identifying data.
For larger deliveries or other language varieties: sales@labari.dev
