CoolFace
Datasetpublic

MEscriva/french-education-speech

French Education Speech - Transcribed Dataset High-quality French educational speech dataset transcribed with OpenAI Whisper API, prepared for training automatic speech recognition (ASR) models. Dataset Summary This dataset contains 3,933 transcribed audio segments from the French educational domain, totaling approximately 12.82 hours of audio. All transcriptions were performed using OpenAI Whisper API (optimized Whisper-1 model) to ensure maximum accuracy… See the full description on the dataset page: https://huggingface.co/datasets/MEscriva/french-education-speech.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
4likes145downloads
8 commits on main
83db6f89mo ago

Fix README: remove non-official task category

MEscriva
2838c8f9mo ago

Add transcriptions: 3,933 segments transcribed with OpenAI Whisper API, hallucinations removed (0.93%)

MEscriva
a7834399mo ago

Update README with transcription methodology

MEscriva
90019c09mo ago

Upload dataset

MEscriva
1f2c72811mo ago

Add pipeline ASCII diagram and duration distribution figure

mathisescriva
a54652911mo ago

Update stats to reflect current dataset

mathisescriva
152322b11mo ago

Add YAML metadata to README

mathisescriva
5d2c0a911mo ago

Initial commit: French Education Speech Corpus (Phase 1)

mathisescriva