MEscriva/french-education-speech
French Education Speech - Transcribed Dataset High-quality French educational speech dataset transcribed with OpenAI Whisper API, prepared for training automatic speech recognition (ASR) models. Dataset Summary This dataset contains 3,933 transcribed audio segments from the French educational domain, totaling approximately 12.82 hours of audio. All transcriptions were performed using OpenAI Whisper API (optimized Whisper-1 model) to ensure maximum accuracy… See the full description on the dataset page: https://huggingface.co/datasets/MEscriva/french-education-speech.
Fix README: remove non-official task category
Add transcriptions: 3,933 segments transcribed with OpenAI Whisper API, hallucinations removed (0.93%)
Update README with transcription methodology
Upload dataset
Add pipeline ASCII diagram and duration distribution figure
Update stats to reflect current dataset
Add YAML metadata to README
Initial commit: French Education Speech Corpus (Phase 1)
