datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Egyptian-Arabic-Lectures
Egyptian Arabic Lectures Dataset
The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts.
Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.feynman-audio-lectures-dataset
The Feynman Lectures on Physics - Audio Dataset
Dataset Description
This dataset contains the complete audio recordings from Richard Feynman's famous physics lectures at the California Institute of Technology, delivered between 1961-1964. The dataset includes all three volumes of "The Feynman Lectures on Physics" with rich metadata and quality metrics.
Dataset Summary
Total Lectures: ~100+ audio recordings
Format: M4A (original format preserved)
Speaker:… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/feynman-audio-lectures-dataset.
