datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biomedical_lectures_v2
Vidore Benchmark 2 - MIT Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_v2.ml-lectures
ML/Math Lecture Archive
Archived lecture videos (1080p MP4) with English subtitles (.vtt) from publicly
available university course recordings on YouTube.
268 videos, ~72 GB.
Contents
Folder
Course
Videos
18.065_Strang/
MIT 18.065 — Matrix Methods in Data Analysis, Signal Processing, and Machine Learning (Gilbert Strang)
36
18.06SC_LinearAlgebra/
MIT 18.06SC — Linear Algebra, Fall 2011 (Gilbert Strang)
74
CS109_Piech/
Stanford CS109 — Introduction… See the full description on the dataset page: https://huggingface.co/datasets/olive5/ml-lectures.biomedical_lectures_eng_v2
Vidore Benchmark 2 - MIT Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_eng_v2.Egyptian-Arabic-Lectures
Egyptian Arabic Lectures Dataset
The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts.
Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.feynman-audio-lectures-dataset
The Feynman Lectures on Physics - Audio Dataset
Dataset Description
This dataset contains the complete audio recordings from Richard Feynman's famous physics lectures at the California Institute of Technology, delivered between 1961-1964. The dataset includes all three volumes of "The Feynman Lectures on Physics" with rich metadata and quality metrics.
Dataset Summary
Total Lectures: ~100+ audio recordings
Format: M4A (original format preserved)
Speaker:… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/feynman-audio-lectures-dataset.karpathy-lectures-transcriptslong_audio_youtube_lecturesAudio dataset with Russian speech of scientific lectures from YouTube.
The dataset contains seven long 20-40 minute Russian audios collected as a test set for our ASR system. The audios belong to different lexical and speech domains; they are parts of several Russian scientific lectures on various subjects: philology, mathematics, history, etc.
All recordings were made in relatively quiet acoustic environments typical of lecture halls; however, some background noises, such as the sound of… See the full description on the dataset page: https://huggingface.co/datasets/dangrebenkin/long_audio_youtube_lectures.Translated_Egyptian_Arabic_Lecturestextbooks_lectures_stemwikibiomedical_lectures_eng_v2
Vidore Benchmark 2 - MIT Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/biomedical_lectures_eng_v2.ViDoRe_biomedical_lectures_v2_multilingualverifiedtutor-lecturesViDoRe_biomedical_lectures_v2biomedical_lectures_eng_v2_reasoningpolish_lectures_wolnelektury-plyoutube-transcripts-as-lecturesvidore_benchmark_2_biomedical_lectures_v2_reranker_adapted
Dataset Card for Vidore Reranker Benchmark : vidore_benchmark_2_biomedical_lectures_v2_reranker_adapted
Dataset Summary
This dataset provides a reranking benchmark based on the VIDORE V2 benchmark, designed to evaluate reranker models in a multimodal retrieval context. The dataset includes a corpus of image data, a set of natural language queries, and the top 25 retrievals (images) returned by a mid-performance multimodal retriever. This setup simulates a realistic… See the full description on the dataset page: https://huggingface.co/datasets/UlrickBL/vidore_benchmark_2_biomedical_lectures_v2_reranker_adapted.textbooks_lectures_glossaries_updatedLecturesgenerated_lecturesncert-lectures-india
license: other
task_categories:
audio-classification
automatic-speech-recognition
language:
en
hi
bn
ta
te
ml
mr
or
as
pa
tags:
ncert,
education,
india,
upsc,
humanities,
recorded-lectures,
government-exams,
science,
long-form-audio
pretty_name: NCERT Lectures India
size_categories:
1K<n<10K
NCERT Lectures India
Dataset Description
A large-scale collection of recorded lectures covering NCERT curriculum
across Science and Humanities streams, specifically… See the full description on the dataset page: https://huggingface.co/datasets/DataOrigin/ncert-lectures-india.lecture_summary_translations_english_spanishai-lectures-spring-24This is a dataset containing lectures where the instructor draws under a document camera.
The dataset can be used to train multimodal representation learning approaches that can be used by Q&A agents.
Questions may come in the form of a natural language query "what are the advantages of CNNs over Dense networks" and the system should respond back with the
video clip / segment that contains the video explanation.
The_Feynman_Lectures_on_Physics_Textbook_by_Matthew_Sands_Richard_Feynman_and_Robert_B_Leighton
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [enesxgrahovac]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/harisnaeem/The_Feynman_Lectures_on_Physics_Textbook_by_Matthew_Sands_Richard_Feynman_and_Robert_B_Leighton.textbooks_lectures_glossaries_stemwikitextbooks_lectures_glossaries_cleanedtm-lecturestextbooks_lectures_glossaries_cleaned_stemwikieduclaw-lectureslectures
