datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biomedical_lectures_v2
Vidore Benchmark 2 - MIT Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_v2.biomedical_lectures_eng_v2
Vidore Benchmark 2 - MIT Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/biomedical_lectures_eng_v2.Egyptian-Arabic-Lectures
Egyptian Arabic Lectures Dataset
The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts.
Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.feynman-audio-lectures-dataset
The Feynman Lectures on Physics - Audio Dataset
Dataset Description
This dataset contains the complete audio recordings from Richard Feynman's famous physics lectures at the California Institute of Technology, delivered between 1961-1964. The dataset includes all three volumes of "The Feynman Lectures on Physics" with rich metadata and quality metrics.
Dataset Summary
Total Lectures: ~100+ audio recordings
Format: M4A (original format preserved)
Speaker:… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/feynman-audio-lectures-dataset.karpathy-lectures-transcriptslong_audio_youtube_lecturesAudio dataset with Russian speech of scientific lectures from YouTube.
The dataset contains seven long 20-40 minute Russian audios collected as a test set for our ASR system. The audios belong to different lexical and speech domains; they are parts of several Russian scientific lectures on various subjects: philology, mathematics, history, etc.
All recordings were made in relatively quiet acoustic environments typical of lecture halls; however, some background noises, such as the sound of… See the full description on the dataset page: https://huggingface.co/datasets/dangrebenkin/long_audio_youtube_lectures.Translated_Egyptian_Arabic_Lecturestextbooks_lectures_stemwikibiomedical_lectures_eng_v2
Vidore Benchmark 2 - MIT Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of MIT courses in anatomy (precisely tissue interactions).
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to MIT biology courses. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/aakash-projects/biomedical_lectures_eng_v2.ViDoRe_biomedical_lectures_v2_multilingualViDoRe_biomedical_lectures_v2biomedical_lectures_eng_v2_reasoningpolish_lectures_wolnelektury-plyoutube-transcripts-as-lecturesvidore_benchmark_2_biomedical_lectures_v2_reranker_adapted
Dataset Card for Vidore Reranker Benchmark : vidore_benchmark_2_biomedical_lectures_v2_reranker_adapted
Dataset Summary
This dataset provides a reranking benchmark based on the VIDORE V2 benchmark, designed to evaluate reranker models in a multimodal retrieval context. The dataset includes a corpus of image data, a set of natural language queries, and the top 25 retrievals (images) returned by a mid-performance multimodal retriever. This setup simulates a realistic… See the full description on the dataset page: https://huggingface.co/datasets/UlrickBL/vidore_benchmark_2_biomedical_lectures_v2_reranker_adapted.textbooks_lectures_glossaries_updatedgenerated_lectureslecture_summary_translations_english_spanishtextbooks_lectures_glossaries_stemwikitextbooks_lectures_glossaries_cleanedtm-lecturestextbooks_lectures_glossaries_cleaned_stemwikiLectures-test-V1textbooks_lectures_glossaries_stemwiki_filteredtextbooks_lectures_glossaries_cleaned_v2lecturestextbooks_lectures_stemwiki_chunks_512textbooks_lectures_stemwiki_chunks_256textbooks_lectures_glossariestextbooks_lectures_glossaries_cleaned_v2_stemwiki_chunks
