datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EuroSpeech
EuroSpeech Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.disrpt
Disrpt is a multilingual, multi-framework unified discourse analysis benchmark.
It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages.
⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB.
To load these datasets, run the following:
pip install disrpt-utils
Then
from disrpt_utils import load_dataset
corpora_paths={
# ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.AIMEfrom datasets import load_dataset
dataset = load_dataset('disco-eth/AIME')
AIME: AI Music Evaluation Dataset
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset.
The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset.
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.equity-perp-price-discovery
Equity and pre-IPO perpetual prices
Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets.
Contents
Table
Record
perpetual_mark_and_index_prices
An instrument's mark price, index price and basis at an observation time
Using the data
market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.EuroSpeech-24kHz
EuroSpeech 24 kHz Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
Dataset Summary
Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.discofuse
Dataset Card for "discofuse"
Dataset Summary
DiscoFuse is a large scale dataset for discourse-based sentence fusion.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
discofuse-sport
Size of downloaded dataset files: 4.33 GB
Size of the generated dataset: 15.04 GB
Total amount of disk used: 19.36 GB
An example of 'train' looks as follows.
{… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.GlobalDISCO
GlobalDISCO
GlobalDISCO is a large-scale dataset consisting of 73k music tracks generated by state-of-the-art commercial generative music models, along with paired links to 93k reference tracks in LAION-DISCO-12M. The dataset spans 147 languages and includes musical style prompts extracted from MusicBrainz and Wikipedia. The dataset is globally balanced, representing musical styles from artists across 79 countries and five continents. It is aimed to support the research community in… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/GlobalDISCO.caml-animal-discourse-2020-present
Reddit Animal-Discourse Corpus — CLEANED (2020–present)
Submissions and comments from animal-relevant subreddits, gathered via
PullPush.io, covering January 2020 to the present.
Built as part of research on AI-mediated value lock-in in human animal-welfare
discourse.
Coverage
Subreddit
Submissions
Comments
Date range (submissions)
r/AnimalRights
15,719
34,686
2020-01-01 → 2025-05-19
r/AntiVegan
17,252
182,890
2020-01-01 → 2025-05-19
r/AskVegans
4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.discovery
Dataset Card for Discovery
Dataset Summary
Discourse marker prediction with 174 markers
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
input : sentence1, sentence2,
label: marker originally between sentence1 and sentence2
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
Train/Val/Test
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/sileod/discovery.Discord-Dialogues
Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format.
This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words.
Nomic Atlas Map
Features
Mixed single and multi-turn exchanges
Human-only dialogues (no bots)
Filtered for ToS and harmful contentLinks… See the full description on the dataset page: https://huggingface.co/datasets/mookiezi/Discord-Dialogues.DisciplineGen-1Mdischarge_target
Dataset Card for "discharge_target"
More Information needed
discovery_discovery_promptsourcediscard_tileomy_f3m_multi_Disconnect-0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_multi",
"total_episodes": 100,
"total_frames": 61744,
"total_tasks": 1,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/omy_f3m_multi_Disconnect-0.LAION-DISCO-12MThe LAION-DISCO-12M dataset contains 12M links to music on YouTube, inspired by the methodology of DISCO-10M. It contains song metadata (song_id, title, artist_names, artist_ids, album_name, album_id, isExplicit, views, duration) and YouTube URL, pointing to the original song on the public web. It does not contain any original audio samples and is thus an index dataset.
Starting from an initial seed list of artists, we can discover new artists by recursively exploring the artists listed in the… See the full description on the dataset page: https://huggingface.co/datasets/laion/LAION-DISCO-12M.MovieTection
Dataset Description 🎬
The MovieTection dataset is a benchmark designed for detecting pretraining data in Large Vision-Language Models (VLMs). It serves as a resource for analyzing model exposure to Copyrighted Visual Content ©️.
Paper: DIS-CO: Discovering Copyrighted Content in VLMs Training Data
Direct Use 🖥️
The dataset is designed for image/caption-based question-answering, where models predict the movie title given a frame or its corresponding textual… See the full description on the dataset page: https://huggingface.co/datasets/DIS-CO/MovieTection.reddit-control-discourse-2016-present-pretau
Reddit Control Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records
(75.1%) from 1,030,104 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.informes_discriminacion_gitana
Resumen del dataset
Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.SWE-bench_Verified-discriminative
SWE-bench Verified Discriminative Subsets
Dataset Description
This dataset contains discriminative subsets of SWE-bench Verified designed to provide more sensitive evaluation of SWE-agent capabilities. As top-performing agents achieve 73%+ on the full benchmark, these subsets focus on truly challenging problems to better discriminate between cutting-edge systems.
Key Features
4 discriminative splits targeting different evaluation needs
335 carefully selected… See the full description on the dataset page: https://huggingface.co/datasets/jatinganhotra/SWE-bench_Verified-discriminative.reddit-animal-discourse-2016-present-pretau
Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records
(78.0%) from 343,756 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.AgentsNet
AgentsNet
This repository contains the graph instances used in the AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs paper.
AgentsNet is a new benchmark for multi-agent reasoning, designed to measure the ability of multi-agent systems to collaboratively form strategies for problem-solving, self-organization, and effective communication given a network topology. It draws inspiration from classical problems in distributed systems and graph theory.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AgentsNet.disco-model-outputs
DISCO model outputs
Tabular release of per-model, per-item correctness and answer scores used to train and evaluate DISCO: Diversifying Sample Condensation for Efficient Model Evaluation. The paper studies cheap benchmark performance prediction from a small subset of evaluation items; this dataset supplies the raw harness-style outputs for MMLU (57 subjects), HellaSwag, Winogrande, ARC, and related tasks from the Open LLM Leaderboard ecosystem.
Paper
Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/arubique/disco-model-outputs.eval-deployment-discriminationcoarse_discourse
Dataset Card for "coarse_discourse"
Dataset Summary
A large corpus of discourse annotations and relations on ~10K forum threads.
We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.ontario-lobbying-disclosure-graph
Ontario Lobbying & MPP Disclosure Graph
A structured, entity-resolved projection of three public Ontario government
records sources, exported as flat, documented parquet tables:
Ontario Lobbyist Registry (Office of the Integrity Commissioner of
Ontario, lobbyist.oico.on.ca) — lobbyist registrations: who is registered to
lobby, for which client, about what, aimed at which offices.
MPP Public Disclosure Statements (Office of the Integrity Commissioner
of Ontario, PDS) — annual… See the full description on the dataset page: https://huggingface.co/datasets/agoulah/ontario-lobbying-disclosure-graph.reddit-animal-discourse-2016-present
Reddit Animal Discourse (2016-present)
Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed.
Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.phased-self-discover-mistral-structured-5-shot-bbh-evalpanda_pick_cube_demos_sim_discrete_newThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 30,
"total_frames": 3570,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/KeWangRobotics/panda_pick_cube_demos_sim_discrete_new.franka-discrete-reach4absThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka_panda",
"total_episodes": 371,
"total_frames": 27053,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:371"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ajacquet/franka-discrete-reach4abs.
