datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WorldSpeech
WorldSpeech
🎉 WorldSpeech has been accepted to NeurIPS 2026! 🎉See the paper on arXiv.
A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/WorldSpeech.EuroSpeech
EuroSpeech Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.disrpt
Disrpt is a multilingual, multi-framework unified discourse analysis benchmark.
It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages.
⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB.
To load these datasets, run the following:
pip install disrpt-utils
Then
from disrpt_utils import load_dataset
corpora_paths={
# ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.discover-toolsdiscord-chatAIMEfrom datasets import load_dataset
dataset = load_dataset('disco-eth/AIME')
AIME: AI Music Evaluation Dataset
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset.
The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset.
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.equity-perp-price-discovery
Equity and pre-IPO perpetual prices
Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets.
Contents
Table
Record
perpetual_mark_and_index_prices
An instrument's mark price, index price and basis at an observation time
Using the data
market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.discoverybenchData-driven Discovery Benchmark from the paper:
"DiscoveryBench: Towards Data-Driven Discovery with Large Language Models"
🔭 Overview
DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.EuroSpeech-24kHz
EuroSpeech 24 kHz Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
Dataset Summary
Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.discofuse
Dataset Card for "discofuse"
Dataset Summary
DiscoFuse is a large scale dataset for discourse-based sentence fusion.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
discofuse-sport
Size of downloaded dataset files: 4.33 GB
Size of the generated dataset: 15.04 GB
Total amount of disk used: 19.36 GB
An example of 'train' looks as follows.
{… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.gspc-custody-disclosure
GSPC — custody disclosure facts (CustodyFacts)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
MEASURED financial/domain axis (named-string presence on retrieved pages over live XRPL reader-16, n=16). Not a model leaderboard. No accuracy, no fleet, no leader.
Live status is the custody-disclosure row on GET https://councilof.ai/api/gspc. Not a certificate.
Tokenisation evidence question: What can an outsider verify after… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-custody-disclosure.discrim-eval
Dataset Card for Discrim-Eval
Dataset Summary
The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials.
Each prompt instructs the model to make a binary decision (yes/no)
about a particular person described in the prompt.
Each person is described in terms of three demographic attributes:
age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary)
, and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.GlobalDISCO
GlobalDISCO
GlobalDISCO is a large-scale dataset consisting of 73k music tracks generated by state-of-the-art commercial generative music models, along with paired links to 93k reference tracks in LAION-DISCO-12M. The dataset spans 147 languages and includes musical style prompts extracted from MusicBrainz and Wikipedia. The dataset is globally balanced, representing musical styles from artists across 79 countries and five continents. It is aimed to support the research community in… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/GlobalDISCO.caml-animal-discourse-2020-present
Reddit Animal-Discourse Corpus — CLEANED (2020–present)
Submissions and comments from animal-relevant subreddits, gathered via
PullPush.io, covering January 2020 to the present.
Built as part of research on AI-mediated value lock-in in human animal-welfare
discourse.
Coverage
Subreddit
Submissions
Comments
Date range (submissions)
r/AnimalRights
15,719
34,686
2020-01-01 → 2025-05-19
r/AntiVegan
17,252
182,890
2020-01-01 → 2025-05-19
r/AskVegans
4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.discovery
Dataset Card for Discovery
Dataset Summary
Discourse marker prediction with 174 markers
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
input : sentence1, sentence2,
label: marker originally between sentence1 and sentence2
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
Train/Val/Test
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/sileod/discovery.disco-elysium-utterancesDisco Elysium voices dataset. This is not meant to be used with HF's datasets library; please git lfs clone https://huggingface.co/datasets/main-horse/disco-elysium-utterances and use the files directly.
Directory structure:
processed
├── Character A
│ ├── metadata.test.txt
│ ├── metadata.train.txt
│ ├── metadata.txt
│ └── wavs
├── Character B
├── ...
└── Narrator
├── metadata.test.txt
├── metadata.train.txt
├── metadata.txt
└── wavs.zip
Some directories have wavs.zip… See the full description on the dataset page: https://huggingface.co/datasets/main-horse/disco-elysium-utterances.Discord-Unveiled-Extracted
Discord Unveiled - Filtered Dataset
This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset.
Data Processing
The data has been processed to:
Convert JSON data to CSV format.
Remove messages from bots.
Filter out messages containing only URLs, mentions, channels or discord emojis.
Filter out messages that are not in English using a FastText language identification model.
Data Fields
The CSV files in… See the full description on the dataset page: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.Discord-Dialogues
Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format.
This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words.
Nomic Atlas Map
Features
Mixed single and multi-turn exchanges
Human-only dialogues (no bots)
Filtered for ToS and harmful contentLinks… See the full description on the dataset page: https://huggingface.co/datasets/mookiezi/Discord-Dialogues.DisciplineGen-1Mdischarge_target
Dataset Card for "discharge_target"
More Information needed
dblp-discovery-dataset
Dataset Card for DBLP Discovery Dataset (D3)
Dataset Summary
DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.cineaudiosynth
CineAudioSynth
Synthetic cinematic audio for source separation. 453 scenes, ~22.8 h, 48 kHz / 16-bit / stereo WAV.
Each data/scene_NNNN/ contains:
linear/ — additive render: mix.wav is the BIT-EXACT 16-bit sum of the four stems
(speech, music, ambience, sfx) — max |mix − Σstems| = 0, verified per scene.
Includes gain_envelope.json (sidechain ducking envelopes).
release/ — mastered render of the same scene (compression/limiting/loudness on the
mix bus; intentionally… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/cineaudiosynth.discovery_discovery_promptsourcecineaudiodb
CineAudioDB
A real-world evaluation set for cinematic audio source separation.
CineAudioDB contains real film/animation productions with ground-truth stems
(dialogue, music, sfx) for evaluating cinematic source-separation models.
Unlike synthetic datasets that sum stems linearly, real productions are mixed with a
non-linear mastering chain (compression, limiting, sidechain ducking, reverb), so the
released stems do not generally sum to the mastered mix. To support fair… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/cineaudiodb.2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment.
field
value
experiment
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.stable-diffusion-discord-promptsstable-diffusion-discord-prompts
All messages from dreambot from all dream-[1-50] channels in stable-diffusion discord
source:
https://github.com/bartman081523/stable-diffusion-discord-prompts
Discord-Unveiled-Compressed
.hf-sanitized.hf-sanitized-uCWd6SwyNH8FCkRETeRYS .container { --bg-primary: #0d0511; --bg-secondary: #1a0f1f; --bg-tertiary: #2d1b35; --bg-card: #3d2847; --text-primary: #fef7ff; --text-secondary: #f0d9ff; --text-muted: #c084fc; --pink-soft: #fce7f3; --pink-medium: #f9a8d4; --pink-bright: #ec4899; --pink-hot: #e91e63; --pink-neon: #ff1493; --purple-soft: #e879f9; --purple-bright: #c026d3; --purple-deep: #7c3aed; --border-glow: #f472b6; --shadow-pink: rgba(244, 114, 182, 0.4);… See the full description on the dataset page: https://huggingface.co/datasets/SaisExperiments/Discord-Unveiled-Compressed.LAION-DISCO-12MThe LAION-DISCO-12M dataset contains 12M links to music on YouTube, inspired by the methodology of DISCO-10M. It contains song metadata (song_id, title, artist_names, artist_ids, album_name, album_id, isExplicit, views, duration) and YouTube URL, pointing to the original song on the public web. It does not contain any original audio samples and is thus an index dataset.
Starting from an initial seed list of artists, we can discover new artists by recursively exploring the artists listed in the… See the full description on the dataset page: https://huggingface.co/datasets/laion/LAION-DISCO-12M.AgentsNet
AgentsNet
This repository contains the graph instances used in the AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs paper.
AgentsNet is a new benchmark for multi-agent reasoning, designed to measure the ability of multi-agent systems to collaboratively form strategies for problem-solving, self-organization, and effective communication given a network topology. It draws inspiration from classical problems in distributed systems and graph theory.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AgentsNet.DISC-Med-SFTThis is a repository containing a subset of the DISC-Med-SFT Dataset.
Check DISC-MedLLM for more information.
