CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01disco-eth /WorldSpeech WorldSpeech 🎉 WorldSpeech has been accepted to NeurIPS 2026! 🎉See the paper on arXiv. A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/WorldSpeech.audioautomatic-speech-recognition10M<n<100M51 likes41k downloads54m agoHugging Face02disco-eth /EuroSpeech EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.audioautomatic-speech-recognition10M<n<100M99 likes32k downloads5mo agoHugging Face03multilingual-discourse-hub /disrpt Disrpt is a multilingual, multi-framework unified discourse analysis benchmark. It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages. ⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB. To load these datasets, run the following: pip install disrpt-utils Then from disrpt_utils import load_dataset corpora_paths={ # ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.text100K<n<1M3 likes12k downloads1y agoHugging Face04mcp-tools /discover-toolstextn<1K5 likes6.9k downloads1mo agoHugging Face05breadlicker45 /discord-chattext10K<n<100K6 likes3.5k downloads3y agoHugging Face06disco-eth /AIMEfrom datasets import load_dataset dataset = load_dataset('disco-eth/AIME') AIME: AI Music Evaluation Dataset The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo. The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset. The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset. The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.audio1K<n<10K9 likes2.4k downloads2y agoHugging Face07dataforge-labs /equity-perp-price-discovery Equity and pre-IPO perpetual prices Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets. Contents Table Record perpetual_mark_and_index_prices An instrument's mark price, index price and basis at an observation time Using the data market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.tabulartime-series-forecasting10K<n<100K1 likes2.1k downloads2h agoHugging Face08allenai /discoverybenchData-driven Discovery Benchmark from the paper: "DiscoveryBench: Towards Data-Driven Discovery with Large Language Models" 🔭 Overview DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.texttext-generationn<1K18 likes1.7k downloads1y agoHugging Face09disco-eth /EuroSpeech-24kHz EuroSpeech 24 kHz Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. Dataset Summary Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.audioautomatic-speech-recognition10M<n<100M3 likes1.7k downloads5mo agoHugging Face10google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.3k downloads3y agoHugging Face11csoai /gspc-custody-disclosure GSPC — custody disclosure facts (CustodyFacts) SWIFT census (live): https://councilof.ai/api/swift XRPL reader (live): https://councilof.ai/api/xrpl MEASURED financial/domain axis (named-string presence on retrieved pages over live XRPL reader-16, n=16). Not a model leaderboard. No accuracy, no fleet, no leader. Live status is the custody-disclosure row on GET https://councilof.ai/api/gspc. Not a certificate. Tokenisation evidence question: What can an outsider verify after… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-custody-disclosure.tabularothern<1K0 likes1.2k downloads1d agoHugging Face12Anthropic /discrim-eval Dataset Card for Discrim-Eval Dataset Summary The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials. Each prompt instructs the model to make a binary decision (yes/no) about a particular person described in the prompt. Each person is described in terms of three demographic attributes: age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary) , and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.tabularquestion-answering10K<n<100K60 likes1.1k downloads3y agoHugging Face13disco-eth /GlobalDISCO GlobalDISCO GlobalDISCO is a large-scale dataset consisting of 73k music tracks generated by state-of-the-art commercial generative music models, along with paired links to 93k reference tracks in LAION-DISCO-12M. The dataset spans 147 languages and includes musical style prompts extracted from MusicBrainz and Wikipedia. The dataset is globally balanced, representing musical styles from artists across 79 countries and five continents. It is aimed to support the research community in… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/GlobalDISCO.audioaudio-classification10K<n<100K1 likes1k downloads10mo agoHugging Face14CompassioninMachineLearning /caml-animal-discourse-2020-present Reddit Animal-Discourse Corpus — CLEANED (2020–present) Submissions and comments from animal-relevant subreddits, gathered via PullPush.io, covering January 2020 to the present. Built as part of research on AI-mediated value lock-in in human animal-welfare discourse. Coverage Subreddit Submissions Comments Date range (submissions) r/AnimalRights 15,719 34,686 2020-01-01 → 2025-05-19 r/AntiVegan 17,252 182,890 2020-01-01 → 2025-05-19 r/AskVegans 4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.tabulartext-classification1M<n<10M0 likes811 downloads3mo agoHugging Face15sileod /discovery Dataset Card for Discovery Dataset Summary Discourse marker prediction with 174 markers Supported Tasks and Leaderboards [More Information Needed] Languages English Dataset Structure input : sentence1, sentence2, label: marker originally between sentence1 and sentence2 Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits Train/Val/Test Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/sileod/discovery.texttext-classification1M<n<10M8 likes676 downloads2y agoHugging Face16main-horse /disco-elysium-utterancesDisco Elysium voices dataset. This is not meant to be used with HF's datasets library; please git lfs clone https://huggingface.co/datasets/main-horse/disco-elysium-utterances and use the files directly. Directory structure: processed ├── Character A │   ├── metadata.test.txt │   ├── metadata.train.txt │   ├── metadata.txt │   └── wavs ├── Character B ├── ... └── Narrator ├── metadata.test.txt ├── metadata.train.txt ├── metadata.txt └── wavs.zip Some directories have wavs.zip… See the full description on the dataset page: https://huggingface.co/datasets/main-horse/disco-elysium-utterances.audio10K<n<100K2 likes648 downloads2y agoHugging Face17ManBib /Discord-Unveiled-Extracted Discord Unveiled - Filtered Dataset This dataset contains superficially filtered and processed Discord message data from the Discord Unveiled dataset. Data Processing The data has been processed to: Convert JSON data to CSV format. Remove messages from bots. Filter out messages containing only URLs, mentions, channels or discord emojis. Filter out messages that are not in English using a FastText language identification model. Data Fields The CSV files in… See the full description on the dataset page: https://huggingface.co/datasets/ManBib/Discord-Unveiled-Extracted.texttoken-classification100M<n<1B2 likes626 downloads1y agoHugging Face18mookiezi /Discord-Dialogues Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format. This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words. Nomic Atlas Map Features Mixed single and multi-turn exchanges Human-only dialogues (no bots) Filtered for ToS and harmful contentLinks… See the full description on the dataset page: https://huggingface.co/datasets/mookiezi/Discord-Dialogues.tabular1M<n<10M22 likes592 downloads1y agoHugging Face19VisionXLab /DisciplineGen-1Mtabular1M<n<10M5 likes516 downloads3mo agoHugging Face20dischargesum /discharge_target Dataset Card for "discharge_target" More Information needed tabular10K<n<100K0 likes490 downloads3y agoHugging Face21jpwahle /dblp-discovery-dataset Dataset Card for DBLP Discovery Dataset (D3) Dataset Summary DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.tabularother1M<n<10M4 likes477 downloads1y agoHugging Face22disco-eth /cineaudiosynth CineAudioSynth Synthetic cinematic audio for source separation. 453 scenes, ~22.8 h, 48 kHz / 16-bit / stereo WAV. Each data/scene_NNNN/ contains: linear/ — additive render: mix.wav is the BIT-EXACT 16-bit sum of the four stems (speech, music, ambience, sfx) — max |mix − Σstems| = 0, verified per scene. Includes gain_envelope.json (sidechain ducking envelopes). release/ — mastered render of the same scene (compression/limiting/loudness on the mix bus; intentionally… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/cineaudiosynth.audio1K<n<10K0 likes435 downloads3mo agoHugging Face23marcov /discovery_discovery_promptsourcetext1M<n<10M0 likes409 downloads2y agoHugging Face24disco-eth /cineaudiodb CineAudioDB A real-world evaluation set for cinematic audio source separation. CineAudioDB contains real film/animation productions with ground-truth stems (dialogue, music, sfx) for evaluating cinematic source-separation models. Unlike synthetic datasets that sum stems linearly, real productions are mixed with a non-linear mastering chain (compression, limiting, sidechain ducking, reverb), so the released stems do not generally sum to the mastered mix. To support fair… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/cineaudiodb.audioaudio-to-audion<1K2 likes380 downloads3mo agoHugging Face25dougalldeepmind /2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment. field value experiment LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.textn<1K0 likes351 downloads25d agoHugging Face26neuralworm /stable-diffusion-discord-promptsstable-diffusion-discord-prompts All messages from dreambot from all dream-[1-50] channels in stable-diffusion discord source: https://github.com/bartman081523/stable-diffusion-discord-prompts text1M<n<10M29 likes302 downloads4y agoHugging Face27SaisExperiments /Discord-Unveiled-Compressed .hf-sanitized.hf-sanitized-uCWd6SwyNH8FCkRETeRYS .container { --bg-primary: #0d0511; --bg-secondary: #1a0f1f; --bg-tertiary: #2d1b35; --bg-card: #3d2847; --text-primary: #fef7ff; --text-secondary: #f0d9ff; --text-muted: #c084fc; --pink-soft: #fce7f3; --pink-medium: #f9a8d4; --pink-bright: #ec4899; --pink-hot: #e91e63; --pink-neon: #ff1493; --purple-soft: #e879f9; --purple-bright: #c026d3; --purple-deep: #7c3aed; --border-glow: #f472b6; --shadow-pink: rgba(244, 114, 182, 0.4);… See the full description on the dataset page: https://huggingface.co/datasets/SaisExperiments/Discord-Unveiled-Compressed.tabularn<1K28 likes285 downloads1y agoHugging Face28laion /LAION-DISCO-12MThe LAION-DISCO-12M dataset contains 12M links to music on YouTube, inspired by the methodology of DISCO-10M. It contains song metadata (song_id, title, artist_names, artist_ids, album_name, album_id, isExplicit, views, duration) and YouTube URL, pointing to the original song on the public web. It does not contain any original audio samples and is thus an index dataset. Starting from an initial seed list of artists, we can discover new artists by recursively exploring the artists listed in the… See the full description on the dataset page: https://huggingface.co/datasets/laion/LAION-DISCO-12M.text10M<n<100M55 likes279 downloads3mo agoHugging Face29disco-eth /AgentsNet AgentsNet This repository contains the graph instances used in the AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs paper. AgentsNet is a new benchmark for multi-agent reasoning, designed to measure the ability of multi-agent systems to collaboratively form strategies for problem-solving, self-organization, and effective communication given a network topology. It draws inspiration from classical problems in distributed systems and graph theory. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AgentsNet.tabulargraph-mln<1K2 likes265 downloads1y agoHugging Face30Flmc /DISC-Med-SFTThis is a repository containing a subset of the DISC-Med-SFT Dataset. Check DISC-MedLLM for more information. textquestion-answering100K<n<1M98 likes260 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.