CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01disco-eth /EuroSpeech EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.audioautomatic-speech-recognition10M<n<100M99 likes32k downloads5mo agoHugging Face02multilingual-discourse-hub /disrpt Disrpt is a multilingual, multi-framework unified discourse analysis benchmark. It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages. ⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB. To load these datasets, run the following: pip install disrpt-utils Then from disrpt_utils import load_dataset corpora_paths={ # ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.text100K<n<1M3 likes12k downloads1y agoHugging Face03disco-eth /AIMEfrom datasets import load_dataset dataset = load_dataset('disco-eth/AIME') AIME: AI Music Evaluation Dataset The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo. The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset. The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset. The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.audio1K<n<10K9 likes2.4k downloads2y agoHugging Face04dataforge-labs /equity-perp-price-discovery Equity and pre-IPO perpetual prices Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets. Contents Table Record perpetual_mark_and_index_prices An instrument's mark price, index price and basis at an observation time Using the data market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.tabulartime-series-forecasting10K<n<100K1 likes2.2k downloads3h agoHugging Face05disco-eth /EuroSpeech-24kHz EuroSpeech 24 kHz Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. Dataset Summary Languages: 22 European languages (see detailed… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech-24kHz.audioautomatic-speech-recognition10M<n<100M3 likes1.6k downloads5mo agoHugging Face06google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.2k downloads3y agoHugging Face07disco-eth /GlobalDISCO GlobalDISCO GlobalDISCO is a large-scale dataset consisting of 73k music tracks generated by state-of-the-art commercial generative music models, along with paired links to 93k reference tracks in LAION-DISCO-12M. The dataset spans 147 languages and includes musical style prompts extracted from MusicBrainz and Wikipedia. The dataset is globally balanced, representing musical styles from artists across 79 countries and five continents. It is aimed to support the research community in… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/GlobalDISCO.audioaudio-classification10K<n<100K1 likes1.1k downloads10mo agoHugging Face08CompassioninMachineLearning /caml-animal-discourse-2020-present Reddit Animal-Discourse Corpus — CLEANED (2020–present) Submissions and comments from animal-relevant subreddits, gathered via PullPush.io, covering January 2020 to the present. Built as part of research on AI-mediated value lock-in in human animal-welfare discourse. Coverage Subreddit Submissions Comments Date range (submissions) r/AnimalRights 15,719 34,686 2020-01-01 → 2025-05-19 r/AntiVegan 17,252 182,890 2020-01-01 → 2025-05-19 r/AskVegans 4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.tabulartext-classification1M<n<10M0 likes853 downloads3mo agoHugging Face09sileod /discovery Dataset Card for Discovery Dataset Summary Discourse marker prediction with 174 markers Supported Tasks and Leaderboards [More Information Needed] Languages English Dataset Structure input : sentence1, sentence2, label: marker originally between sentence1 and sentence2 Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits Train/Val/Test Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/sileod/discovery.texttext-classification1M<n<10M8 likes721 downloads2y agoHugging Face10mookiezi /Discord-Dialogues Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format. This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words. Nomic Atlas Map Features Mixed single and multi-turn exchanges Human-only dialogues (no bots) Filtered for ToS and harmful contentLinks… See the full description on the dataset page: https://huggingface.co/datasets/mookiezi/Discord-Dialogues.tabular1M<n<10M22 likes597 downloads1y agoHugging Face11VisionXLab /DisciplineGen-1Mtabular1M<n<10M5 likes492 downloads3mo agoHugging Face12dischargesum /discharge_target Dataset Card for "discharge_target" More Information needed tabular10K<n<100K0 likes464 downloads3y agoHugging Face13marcov /discovery_discovery_promptsourcetext1M<n<10M0 likes422 downloads2y agoHugging Face14Miyalinsky /discard_tiletabular10K<n<100K0 likes320 downloads10mo agoHugging Face15AivexRoboticsGroup /omy_f3m_multi_Disconnect-0This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m_multi", "total_episodes": 100, "total_frames": 61744, "total_tasks": 1, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AivexRoboticsGroup/omy_f3m_multi_Disconnect-0.tabularrobotics10K<n<100K0 likes305 downloads8mo agoHugging Face16laion /LAION-DISCO-12MThe LAION-DISCO-12M dataset contains 12M links to music on YouTube, inspired by the methodology of DISCO-10M. It contains song metadata (song_id, title, artist_names, artist_ids, album_name, album_id, isExplicit, views, duration) and YouTube URL, pointing to the original song on the public web. It does not contain any original audio samples and is thus an index dataset. Starting from an initial seed list of artists, we can discover new artists by recursively exploring the artists listed in the… See the full description on the dataset page: https://huggingface.co/datasets/laion/LAION-DISCO-12M.text10M<n<100M55 likes276 downloads3mo agoHugging Face17DIS-CO /MovieTection Dataset Description 🎬 The MovieTection dataset is a benchmark designed for detecting pretraining data in Large Vision-Language Models (VLMs). It serves as a resource for analyzing model exposure to Copyrighted Visual Content ©️. Paper: DIS-CO: Discovering Copyrighted Content in VLMs Training Data Direct Use 🖥️ The dataset is designed for image/caption-based question-answering, where models predict the movie title given a frame or its corresponding textual… See the full description on the dataset page: https://huggingface.co/datasets/DIS-CO/MovieTection.imagevisual-question-answering10K<n<100K8 likes258 downloads1y agoHugging Face18CompassioninMachineLearning /reddit-control-discourse-2016-present-pretau Reddit Control Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records (75.1%) from 1,030,104 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.tabular1M<n<10M0 likes256 downloads3mo agoHugging Face19somosnlp-hackathon-2023 /informes_discriminacion_gitana Resumen del dataset Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.imagetext-classification1K<n<10K8 likes236 downloads3y agoHugging Face20jatinganhotra /SWE-bench_Verified-discriminative SWE-bench Verified Discriminative Subsets Dataset Description This dataset contains discriminative subsets of SWE-bench Verified designed to provide more sensitive evaluation of SWE-agent capabilities. As top-performing agents achieve 73%+ on the full benchmark, these subsets focus on truly challenging problems to better discriminate between cutting-edge systems. Key Features 4 discriminative splits targeting different evaluation needs 335 carefully selected… See the full description on the dataset page: https://huggingface.co/datasets/jatinganhotra/SWE-bench_Verified-discriminative.textn<1K0 likes236 downloads1y agoHugging Face21CompassioninMachineLearning /reddit-animal-discourse-2016-present-pretau Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records (78.0%) from 343,756 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.tabular1M<n<10M0 likes233 downloads3mo agoHugging Face22disco-eth /AgentsNet AgentsNet This repository contains the graph instances used in the AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs paper. AgentsNet is a new benchmark for multi-agent reasoning, designed to measure the ability of multi-agent systems to collaboratively form strategies for problem-solving, self-organization, and effective communication given a network topology. It draws inspiration from classical problems in distributed systems and graph theory. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AgentsNet.tabulargraph-mln<1K2 likes218 downloads1y agoHugging Face23arubique /disco-model-outputs DISCO model outputs Tabular release of per-model, per-item correctness and answer scores used to train and evaluate DISCO: Diversifying Sample Condensation for Efficient Model Evaluation. The paper studies cheap benchmark performance prediction from a small subset of evaluation items; this dataset supplies the raw harness-style outputs for MMLU (57 subjects), HellaSwag, Winogrande, ARC, and related tasks from the Open LLM Leaderboard ecosystem. Paper Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/arubique/disco-model-outputs.tabularother10M<n<100M0 likes217 downloads6mo agoHugging Face24geodesic-research /eval-deployment-discriminationtabular100K<n<1M0 likes183 downloads3mo agoHugging Face25google-research-datasets /coarse_discourse Dataset Card for "coarse_discourse" Dataset Summary A large corpus of discourse annotations and relations on ~10K forum threads. We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.texttext-classification100K<n<1M6 likes182 downloads3y agoHugging Face26agoulah /ontario-lobbying-disclosure-graph Ontario Lobbying & MPP Disclosure Graph A structured, entity-resolved projection of three public Ontario government records sources, exported as flat, documented parquet tables: Ontario Lobbyist Registry (Office of the Integrity Commissioner of Ontario, lobbyist.oico.on.ca) — lobbyist registrations: who is registered to lobby, for which client, about what, aimed at which offices. MPP Public Disclosure Statements (Office of the Integrity Commissioner of Ontario, PDS) — annual… See the full description on the dataset page: https://huggingface.co/datasets/agoulah/ontario-lobbying-disclosure-graph.tabular100K<n<1M0 likes165 downloads4mo agoHugging Face27CompassioninMachineLearning /reddit-animal-discourse-2016-present Reddit Animal Discourse (2016-present) Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed. Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.tabular1M<n<10M0 likes160 downloads3mo agoHugging Face28sachithgunasekara /phased-self-discover-mistral-structured-5-shot-bbh-evaltext1K<n<10K0 likes159 downloads2y agoHugging Face29KeWangRobotics /panda_pick_cube_demos_sim_discrete_newThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 30, "total_frames": 3570, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/KeWangRobotics/panda_pick_cube_demos_sim_discrete_new.tabularrobotics1K<n<10K0 likes154 downloads1y agoHugging Face30ajacquet /franka-discrete-reach4absThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "franka_panda", "total_episodes": 371, "total_frames": 27053, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:371" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ajacquet/franka-discrete-reach4abs.imagerobotics10K<n<100K0 likes148 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.