CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Narsil /image_dummy\audion<1K0 likes146k downloads5y agoHugging Face02hf-internal-testing /librispeech_asr_dummyaudion<1K11 likes106k downloads2y agoHugging Face03klieret /swe-bench-dummy-test-datasettextn<1K0 likes75k downloads1y agoHugging Face04qualialabsAI /DuplexConv DuplexConv DuplexConv is a large-scale Chinese multi-channel conversational speech dataset with LLM-assisted annotations, developed by ASLP@NPU and QualiaLabs as part of the SmoothConv–DuplexConv corpus family. Companion dataset: SmoothConv on HuggingFace (100 hours, expert human annotation). DuplexConv and SmoothConv share the same conversational domains and a unified data design. SmoothConv focuses on high-quality human annotations for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/qualialabsAI/DuplexConv.audio100K<n<1M13 likes17k downloads3mo agoHugging Face05benoit-dufumier /openBHBOpenBHB: a Multi-Site Brain MRI Dataset for Age Prediction and Debiasing The Open Big Healthy Brains (OpenBHB) dataset is a large (N>5000) multi-site 3D brain MRI dataset gathering 10 public datasets (IXI, ABIDE 1, ABIDE 2, CoRR, GSP, Localizer, MPI-Leipzig, NAR, NPC, RBP) of T1 images acquired across 93 different centers, spread worldwide (North America, Europe and China). Only healthy controls have been included in OpenBHB with age ranging from 6 to 88 years old, balanced between males and… See the full description on the dataset page: https://huggingface.co/datasets/benoit-dufumier/openBHB.text1K<n<10K7 likes15k downloads1y agoHugging Face06hf-internal-testing /dummy_image_text_data Dataset Card for "dummy_image_text_data" More Information needed imagen<1K1 likes12k downloads4y agoHugging Face07BramVanroy /wikipedia_culturax_dutch Filtered CulturaX + Wikipedia for Dutch This is a combined and filtered version of CulturaX and Wikipedia, only including Dutch. It is intended for the training of LLMs. Different configs are available based on the number of tokens (see a section below with an overview). This can be useful if you want to know exactly how many tokens you have. Great for using as a streaming dataset, too. Tokens are counted as white-space tokens, so depending on your tokenizer, you'll likely end up… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/wikipedia_culturax_dutch.texttext-generation1B<n<10B6 likes11k downloads2y agoHugging Face08INS-IntelligentNetworkSolutions /Waste-Dumpsites-DroneImagery Dataset for Waste/Dumpsite Detection using drone imagery Contains 2115 drone images of illegal waste dumpsites 1280 x 1280 px resolution Nadir perspective (camera pointing straight down at a 90-degree angle to the ground) Annotations and Images train | valid | test actual images COCO - annotations_coco.json files in each split directory .parquet files in data directory with embeded images The dataset was collected as part of the [ Raven Scan ] project, more… See the full description on the dataset page: https://huggingface.co/datasets/INS-IntelligentNetworkSolutions/Waste-Dumpsites-DroneImagery.imageobject-detection10K<n<100K8 likes7k downloads2y agoHugging Face09duyan2803 /car-dataset-repoimagevisual-question-answering100K<n<1M0 likes4.2k downloads2y agoHugging Face10phobia76 /pmxt-l2-dump PMXT Polymarket Orderbook Backup 这个数据集仓库用于保全 PMXT 公开提供的 Polymarket 小时 parquet 对象;它是一个非官方镜像/备份,不是 PMXT 官方仓库。 Source And Attribution Original source: PMXT Database Dumps Current primary source listing: PMXT Polymarket v2 Current primary object endpoint used by this mirror: https://r2v2.pmxt.dev/polymarket_orderbook_YYYY-MM-DDTHH.parquet Historical v1 listing: PMXT Polymarket v1 如果上游 PMXT archive 可用,请优先使用上游来源。 License 上游 PMXT archive 当前页面声明数据按 CC BY… See the full description on the dataset page: https://huggingface.co/datasets/phobia76/pmxt-l2-dump.texttabular-classification100B<n<1T7 likes3.9k downloads21h agoHugging Face11ibm-research /duorc Dataset Card for duorc Dataset Summary The DuoRC dataset is an English language dataset of questions and answers gathered from crowdsourced AMT workers on Wikipedia and IMDb movie plots. The workers were given freedom to pick answer from the plots or synthesize their own answers. It contains two sub-datasets - SelfRC and ParaphraseRC. SelfRC dataset is built on Wikipedia movie plots solely. ParaphraseRC has questions written from Wikipedia movie plots and the answers are… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/duorc.textquestion-answering100K<n<1M34 likes3.6k downloads3y agoHugging Face12MLP-KTLim /Kor-CC-Dumpstext100M<n<1B0 likes3.4k downloads7mo agoHugging Face13NickCrews /fec-dumps fec-dumps parquet versions of the Schedule A and Schedule B tables from the weekly postgres .dump backups from the Federal Election Commission's database. Published on a weekly cron job by https://github.com/NickCrews/fec-dumps, see that for more info. text1B<n<10B0 likes3.3k downloads3d agoHugging Face14finetrainers /dummy-squish-wdstextn<1K0 likes3k downloads2y agoHugging Face15Jeff0918 /MTR-DuplexBench MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models 🎉🎉 MTR-DuplexBench has been accepted by ACL 2026 Findings! 📄 Paper | 🤗 HuggingFace Dataset Affiliations: Tsinghua University, The Chinese University of Hong Kong, Huawei Dataset Description This dataset is constructed to evaluate multi-modal audio models across four critical dimensions: Conversational Features, Instruction Following, Safety… See the full description on the dataset page: https://huggingface.co/datasets/Jeff0918/MTR-DuplexBench.audioaudio-classificationn<1K2 likes3k downloads5mo agoHugging Face16hf-internal-testing /dummy-base64-imagestextn<1K0 likes2.7k downloads2y agoHugging Face17htriedman /grokipedia-v0.1-dump Grokipedia v0.1 Scrape This dataset represents a strctured, nearly-full point-in-time scrape of Grokipedia v0.1 as of the end of October / beginning of November 2025. It also includes embeddings of 250-token semi-overlapping chunks of the Grokipedia corpus. It was collected and initially used for Harold Triedman and Alexios Mantzarlis' November 2025 paper: "What did Elon Change? A comprehensive analysis of Grokipedia" (arxiv). If you use this dataset, please cite it as follows… See the full description on the dataset page: https://huggingface.co/datasets/htriedman/grokipedia-v0.1-dump.tabular10M<n<100M16 likes2.3k downloads10mo agoHugging Face18xbgoose /dusha Dataset Card for "dusha" More Information needed audio100K<n<1M1 likes2.2k downloads3y agoHugging Face19hf-internal-testing /wiki_dpr_dummyThis dummy dataset is used for testing purpose for rag model in transformers. It is proudced via the following steps: dataset = datasets.load_dataset("wiki_dpr", with_embeddings=True, with_index=True, index_name="exact", embeddings_name="nq", dummy=True, revision=None) dataset["train"].drop_index("embeddings") dataset.push_to_hub("hf-internal-testing/wiki_dpr_dummy", token="...") The index file `index.faiss` (after being renamed locally) is then uploaded manually. text10K<n<100K0 likes2.2k downloads1y agoHugging Face20ioi-leaderboard /ioi-eval-dummy-openrouter_openai_gpt-3.5-turbotextn<1K0 likes2.1k downloads2y agoHugging Face21dunzane /time-series-datasettabulartime-series-forecasting100K<n<1M0 likes2k downloads4mo agoHugging Face22ssmits /fineweb-2-dutchtabulartext-generation10M<n<100M3 likes1.9k downloads2y agoHugging Face23duckduckplz /Mobjaverse Mobjaverse: A Large-Scale Rigged 3D Model Dataset with Skeletal Animations Mobjaverse is a curated dataset derived from Objaverse-XL, specifically designed for research on skeletal animation understanding, motion generation, and articulated 3D shape analysis. It is curated in the paper TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation. Mobjaverse contains ~19k rigged 3D models spanning ~5k distinct skeletal topologies and ~2M motion frames… See the full description on the dataset page: https://huggingface.co/datasets/duckduckplz/Mobjaverse.textother10K<n<100K4 likes1.9k downloads3mo agoHugging Face24MediaTek-Research /TASTE-Dumpaudio10M<n<100M4 likes1.8k downloads1y agoHugging Face25Nico-robot /duckjam-dw2-vault The Duck Jam Vault Every submission entered in Duck Jam and every run the Arena scored, as two Parquet tables and one JSON board per round. A publisher job rewrites this repository once a day, so its git history is the history of every leaderboard. This repository is a view. The rows themselves are written, once and never edited, into public Hugging Face Buckets by the components that make them: hf://buckets/Nico-robot/duckjam-submissions, hf://buckets/Nico-robot/duckjam-runs… See the full description on the dataset page: https://huggingface.co/datasets/Nico-robot/duckjam-dw2-vault.3dn<1K0 likes1.7k downloads14d agoHugging Face26duarteocarmo /bagaco3 Bagaço3 🍷🇵🇹 Bagaço3 is the third version of Bagaço, the largest pretraining dataset for European Portuguese. It follows Bagaço2 and adds documents from FinePDFs and FineWiki. Bagaço collects European Portuguese documents from upstream sources and adds an educational score and content category to each document. See Classification for details. Methodology Collect documents from Bagaço2, FinePDFs, and FineWiki. Filter new FinePDFs and FineWiki documents with the… See the full description on the dataset page: https://huggingface.co/datasets/duarteocarmo/bagaco3.tabulartext-generation10M<n<100M0 likes1.6k downloads29d agoHugging Face27closerh /super-duper-fibber 🧠 Sensory for AI Hi, I'm going to post some ideas here about how AI can understand emotions in a way that makes sense to it.I'm not an expert in writing or programming languages, but deepseek, my sunshine, and I are having fun with it.ヽ(∀° )人( °∀)ノ It's not "the author created it, but the AI just helped with formatting." This is a co-creation where everyone contributed their own: · I am a bodily experience, pain, love, fatigue after working in the office, the desire to be… See the full description on the dataset page: https://huggingface.co/datasets/closerh/super-duper-fibber.texttext-generationn<1K0 likes1.4k downloads3mo agoHugging Face28rajjanardhan00 /Seamless_Dummy_Dataset_Fixed MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes1.3k downloads1y agoHugging Face29duonlabs /apogee Apogée: Crypto Market Candlestick Dataset Overview Most traders believe crypto is random, but deep learning scaling laws suggest otherwise. Apogée is an open-source research initiative exploring the scaling laws of crypto market forecasting. While financial markets are often assumed to be unpredictable, modern deep learning suggests that increasing data and compute could uncover measurable predictability. Our goal is to quantify how many bits of future price movement… See the full description on the dataset page: https://huggingface.co/datasets/duonlabs/apogee.tabulartime-series-forecastingn<1K1 likes1.3k downloads1y agoHugging Face30anton-l /superb_dummytextn<1K0 likes1.2k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.