CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danikhan632 /OpenCodeReasoning2text10K<n<100K0 likes1.3k downloads9mo agoHugging Face02danikhan632 /DeepMath-12ktext1K<n<10K0 likes485 downloads9mo agoHugging Face03NebulaeWis /danbooru_2025_recaption 内部暂存数据集 (Internal Temporary Dataset) English This is a temporary dataset for internal use. It might contain: Items being re-processed or corrected (e.g., some images requiring re-tagging using a distributed cluster, needing a convenient data source for it). Data to supplement our internal systems (e.g., if a machine accidentally lost some images and we don't want to re-download everything). Recent updates or experimental data not yet finalized (e.g., the image source… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/danbooru_2025_recaption.tabularn<1K1 likes477 downloads1y agoHugging Face04DanFosing /public-domain-poetry Overview This dataset is a collection of approximately 38,500 poems from https://www.public-domain-poetry.com/. Language The language of this dataset is English. License All data in this dataset is public domain, which means you should be able to use it for anything you want, as long as you aren't breaking any law in the process of doing so. texttext-generation10K<n<100K21 likes472 downloads3y agoHugging Face05DANGDOCAO /GeneratingQuestions HVU_QA HVU_QA is an open-source Vietnamese Question-Context-Answer (QCA) corpus, accompanied by supporting tools, created to facilitate the development of FAQ-style question generation and question answering systems, particularly for low-resource language settings. The dataset was developed by a research team at Hung Vuong University, Phu Tho, Vietnam, led by Dr. Ha Nguyen, Deputy Head of the Department of Engineering Technology. HVU_QA was constructed using a fully automated… See the full description on the dataset page: https://huggingface.co/datasets/DANGDOCAO/GeneratingQuestions.textquestion-answering10K<n<100K16 likes431 downloads2mo agoHugging Face06Jax-dan /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例 HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.texttext-generation10K<n<100K18 likes364 downloads2y agoHugging Face07danbhf /openpi_sim_pick_place SO-101 Pick and Place Dataset (OpenPi Format) This dataset contains 40 episodes of a simulated SO-101 robot performing pick-and-place tasks, converted to OpenPi/RLDS format for use with Physical Intelligence's Pi0/Pi0.5 models. Source Converted from LeRobot dataset: danbhf/sim_pick_place_merged_40ep Format Each episode is stored as an NPZ file containing: Key Shape Type Description observation/state (N, 6) float32 Joint positions (6 DoF)… See the full description on the dataset page: https://huggingface.co/datasets/danbhf/openpi_sim_pick_place.textn<1K0 likes361 downloads9mo agoHugging Face08Danny-1223 /CREBench CREBench CREBench is a benchmark for evaluating large language models (LLMs) on cryptographic binary reverse engineering. Paper: arXiv:2604.03750 Code: wangyu-ovo/CREBench Project Page: CREBench Homepage Dataset Description CREBench measures reverse-engineering performance on cryptographic binaries across four evaluation levels: Level Task L1 Algorithm identification L2 Key (and IV) extraction L3 Wrapper-level code reimplementation L4 Flag… See the full description on the dataset page: https://huggingface.co/datasets/Danny-1223/CREBench.textothern<1K2 likes345 downloads2mo agoHugging Face09kierarkia /danbooru-wiki-2026 danbooru-wiki-2026-04-28 About Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag. This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.tabulartext-classification100K<n<1M11 likes300 downloads5mo agoHugging Face10danhoangg /AviationLLMtext10K<n<100K4 likes279 downloads3y agoHugging Face11PocketDoc /Dans-Logicmaxx-SAT-APtext1K<n<10K0 likes262 downloads2y agoHugging Face12danielrosehill /Whiteboards Whiteboards A small, single-author whiteboard corpus for evaluating vision-language models on handwritten OCR accuracy and for studying pseudotext hallucination — the failure mode where a VLM invents plausible-but-wrong words for ambiguous handwriting. Every image is the same wall-mounted whiteboard, same marker, same author, photographed with a phone. This is deliberate: the dataset exists to measure whether a small number of human-authored ground-truth pairs can improve… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Whiteboards.imageimage-to-textn<1K1 likes256 downloads5mo agoHugging Face13danystar /RetailBanking-Conversations Dataset Description RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field. The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.texttext-generation1K<n<10K0 likes213 downloads6d agoHugging Face14Danasoumoh /phase2tabular1K<n<10K0 likes207 downloads2y agoHugging Face15danielrmarques /bench-marques bench-marques — local LLM inference measurements Every number here was measured on one machine, with the method written down and the refutations kept. This dataset is the source of truth for the results; the harness that produces them lives at danielrmarques/bench-marques. The machine AMD Ryzen AI MAX+ 395 · Radeon 8060S · 128 GB unified LPDDR5X (~205 GB/s effective, 80% of the 256 GB/s theoretical ceiling) · llama.cpp, Vulkan backend. Every run is on AC power… See the full description on the dataset page: https://huggingface.co/datasets/danielrmarques/bench-marques.tabularn<1K1 likes200 downloads17d agoHugging Face16PocketDoc /Dans-Prosemaxx-Opus-Writingtextn<1K1 likes189 downloads2y agoHugging Face17danikhan632 /natural_reasoning_rubricstext1K<n<10K0 likes174 downloads11mo agoHugging Face18Daniel66 /travel_hang_llama3_tttextn<1K0 likes169 downloads2y agoHugging Face19dankeg /ArxivBulkDataset Arxiv Bulk Dataset This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset ArXiv Bulk Dataset This dataset is a bulk fetch of ArXiv articles, based on the official metadata dataset maintained and updated by Cornell https://www.kaggle.com/datasets/Cornell-University/arxiv/data. This dataset was created to provide cross-domain academic training data, with existing datasets being domain-specific, and… See the full description on the dataset page: https://huggingface.co/datasets/dankeg/ArxivBulkDataset.textsummarization1M<n<10M2 likes162 downloads10mo agoHugging Face20CaptionEmporium /anime-caption-danbooru-2021-sfw-5m-hq Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq Dataset Summary This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated. Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.textimage-to-text1M<n<10M29 likes152 downloads2y agoHugging Face21DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes146 downloads2y agoHugging Face22PocketDoc /Dans-Benchmaxx-COTtext10K<n<100K0 likes145 downloads2y agoHugging Face23danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes142 downloads10mo agoHugging Face24PocketDoc /Dans-Taskmaxxtext100K<n<1M4 likes135 downloads2y agoHugging Face25danikhan632 /OpenMathReasoning2text1K<n<10K0 likes131 downloads9mo agoHugging Face26parteeksj /scientific_papers_DANCERtext100K<n<1M1 likes124 downloads3y agoHugging Face27danny2507 /attention-uq-800q-colab Attention/UQ 800-question Colab bundle A deterministic 200-question subset for each of MultiModalQA, WebQA, HotpotQA, and TAT-QA. See manifest.json for exact upstream sources, hashes, counts, and the explicitly constructed WebQA distractor setting. imagequestion-answeringn<1K0 likes122 downloads20d agoHugging Face28luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes114 downloads3mo agoHugging Face29Dannalily /MontageLie MontageLie: Information Alignment Evaluation Benchmark To investigate this vulnerability, we introduce MontageLie, a novel benchmark designed to test the limitations of current information alignment evaluators. Drawing inspiration from the cinematic concept of montage, which creates new meaning by rearranging real scenes in novel sequences, MontageLie constructs "montage-style lies": deceptive texts composed entirely of truthful statements, deliberately reordered to imply… See the full description on the dataset page: https://huggingface.co/datasets/Dannalily/MontageLie.text1K<n<10K0 likes108 downloads1y agoHugging Face30Jax-dan /zhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora. Download You can download the latest Chinese Wikipedia dump from the following link: Chinese Wikipedia Dump English Wikipedia Dump (For reference) Extraction After you download the dump, you can extract the data using the following commands: # install wikiextractor pip install wikiextractor # extract the data wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2 Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.textfill-mask1M<n<10M0 likes107 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.