CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01danikhan632 /OpenCodeReasoning2text10K<n<100K0 likes1.3k downloads9mo agoHugging Face02danikhan632 /DeepMath-12ktext1K<n<10K0 likes515 downloads9mo agoHugging Face03DanFosing /public-domain-poetry Overview This dataset is a collection of approximately 38,500 poems from https://www.public-domain-poetry.com/. Language The language of this dataset is English. License All data in this dataset is public domain, which means you should be able to use it for anything you want, as long as you aren't breaking any law in the process of doing so. texttext-generation10K<n<100K21 likes480 downloads3y agoHugging Face04DANGDOCAO /GeneratingQuestions HVU_QA HVU_QA is an open-source Vietnamese Question-Context-Answer (QCA) corpus, accompanied by supporting tools, created to facilitate the development of FAQ-style question generation and question answering systems, particularly for low-resource language settings. The dataset was developed by a research team at Hung Vuong University, Phu Tho, Vietnam, led by Dr. Ha Nguyen, Deputy Head of the Department of Engineering Technology. HVU_QA was constructed using a fully automated… See the full description on the dataset page: https://huggingface.co/datasets/DANGDOCAO/GeneratingQuestions.textquestion-answering10K<n<100K16 likes437 downloads2mo agoHugging Face05Jax-dan /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例 HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.texttext-generation10K<n<100K18 likes369 downloads2y agoHugging Face06danbhf /openpi_sim_pick_place SO-101 Pick and Place Dataset (OpenPi Format) This dataset contains 40 episodes of a simulated SO-101 robot performing pick-and-place tasks, converted to OpenPi/RLDS format for use with Physical Intelligence's Pi0/Pi0.5 models. Source Converted from LeRobot dataset: danbhf/sim_pick_place_merged_40ep Format Each episode is stored as an NPZ file containing: Key Shape Type Description observation/state (N, 6) float32 Joint positions (6 DoF)… See the full description on the dataset page: https://huggingface.co/datasets/danbhf/openpi_sim_pick_place.textn<1K0 likes363 downloads9mo agoHugging Face07Danny-1223 /CREBench CREBench CREBench is a benchmark for evaluating large language models (LLMs) on cryptographic binary reverse engineering. Paper: arXiv:2604.03750 Code: wangyu-ovo/CREBench Project Page: CREBench Homepage Dataset Description CREBench measures reverse-engineering performance on cryptographic binaries across four evaluation levels: Level Task L1 Algorithm identification L2 Key (and IV) extraction L3 Wrapper-level code reimplementation L4 Flag… See the full description on the dataset page: https://huggingface.co/datasets/Danny-1223/CREBench.textothern<1K2 likes357 downloads2mo agoHugging Face08kierarkia /danbooru-wiki-2026 danbooru-wiki-2026-04-28 About Wiki pages about the danbooru tags on danbooru.donmai.us. The wiki contains the description of each tag. This dataset was collected from the wiki data using danbooru API and then merged with tags data (category_name & post_count). No entries were removed, no matter how many posts they have. The resulting set was manually filtered. Originally some of the lists, tag_group:, about:, help:, howto:, etc, pages were labled as "general"… See the full description on the dataset page: https://huggingface.co/datasets/kierarkia/danbooru-wiki-2026.tabulartext-classification100K<n<1M11 likes283 downloads5mo agoHugging Face09danhoangg /AviationLLMtext10K<n<100K4 likes279 downloads3y agoHugging Face10PocketDoc /Dans-Logicmaxx-SAT-APtext1K<n<10K0 likes261 downloads2y agoHugging Face11danielrosehill /Whiteboards Whiteboards A small, single-author whiteboard corpus for evaluating vision-language models on handwritten OCR accuracy and for studying pseudotext hallucination — the failure mode where a VLM invents plausible-but-wrong words for ambiguous handwriting. Every image is the same wall-mounted whiteboard, same marker, same author, photographed with a phone. This is deliberate: the dataset exists to measure whether a small number of human-authored ground-truth pairs can improve… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Whiteboards.imageimage-to-textn<1K1 likes261 downloads5mo agoHugging Face12danystar /RetailBanking-Conversations Dataset Description RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field. The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.texttext-generation1K<n<10K0 likes213 downloads7d agoHugging Face13Danasoumoh /phase2tabular1K<n<10K0 likes211 downloads2y agoHugging Face14danielrmarques /bench-marques bench-marques — local LLM inference measurements Every number here was measured on one machine, with the method written down and the refutations kept. This dataset is the source of truth for the results; the harness that produces them lives at danielrmarques/bench-marques. The machine AMD Ryzen AI MAX+ 395 · Radeon 8060S · 128 GB unified LPDDR5X (~205 GB/s effective, 80% of the 256 GB/s theoretical ceiling) · llama.cpp, Vulkan backend. Every run is on AC power… See the full description on the dataset page: https://huggingface.co/datasets/danielrmarques/bench-marques.tabularn<1K1 likes204 downloads18d agoHugging Face15PocketDoc /Dans-Prosemaxx-Opus-Writingtextn<1K1 likes187 downloads2y agoHugging Face16danikhan632 /natural_reasoning_rubricstext1K<n<10K0 likes174 downloads11mo agoHugging Face17Daniel66 /travel_hang_llama3_tttextn<1K0 likes169 downloads2y agoHugging Face18dankeg /ArxivBulkDataset Arxiv Bulk Dataset This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset ArXiv Bulk Dataset This dataset is a bulk fetch of ArXiv articles, based on the official metadata dataset maintained and updated by Cornell https://www.kaggle.com/datasets/Cornell-University/arxiv/data. This dataset was created to provide cross-domain academic training data, with existing datasets being domain-specific, and… See the full description on the dataset page: https://huggingface.co/datasets/dankeg/ArxivBulkDataset.textsummarization1M<n<10M2 likes160 downloads11mo agoHugging Face19DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes148 downloads2y agoHugging Face20PocketDoc /Dans-Benchmaxx-COTtext10K<n<100K0 likes146 downloads2y agoHugging Face21danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes144 downloads10mo agoHugging Face22danikhan632 /OpenMathReasoning2text1K<n<10K0 likes138 downloads9mo agoHugging Face23CaptionEmporium /anime-caption-danbooru-2021-sfw-5m-hq Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq Dataset Summary This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated. Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.textimage-to-text1M<n<10M29 likes136 downloads2y agoHugging Face24luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes136 downloads3mo agoHugging Face25PocketDoc /Dans-Taskmaxxtext100K<n<1M4 likes132 downloads2y agoHugging Face26parteeksj /scientific_papers_DANCERtext100K<n<1M1 likes124 downloads3y agoHugging Face27danny2507 /attention-uq-800q-colab Attention/UQ 800-question Colab bundle A deterministic 200-question subset for each of MultiModalQA, WebQA, HotpotQA, and TAT-QA. See manifest.json for exact upstream sources, hashes, counts, and the explicitly constructed WebQA distractor setting. imagequestion-answeringn<1K0 likes123 downloads21d agoHugging Face28Dannalily /MontageLie MontageLie: Information Alignment Evaluation Benchmark To investigate this vulnerability, we introduce MontageLie, a novel benchmark designed to test the limitations of current information alignment evaluators. Drawing inspiration from the cinematic concept of montage, which creates new meaning by rearranging real scenes in novel sequences, MontageLie constructs "montage-style lies": deceptive texts composed entirely of truthful statements, deliberately reordered to imply… See the full description on the dataset page: https://huggingface.co/datasets/Dannalily/MontageLie.text1K<n<10K0 likes106 downloads1y agoHugging Face29Jax-dan /zhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora. Download You can download the latest Chinese Wikipedia dump from the following link: Chinese Wikipedia Dump English Wikipedia Dump (For reference) Extraction After you download the dump, you can extract the data using the following commands: # install wikiextractor pip install wikiextractor # extract the data wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2 Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.textfill-mask1M<n<10M0 likes100 downloads1y agoHugging Face30Daniel4190 /filtered_models_swe_smithtext1K<n<10K0 likes98 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.