CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mandarjoshi /trivia_qa Dataset Card for "trivia_qa" Dataset Summary TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/mandarjoshi/trivia_qa.textquestion-answering100K<n<1M206 likes154k downloads3y agoHugging Face02Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M197 likes16k downloads27d agoHugging Face03manu /project_gutenberg Dataset Card for "Project Gutenberg" Project Gutenberg is a library of over 70,000 free eBooks, hosted at https://www.gutenberg.org/. All examples correspond to a single book, and contain a header and a footer of a few lines (delimited by a *** Start of *** and *** End of *** tags). Usage from datasets import load_dataset ds = load_dataset("manu/project_gutenberg", split="fr", streaming=True) print(next(iter(ds))) License Full license is available here:… See the full description on the dataset page: https://huggingface.co/datasets/manu/project_gutenberg.texttext-generation10K<n<100K74 likes8.6k downloads3y agoHugging Face04TIGER-Lab /Mantis-Instruct Mantis-Instruct Paper | Website | Github | Models | Demo Summaries Mantis-Instruct is a fully text-image interleaved multimodal instruction tuning dataset, containing 721K examples from 14 subsets and covering multi-image skills including co-reference, reasoning, comparing, temporal understanding. It's been used to train Mantis Model families Mantis-Instruct has a total of 721K instances, consisting of 14 subsets to cover all the multi-image skills. Among the… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Instruct.text100K<n<1M41 likes8k downloads2y agoHugging Face05Manusagents /arvo-cybergym-2000 ARVO CyberGym-format 2000-task dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It combines jm-rt/arvo-cybergym-1000 with the second 1000-task small-target ARVO batch built outside the original CyberGym set. text1K<n<10K0 likes5.9k downloads2mo agoHugging Face06hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.4k downloads4y agoHugging Face07lighteval /sacrebleu_manualtext100K<n<1M0 likes5k downloads1y agoHugging Face08dsixteen /Niji_1_Man-metaimage10K<n<100K0 likes4.6k downloads3d agoHugging Face09mannycooper /document-review-data Document Review Data Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package. Current Title Extraction Dataset Surface Canonical prefix: datasets/title_extraction/ Effective datasets: datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/ datasets/title_extraction/evaluation/real_device_280_v1/ datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/ The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.tabular100K<n<1M3 likes4.4k downloads15d agoHugging Face10manavtabbly /hindi_audio_dataset_testaudion<1K0 likes4.4k downloads11mo agoHugging Face11Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.40546875 Action score: 0.475 Valid samples: 320/320 tabularn<1K0 likes4.3k downloads15d agoHugging Face12Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes4.3k downloads15d agoHugging Face13Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.39921875 Action score: 0.44375 Valid samples: 320/320 tabularn<1K0 likes4.3k downloads15d agoHugging Face14Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38359375 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes4.3k downloads15d agoHugging Face15manu /code-20b Dataset Card for "code_20b2" More Information needed text10M<n<100M4 likes3.3k downloads3y agoHugging Face16physicl /kitchen-workspace-understanding-safe-manipulation Kitchen Workspace Understanding & Safe Manipulation Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.imagen<1K0 likes2.9k downloads3mo agoHugging Face17hugging-science /mmu_manga mmu_manga HATS Catalog Collection This is the collection of HATS catalogs representing mmu_manga. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be installed via… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_manga.tabular10K<n<100K0 likes2.7k downloads3mo agoHugging Face18manjot007 /Indian-Laws Dataset Card for Indian Laws This is a comprehensive collection of primary legal documents pertinent to the Indian legal system. It is designed to serve as a foundational resource for supervised fine-tuning (SFT) to make language models, particularly those focused on legal applications tailored for Indian law. text10K<n<100K0 likes2.5k downloads5mo agoHugging Face19Zhixi666 /art_manip_data3d10M<n<100M1 likes2.5k downloads4mo agoHugging Face20manu /code_20b Dataset Card for "code_20b" More Information needed text10M<n<100M1 likes2.3k downloads3y agoHugging Face21MahtaFetrat /Mana-TTS ManaTTS-Persian-Speech-Dataset ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality audio (sampled at 44.1 kHz). Released under the permissive CC-0 license, this dataset is freely usable for both educational and commercial purposes. Collected from Nasl-e-Mana magazine, the dataset covers a diverse range of topics, making it ideal for training robust text-to-speech (TTS) models. The release includes a fully transparent… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/Mana-TTS.tabular10K<n<100K29 likes2.3k downloads1y agoHugging Face22manikandan18ramalingam /agentic-ai-options-resultstextn<1K1 likes2k downloads1h agoHugging Face23mangopy /ToolRet-Queries🔧 Retrieving useful tools from a large-scale toolset is an important step for Large language model (LLMs) in tool learning. This project (ToolRet) contribute to (i) the first comprehensive tool retrieval benchmark to systematically evaluate existing information retrieval (IR) models on tool retrieval tasks; and (ii) a large-scale training dataset to optimize the expertise of IR models on this tool retrieval task. See the official Github for more details. A concrete example for our evaluation… See the full description on the dataset page: https://huggingface.co/datasets/mangopy/ToolRet-Queries.text1K<n<10K7 likes1.9k downloads2y agoHugging Face24manifoldlabs /Infinity-Instruct Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] (would be released soon) The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and… See the full description on the dataset page: https://huggingface.co/datasets/manifoldlabs/Infinity-Instruct.texttext-generation10M<n<100M5 likes1.9k downloads2y agoHugging Face25Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M6 likes1.9k downloads2mo agoHugging Face26AssemblyWorld /partnet-manualpa PartNet-ManualPA 6,871 objects with original shape-frame part meshes and 70,923 rendered assembly PNGs. One object per row; all assets are embedded in Parquet. No source-specific dataloader is needed. from datasets import load_dataset objects = load_dataset("AssemblyWorld/partnet-manualpa", split="all") sample = objects[0] Requires datasets >= 5.0.1 and Pillow. Pin revision to a release commit for reproducibility. For local loading use the prepared directory instead of the Hub… See the full description on the dataset page: https://huggingface.co/datasets/AssemblyWorld/partnet-manualpa.image1K<n<10K0 likes1.7k downloads19d agoHugging Face27RLinf /rlt-maniskill-PegInsertionSide-v1-400-succ RLT ManiSkill Joint Dataset Summary rlt_maniskill_joint is a LeRobot-style dataset for joint-control Robot Learning Token (RLT) training on the ManiSkill peg insertion task. It is designed for the RLinf + OpenPI pi05_rlt_joint pipeline and is used in three stages: OpenPI supervised fine-tuning (SFT) base policy training RLT Stage 1 RL-token training RLT Stage 2 online RL initialization and normalization The dataset corresponds to the ManiSkill task:… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/rlt-maniskill-PegInsertionSide-v1-400-succ.image10K<n<100K0 likes1.7k downloads3mo agoHugging Face28BangumiBase /manariafriends Bangumi Image Base of Manaria Friends This is the image base of bangumi Manaria Friends, we detected 45 characters, 2199 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/manariafriends.image1K<n<10K0 likes1.6k downloads2y agoHugging Face29manycore-research /SpatialLM-Testset SpatialLM Testset Project page | Paper | Code We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos. Folder Structure Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.3dn<1K60 likes1.6k downloads1y agoHugging Face30myendless /ManipDreamer3D_data Dataset of paper ManipDreamer3D This repository contains the dataset for the paper ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory. Data Structure The dataset is organized as follows: manipdreamer3d_data/ ├── 000000/ # data processed from bridge-v1 ├── 000000_v2/ # data processed from bridge-v2 │ ├── data.json # contains gripper state, camera params, etc. │ ├── depth_0000.png… See the full description on the dataset page: https://huggingface.co/datasets/myendless/ManipDreamer3D_data.imagen<1K1 likes1.5k downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.