CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inclusionAI /OpenAoE-2000h Open-AoE — Egocentric Hand Manipulation Dataset Release Roadmap Tier Duration Status nano ~3 h ✅ Released tiny ~100 h ✅ Released full 2000 h 🚧 Uploading Release notes 2026-07-30: Removed samples flagged in PR #1 for camera-intrinsics vs. video-resolution mismatches. 2026-07-31: Uploaded ~323h of data. 2026-08-12: Uploaded ~694h of data. 2026-09-03: Uploaded ~189h of data. Additional data for the full ~2000h release is still… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/OpenAoE-2000h.36 likes503k downloads17d agoHugging Face02orionweller /reddit_mds_incremental0 likes105k downloads2y agoHugging Face03inclusionAI /ConceptEdit-12M ConceptEdit: Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision &nbsp;&nbsp;&nbsp; ConceptEdit-12M is a large-scale image editing dataset. Each sample is stored as a triplet: a source image, an edited image, a JSON metadata file describing the edit instruction, edit category, relative image paths, and VQA-style quality checks. The dataset is packaged as multiple .tar shards. All paths inside the tar files and JSON files are relative paths; no… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ConceptEdit-12M.image-to-image10M<n<100M47 likes61k downloads28d agoHugging Face04Inception3D /GenFusion_Training_Datavideo10K<n<100K1 likes34k downloads1y agoHugging Face05orionweller /cc_en_middle_mds_incremental0 likes17k downloads2y agoHugging Face06skysight-inc /HackerNewsContains all Hacker News posts through April 15th, 2025. 0 likes16k downloads1y agoHugging Face07inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes15k downloads1y agoHugging Face08AlphaDojo /dojo_main_income Languages: 简体中文 · English dojo_main_income — Revenue Breakdown Overview Segment-level main business revenue from listed companies, by industry, product, and region, with amounts and mix ratios. Corresponds to “main business by segment” notes in filings. Files File Description data.parquet Full revenue breakdown detail Key Fields Field Description symbol Stock symbol security_name Company name… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_main_income.tabularn<1K0 likes14k downloads20d agoHugging Face09CohereLabs /include-base-44 INCLUDE-base (44 languages) Dataset Description Paper: http://arxiv.org/abs/2411.19799 Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-base-44.textmultiple-choice10K<n<100K51 likes13k downloads1y agoHugging Face10incognitolm /USA-Map-Tiles Dataset Description The dataset is a set of map tiles for the various states of the USA. License: Open Data Commons Open Database License (ODbL) v1.0 Dataset Sources OpenStreetMap.org Geofabrik Download Server Bounding Boxes per State Dataset Structure Organized into folders by state Structure of graphml files: Top of files have a list of keys/ids that correspond to the properties of the segment: maxspeed: speed limit oneway: if it is a… See the full description on the dataset page: https://huggingface.co/datasets/incognitolm/USA-Map-Tiles.0 likes10k downloads8mo agoHugging Face11inclusionAI /VenusBench-GD VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks Project Page: https://ui-venus.github.io/VenusBench-GD/ Introduction GUI grounding is a critical component in building capable GUI agents. However, existing grounding benchmarks suffer from significant limitations: they either provide insufficient data volume and narrow domain coverage, or focus excessively on a single platform and require highly specialized domain… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-GD.imageimage-text-to-textn<1K14 likes9.5k downloads9mo agoHugging Face12rob9m /inc0 likes5.7k downloads1y agoHugging Face13scikit-learn /adult-census-income Adult Census Income Dataset The following was retrieved from UCI machine learning repository. This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics). A set of reasonably clean records was extracted using the following conditions: ((AAGE>16) && (AGI>100) && (AFNLWGT>1) && (HRSWK>0)). The prediction task is to determine whether a person makes over $50K a year. Description of fnlwgt (final weight)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/adult-census-income.tabular10K<n<100K9 likes4.7k downloads4y agoHugging Face14orionweller /tulu_flan_mds_incremental-tokens0 likes4.6k downloads2y agoHugging Face15inclusionAI /ZwZ-RL-VQA ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use. 📖 Overview The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive: Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V)… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZwZ-RL-VQA.text100K<n<1M17 likes4.5k downloads4mo agoHugging Face16inclusionAI /FinFIRST FinFIRST: Financial Information Retrieval, Sourcing and Traceability Released alongside Ling-3.0-flash-Fin, FinFIRST is an open benchmark for evaluating whether financial search agents can produce answers that are not only correct, but also supported by authoritative, timely, and verifiable evidence. It was developed by Ant Group, with professional support from the investment banking team at China International Capital Corporation Limited (CICC). Financial research requires more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinFIRST.documentquestion-answeringn<1K15 likes3.9k downloads19d agoHugging Face17orionweller /refinedweb_mds_incremental0 likes3.4k downloads2y agoHugging Face18orionweller /c4_mds_incremental0 likes3.4k downloads2y agoHugging Face19orionweller /tulu_flan_mds_incremental0 likes3.1k downloads2y agoHugging Face20orionweller /cc_news_mds_incremental-tokens0 likes2.8k downloads2y agoHugging Face21Emulated-Inc /ogb-full-original OGB full original archives Public, byte-for-byte mirror of 17 official Open Graph Benchmark (OGB) and OGB Large-Scale Challenge archive downloads used by the Emulated-Inc graph benchmark environment. Original ZIP archives are stored under archives//. Each dataset directory includes metadata.json with the authoritative source URL, exact byte size, SHA-256 digest, and repository archive path. Archives were transferred directly from the official SNAP/DGL hosts through ephemeral… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/ogb-full-original.0 likes2.7k downloads22d agoHugging Face22zimingzh /In-cylinder_Flow_Field1 likes2.6k downloads3y agoHugging Face23birkhoffg /folktables-acs-income Dataset Card for "folktables-acs-income" More Information needed tabulartabular-classification1M<n<10M1 likes2.6k downloads3y agoHugging Face24inclusionAI /FinixDocBench FinixDocBench Language: English | 中文 This repository contains a compliance-reviewed public subset of FinixDocBench, the financial-domain document parsing benchmark introduced in the technical report "FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks". The benchmark focuses on document parsing conditions that are common in real financial workflows but underrepresented in saturated clean-document benchmarks: digitally native insurance clauses, noisy… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinixDocBench.imageimage-to-textn<1K14 likes2.5k downloads28d agoHugging Face25CohereLabs /include-lite-44 INCLUDE-lite (44 languages) Dataset Description Paper: http://arxiv.org/abs/2411.19799 Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.textmultiple-choice10K<n<100K16 likes2.5k downloads1y agoHugging Face26inclusionAI /VenusBench-CAPTCHA VenusBench-CAPTCHA: A Real-World CAPTCHA Screenshot–Action Benchmark for GUI Agents Evaluation Code: https://github.com/inclusionAI/UI-Venus/tree/VenusBench-CAPTCHA Introduction CAPTCHA solving is a practical challenge for multimodal GUI agents because it requires more than isolated visual recognition. An agent must understand the challenge instruction, identify the relevant interface region, recognize or reason about the visual target, ground the result… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-CAPTCHA.imageimage-text-to-textn<1K6 likes2.1k downloads20d agoHugging Face27SevatarOoi /Nine-Bus-Load-Increase-Eventtext100M<n<1B0 likes1.8k downloads1y agoHugging Face28Dingdong-Inc /FreshRetailNet-50K FreshRetailNet-50K Dataset Overview FreshRetailNet-50K is the first large-scale benchmark for censored demand estimation in the fresh retail domain, incorporating approximately 20% organically occurring stockout data. It comprises 50,000 store-product 90-day time series of detailed hourly sales data from 898 stores in 18 major cities, encompassing 865 perishable SKUs with meticulous stockout event annotations. The hourly stock status records unique to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dingdong-Inc/FreshRetailNet-50K.tabulartime-series-forecasting1M<n<10M27 likes1.7k downloads9mo agoHugging Face29orionweller /cc_en_head_mds_incremental0 likes1.6k downloads2y agoHugging Face30inclusionAI /Ling-Coder-SFT 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.texttext-generation1M<n<10M45 likes1.4k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.