CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes15k downloads1y agoHugging Face02AlphaDojo /dojo_main_income Languages: 简体中文 · English dojo_main_income — Revenue Breakdown Overview Segment-level main business revenue from listed companies, by industry, product, and region, with amounts and mix ratios. Corresponds to “main business by segment” notes in filings. Files File Description data.parquet Full revenue breakdown detail Key Fields Field Description symbol Stock symbol security_name Company name… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_main_income.tabularn<1K0 likes14k downloads21d agoHugging Face03CohereLabs /include-base-44 INCLUDE-base (44 languages) Dataset Description Paper: http://arxiv.org/abs/2411.19799 Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-base-44.textmultiple-choice10K<n<100K51 likes12k downloads1y agoHugging Face04scikit-learn /adult-census-income Adult Census Income Dataset The following was retrieved from UCI machine learning repository. This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics). A set of reasonably clean records was extracted using the following conditions: ((AAGE>16) && (AGI>100) && (AFNLWGT>1) && (HRSWK>0)). The prediction task is to determine whether a person makes over $50K a year. Description of fnlwgt (final weight)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/adult-census-income.tabular10K<n<100K9 likes4.6k downloads4y agoHugging Face05inclusionAI /ZwZ-RL-VQA ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use. 📖 Overview The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive: Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V)… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZwZ-RL-VQA.text100K<n<1M17 likes4.5k downloads4mo agoHugging Face06inclusionAI /FinFIRST FinFIRST: Financial Information Retrieval, Sourcing and Traceability Released alongside Ling-3.0-flash-Fin, FinFIRST is an open benchmark for evaluating whether financial search agents can produce answers that are not only correct, but also supported by authoritative, timely, and verifiable evidence. It was developed by Ant Group, with professional support from the investment banking team at China International Capital Corporation Limited (CICC). Financial research requires more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinFIRST.documentquestion-answeringn<1K16 likes4k downloads20d agoHugging Face07inclusionAI /FinixDocBench FinixDocBench Language: English | 中文 This repository contains a compliance-reviewed public subset of FinixDocBench, the financial-domain document parsing benchmark introduced in the technical report "FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks". The benchmark focuses on document parsing conditions that are common in real financial workflows but underrepresented in saturated clean-document benchmarks: digitally native insurance clauses, noisy… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/FinixDocBench.imageimage-to-textn<1K14 likes2.6k downloads29d agoHugging Face08CohereLabs /include-lite-44 INCLUDE-lite (44 languages) Dataset Description Paper: http://arxiv.org/abs/2411.19799 Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 11,095 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/include-lite-44.textmultiple-choice10K<n<100K16 likes2.6k downloads1y agoHugging Face09birkhoffg /folktables-acs-income Dataset Card for "folktables-acs-income" More Information needed tabulartabular-classification1M<n<10M1 likes2.3k downloads3y agoHugging Face10inclusionAI /VenusBench-CAPTCHA VenusBench-CAPTCHA: A Real-World CAPTCHA Screenshot–Action Benchmark for GUI Agents Evaluation Code: https://github.com/inclusionAI/UI-Venus/tree/VenusBench-CAPTCHA Introduction CAPTCHA solving is a practical challenge for multimodal GUI agents because it requires more than isolated visual recognition. An agent must understand the challenge instruction, identify the relevant interface region, recognize or reason about the visual target, ground the result… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/VenusBench-CAPTCHA.imageimage-text-to-textn<1K6 likes2.2k downloads21d agoHugging Face11Dingdong-Inc /FreshRetailNet-50K FreshRetailNet-50K Dataset Overview FreshRetailNet-50K is the first large-scale benchmark for censored demand estimation in the fresh retail domain, incorporating approximately 20% organically occurring stockout data. It comprises 50,000 store-product 90-day time series of detailed hourly sales data from 898 stores in 18 major cities, encompassing 865 perishable SKUs with meticulous stockout event annotations. The hourly stock status records unique to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dingdong-Inc/FreshRetailNet-50K.tabulartime-series-forecasting1M<n<10M28 likes1.8k downloads9mo agoHugging Face12SevatarOoi /Nine-Bus-Load-Increase-Eventtext100M<n<1B0 likes1.8k downloads1y agoHugging Face13inclusionAI /Ling-Coder-SFT 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.texttext-generation1M<n<10M45 likes1.4k downloads1y agoHugging Face14incredible45 /Gutenberg-BookCorpus-Cleaned-Data-English Gutenberg-BookCorpus-Cleaned-Data-English This dataset is been cleaned and preprocessed using Gutenberg_English_Preprocessor class method (given below) from preference Kaggle dataset 75,000+ Gutenberg Books and Metadata 2025. This dataset is only specialisation for english contented with rights as "Public domain in the USA" hence you can free used it anywhere. Following reference metadata of Gutenberg is also available and downloaded it using following CLI command below :- pip… See the full description on the dataset page: https://huggingface.co/datasets/incredible45/Gutenberg-BookCorpus-Cleaned-Data-English.text10K<n<100K15 likes1.3k downloads1y agoHugging Face15inclusionAI /AudioMCQ [ICLR 2026] AudioMCQ: Audio Multiple-Choice Question Dataset Also the official repository for the paper "Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models" News [2026.04] Update on MMSU Metric of released models: Based on community feedback, we identified a flaw in our evaluation script that artificially inflated the MMSU scores of our released models by ignoring sequence order. We sincerely apologize for… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AudioMCQ.text100K<n<1M15 likes813 downloads5mo agoHugging Face16Simuletic /CCTV_Incident_Dataset_Fall_Lying_Down_Detection Overview This is an open-source synthetic dataset for Computer Vision (CV) tasks, specifically designed for Fall Detection, Pose Estimation, and Incident Monitoring from overhead CCTV perspectives. Unlike standard object detection datasets, this dataset includes Keypoints (Pose) annotations. This enables models to understand human posture and accurately distinguish between standing and fallen individuals. 🚀 Need more data? This is a sample dataset by Simuletic. We provide… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/CCTV_Incident_Dataset_Fall_Lying_Down_Detection.imagen<1K4 likes787 downloads9mo agoHugging Face17Emulated-Inc /forum-competition-math-training-pool Forum competition mathematics training pool Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.texttext-generation100K<n<1M0 likes736 downloads11d agoHugging Face18Emulated-Inc /olympiad-math-training-pool Olympiad mathematics training pool Public olympiad and competition mathematics, four datasets gathered at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 229052 rows across four folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 225822 rows, every row labelled with the dataset it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/olympiad-math-training-pool.texttext-generation100K<n<1M0 likes668 downloads11d agoHugging Face19Eventual-Inc /sample-parquetSample Parquet dataset for testing purposes textn<1K0 likes661 downloads1y agoHugging Face20Salesforce /FaithEval-inconsistent-v1.0 FaithEval FaithEval is a new and comprehensive benchmark dedicated to evaluating contextual faithfulness in LLMs across three diverse tasks: unanswerable, inconsistent, and counterfactual contexts. [Paper] FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows", ICLR 2025, https://arxiv.org/abs/2410.03727 [Code and Detailed Instructions] https://github.com/SalesforceAIResearch/FaithEval Disclaimer and Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/FaithEval-inconsistent-v1.0.text1K<n<10K3 likes641 downloads2y agoHugging Face21inclusionAI /Ling-Coder-SyntheticQA 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA.texttext-generation10M<n<100M17 likes628 downloads1y agoHugging Face22akankshanc /inception-v1-microscope-data Inception V1 Microscope Data This dataset powers the Inception V1 Microscope, an interactive interface for exploring visual features learned by individual neurons in Inception V1. It combines two complementary interpretability views: Activation maximization: one synthesized visualization optimized to strongly activate each neuron. Top dataset examples: the ten ImageNet examples producing the strongest recorded activations for each neuron, paired with crops associated with the… See the full description on the dataset page: https://huggingface.co/datasets/akankshanc/inception-v1-microscope-data.image10K<n<100K1 likes616 downloads2mo agoHugging Face23inclusionAI /SWE-CARE SWE-CARE: A Comprehensiveness-aware Benchmark for Code Review Evaluation Dataset Description SWE-CARE (Software Engineering - Comprehensive Analysis and Review Evaluation) is a comprehensiveness-aware benchmark for evaluating Large Language Models (LLMs) on repository-level code review tasks. The dataset features real-world code review scenarios from popular open-source Python and Java repositories, with comprehensive metadata and… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/SWE-CARE.texttext-generation1K<n<10K8 likes581 downloads11mo agoHugging Face24gemmozero /ai-agent-security-incidents AI Agent Security Incident Database v0.1 A structured, machine-readable database of 1392 confirmed AI agent security incidents, collected and classified automatically. What is this? Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it. This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.tabulartext-classification1K<n<10K1 likes544 downloads46m agoHugging Face25ai4bharat /INCLUDE Dataset Card for INCLUDE Dataset Summary This dataset contains all videos in the INCLUDE dataset. As huggingface does not support video uploads at this time, the HF dataset contains metadata about each video such as the parent class, the video class, the path to the video and whether its a part of the INCLUDE-50 dataset (use include_50==True to get only include_50 videos). The videos themselves can be downloaded from Zenodo using the provided bash script.… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/INCLUDE.text1K<n<10K5 likes523 downloads2y agoHugging Face26jlh /uci-adult-income Dataset Card for "uci-adult-income" More Information needed tabular10K<n<100K0 likes521 downloads3y agoHugging Face27zhush /incantation-elden-ring-scenes Incantation Elden Ring Combat Captions Paper | Project page | GitHub Preview subset. This repository is a public preview and reference subset of the Incantation dataset. It documents the data format, annotation style, and initial training material used by the project. It should not be interpreted as the final dataset, the complete benchmark, or the full data scale used by the paper. This dataset contains manually collected Elden Ring combat clips paired with structured… See the full description on the dataset page: https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes.textvideo-classification1K<n<10K3 likes511 downloads2mo agoHugging Face28meghana /adult_income_datasettabular10K<n<100K0 likes498 downloads4y agoHugging Face29hheiden /PubChem-124M-SMILES-SELFIES-InChI-IUPAC PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format. Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource: SMILES: Raw and RDKit-Canonicalized. SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.texttext-generation100M<n<1B8 likes496 downloads8mo agoHugging Face30inclusionAI /ZoomBench ZoomBench: A Fine-Grained Multimodal Perception Benchmark 📃 Paper | 🏠 Project | 🤗 Models Overview ZoomBench is a challenging benchmark designed to evaluate the fine-grained multimodal perception capabilities of Multimodal Large Language Models (MLLMs). It specifically targets scenarios where decisive visual evidence is small, subtle, or easily overwhelmed by global context — situations that demand "zooming-level" perception from a single full image. It is… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/ZoomBench.imageimage-text-to-textn<1K8 likes484 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.