CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lego-X /LegoFlow-SWE LegoFlow-SWE · 5,000 verified Harbor SWE tasks and two GLM-5.2 trajectory releases GitHub · Docs · Blog · HuggingFace · LegoX LegoFlow-SWE 5,000 verified Harbor SWE tasks mined by LegoFlow Curator, shipped in original and anti-hack prompt versions, plus two GLM-5.2 trajectory releases under OpenHands SDK and OpenCode, totaling 9,767 trajectories. Release Count What it is tasks/ 5,000 Original prompts tasks-anti-hack/ 5,000 Same task IDs and task files, with… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/LegoFlow-SWE.imagetext-generation1K<n<10K5 likes2.7k downloads8d agoHugging Face02xupy21 /ICPC_Data ICPC World Finals — a discriminative subset, with model traces 24 ICPC World Finals problems (2021–2025), together with the full transcripts of an LLM attempting each of them three times under simulated contest rules. Selection The model Every run in this dataset comes from: nvidia/Nemotron-Cascade-2-30B-A3B The partitions Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.imagetext-generationn<1K0 likes1.1k downloads4d agoHugging Face03Afeng-x /Draw-and-Understand 🎨 Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level comprehension and limit interaction to textual instructions, thereby constraining their flexibility in usage and depth of response. Therefore, we introduce the… See the full description on the dataset page: https://huggingface.co/datasets/Afeng-x/Draw-and-Understand.imagetext-generation8 likes849 downloads10mo agoHugging Face04MixEval /MixEval-X 🚀 Project Page | 📜 arXiv | 👨‍💻 Github | 🏆 Leaderboard | 📝 blog | 🤗 HF Paper | 𝕏 Twitter MixEval-X encompasses eight input-output modality combinations and can be further extended. Its data points reflect real-world task distributions. The last grid presents the scores of frontier organizations’ flagship models on MixEval-X, normalized to a 0-100 scale, with MMG tasks using win rates instead of Elo. Section C of the paper presents example data samples and model responses.… See the full description on the dataset page: https://huggingface.co/datasets/MixEval/MixEval-X.audioimage-to-text1K<n<10K10 likes258 downloads2y agoHugging Face05ferjorosa /bne-hemeroteca-ocr-xix BNE Hemeroteca OCR Dataset (19th Century) This dataset provides full-text OCR and high-resolution page images for 19th-century Spanish publications sourced from the Biblioteca Nacional de España (BNE) - Hemeroteca Digital. It consists of over 830,000 pages from approximately 40,000 documents, totaling roughly 800 million text tokens. Note on Temporal Coverage: Although the selected publications originated in the 19th century, some long-running titles extend into the early 20th… See the full description on the dataset page: https://huggingface.co/datasets/ferjorosa/bne-hemeroteca-ocr-xix.imageimage-to-text100K<n<1M2 likes258 downloads9mo agoHugging Face06xilinghuiye /ChartInt ChartInt ChartInt is a multimodal chart dataset for chart reconstruction, chart editing, style transfer, interaction editing, and data-update tasks. The Hugging Face release is packaged as a datasets-compatible Parquet dataset so that the Dataset Viewer can display rows, text/code fields, and chart screenshots directly. Dataset Structure The dataset contains 2,905 rows in the train split. Task Rows Description native_reconstruction 556 Reconstruct chart code… See the full description on the dataset page: https://huggingface.co/datasets/xilinghuiye/ChartInt.imageimage-to-text1K<n<10K1 likes228 downloads5mo agoHugging Face07XDUImageLab /SandThink SandThink Dataset (v1.0) SandThink 是一个专为具身智能 (Embodied AI) 任务设计的大规模指令微调与偏好对齐数据集。该数据集通过结构化的 Chain-of-Thought (CoT) 推理过程,显著提升了 Vision-Language-Action (VLA) 模型在复杂环境下的任务拆解、路径规划和动作执行能力。 📊 数据集概览 (Dataset Summary) 本数据集包含三个核心组件,总计约 37,000 条高质量数据: 文件名 类型 规模 说明 Mars_CoT.jsonl SFT / CoT ~17,000 条 针对特定复杂环境(如类地/火星沙盘)的高级推理数据。 iter_DPO_data.jsonl DPO / RLHF ~16,000 条 包含 Chosen 和 Rejected 的偏好对,用于模型对齐与优化。 train_CoT.jsonl SFT / CoT ~4,000 条 基础任务的思维链训练数据,涵盖避障、导航等指令。… See the full description on the dataset page: https://huggingface.co/datasets/XDUImageLab/SandThink.imageroboticsn<1K1 likes188 downloads7mo agoHugging Face08xp0123 /GeoBenchimagetext-generation1K<n<10K10 likes119 downloads1y agoHugging Face09xasayi /FLAWS FLAWS: Faults Localization Across Writing in Science FLAWS is a benchmark for evaluating error identification and localization in scientific papers. It currently consists of 713 paper–error examples, including: 265 unique papers with one error inserted using GPT-5 (in ALL_OPENAI.tar.gz) 448 unique papers with one error inserted using Gemini 2.5 Pro (in ALL_GEMINI.tar.gz) The dataset is generated using a systematic, autonomous framework that produces paper–error examples and… See the full description on the dataset page: https://huggingface.co/datasets/xasayi/FLAWS.documentsummarization1 likes78 downloads5mo agoHugging Face10xlr8harder /trellismark-qwen3-4b TrellisMark Qwen3-4B confirmation corpus This is the frozen English confirmation corpus for TrellisMark, an experimental many-user AI-text watermark. It includes exact generated text and token IDs, unwatermarked Qwen controls, public-key detector evidence, the public research key, independent encoder vectors, and the result reports used for the reader-facing curves. The standalone implementation, detector, and reproduction instructions are in the TrellisMark GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b.imagetext-generation100K<n<1M0 likes55 downloads1mo agoHugging Face11xunyoyo /clawkeeper ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers aka The Norton for OpenClaw SAFETY EXFOLIATE! SAFETY EXFOLIATE! Overview ClawKeeper is a comprehensive real-time security framework designed for autonomous agent systems such as OpenClaw. It provides unified protection through three complementary approaches: Skill-based Protection — Operates at the instruction… See the full description on the dataset page: https://huggingface.co/datasets/xunyoyo/clawkeeper.imagetext-classificationn<1K0 likes51 downloads6mo agoHugging Face12baconnier /deepsynth-en-xsum DeepSynth - XSum BBC News Summarization Dataset Description BBC news articles with single-sentence summaries. Focused on extreme summarization where the summary is a single sentence capturing the essence of the article. This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and processed through a frozen 380M parameter visual encoder… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-en-xsum.imagesummarization100K<n<1M0 likes40 downloads11mo agoHugging Face13xaddh /multimodal-privacy Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework Recent advances in multi-modal Large Language Models (M-LLMs) have demonstrated a powerful ability to synthesize implicit information from disparate sources, including images and text. These resourceful data from social media also introduce a significant and underexplored privacy risk: the inference of sensitive personal attributes from seemingly daily media content. However, the lack of benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/xaddh/multimodal-privacy.imagequestion-answering1K<n<10K1 likes37 downloads11mo agoHugging Face14xcodemind /CC-HARD 🧠 CC-HARD: A Challenging Dataset for Design-to-Code Generation 📄 Paper on arXiv 📄 Paper on ACM CC-HARD is a challenging benchmark dataset introduced in the KDD 2025 paper LaTCoder: Converting Webpage Design to Code with Layout-as-Thought. It was specifically designed to evaluate layout fidelity in webpage design-to-code generation. The dataset consists of 128 webpage screenshots and their corresponding HTML/CSS code, manually curated from the Common Crawl corpus. Unlike prior… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/CC-HARD.imagetext-generationn<1K1 likes34 downloads1y agoHugging Face15xma8 /MedATLAS-BenchThe public subset of MedATLAS-Bench. Unzip public_data.zip into data/ folder. It contains a list of folders with names matching sample_id in test.json. We include public_data_samples.zip as a smaller sample for people to quickly review the quality. Data distributions of the public subset: imageimage-text-to-textn<1K0 likes25 downloads5mo agoHugging Face16Xalitobeirut /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다. Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/Xalitobeirut/Nemotron-Personas-Korea.imagetext-generation1M<n<10M0 likes24 downloads5mo agoHugging Face17lingkoai /Flux-Anime-x-Realistic-Miximagetext-generationn<1K0 likes15 downloads2y agoHugging Face18Anonymousab /Dog-XraygatedPlease provide your name, institution, and a brief explanation of how you intend to use the data to our email: anonymous_ab6@hotmail.com, alongside the automatic verification steps on the website. Access will be granted to researchers with a valid use case. imagetext-generation100M<n<1B0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.