datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LegoFlow-SWE
LegoFlow-SWE · 5,000 verified Harbor SWE tasks and two GLM-5.2 trajectory releases
GitHub · Docs · Blog · HuggingFace · LegoX
LegoFlow-SWE
5,000 verified Harbor SWE tasks mined by LegoFlow Curator, shipped in original and anti-hack prompt versions, plus two GLM-5.2 trajectory releases under OpenHands SDK and OpenCode, totaling 9,767 trajectories.
Release
Count
What it is
tasks/
5,000
Original prompts
tasks-anti-hack/
5,000
Same task IDs and task files, with… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/LegoFlow-SWE.ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with the full transcripts of an
LLM attempting each of them three times under simulated contest rules.
Selection
The model
Every run in this dataset comes from:
nvidia/Nemotron-Cascade-2-30B-A3B
The partitions
Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then
placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.Draw-and-Understand
🎨 Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level comprehension and limit interaction to textual instructions, thereby constraining their flexibility in usage and depth of response. Therefore, we introduce the… See the full description on the dataset page: https://huggingface.co/datasets/Afeng-x/Draw-and-Understand.MixEval-X
🚀 Project Page | 📜 arXiv | 👨💻 Github | 🏆 Leaderboard | 📝 blog | 🤗 HF Paper | 𝕏 Twitter
MixEval-X encompasses eight input-output modality combinations and can be further extended. Its data points reflect real-world task distributions. The last grid presents the scores of frontier organizations’ flagship models on MixEval-X, normalized to a 0-100 scale, with MMG tasks using win rates instead of Elo. Section C of the paper presents example data samples and model responses.… See the full description on the dataset page: https://huggingface.co/datasets/MixEval/MixEval-X.bne-hemeroteca-ocr-xix
BNE Hemeroteca OCR Dataset (19th Century)
This dataset provides full-text OCR and high-resolution page images for 19th-century Spanish publications sourced from the Biblioteca Nacional de España (BNE) - Hemeroteca Digital. It consists of over 830,000 pages from approximately 40,000 documents, totaling roughly 800 million text tokens.
Note on Temporal Coverage: Although the selected publications originated in the 19th century, some long-running titles extend into the early 20th… See the full description on the dataset page: https://huggingface.co/datasets/ferjorosa/bne-hemeroteca-ocr-xix.ChartInt
ChartInt
ChartInt is a multimodal chart dataset for chart reconstruction, chart editing, style transfer, interaction editing, and data-update tasks. The Hugging Face release is packaged as a datasets-compatible Parquet dataset so that the Dataset Viewer can display rows, text/code fields, and chart screenshots directly.
Dataset Structure
The dataset contains 2,905 rows in the train split.
Task
Rows
Description
native_reconstruction
556
Reconstruct chart code… See the full description on the dataset page: https://huggingface.co/datasets/xilinghuiye/ChartInt.SandThink
SandThink Dataset (v1.0)
SandThink 是一个专为具身智能 (Embodied AI) 任务设计的大规模指令微调与偏好对齐数据集。该数据集通过结构化的 Chain-of-Thought (CoT) 推理过程,显著提升了 Vision-Language-Action (VLA) 模型在复杂环境下的任务拆解、路径规划和动作执行能力。
📊 数据集概览 (Dataset Summary)
本数据集包含三个核心组件,总计约 37,000 条高质量数据:
文件名
类型
规模
说明
Mars_CoT.jsonl
SFT / CoT
~17,000 条
针对特定复杂环境(如类地/火星沙盘)的高级推理数据。
iter_DPO_data.jsonl
DPO / RLHF
~16,000 条
包含 Chosen 和 Rejected 的偏好对,用于模型对齐与优化。
train_CoT.jsonl
SFT / CoT
~4,000 条
基础任务的思维链训练数据,涵盖避障、导航等指令。… See the full description on the dataset page: https://huggingface.co/datasets/XDUImageLab/SandThink.GeoBenchFLAWS
FLAWS: Faults Localization Across Writing in Science
FLAWS is a benchmark for evaluating error identification and localization in scientific papers. It currently consists of 713 paper–error examples, including:
265 unique papers with one error inserted using GPT-5 (in ALL_OPENAI.tar.gz)
448 unique papers with one error inserted using Gemini 2.5 Pro (in ALL_GEMINI.tar.gz)
The dataset is generated using a systematic, autonomous framework that produces paper–error examples and… See the full description on the dataset page: https://huggingface.co/datasets/xasayi/FLAWS.trellismark-qwen3-4b
TrellisMark Qwen3-4B confirmation corpus
This is the frozen English confirmation corpus for
TrellisMark, an experimental
many-user AI-text watermark. It includes exact generated text and token IDs,
unwatermarked Qwen controls, public-key detector evidence, the public research
key, independent encoder vectors, and the result reports used for the
reader-facing curves. The standalone implementation, detector, and
reproduction instructions are in the
TrellisMark GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/trellismark-qwen3-4b.clawkeeper
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
aka The Norton for OpenClaw
SAFETY EXFOLIATE! SAFETY EXFOLIATE!
Overview
ClawKeeper is a comprehensive real-time security framework designed for autonomous agent systems such as OpenClaw. It provides unified protection through three complementary approaches:
Skill-based Protection — Operates at the instruction… See the full description on the dataset page: https://huggingface.co/datasets/xunyoyo/clawkeeper.deepsynth-en-xsum
DeepSynth - XSum BBC News Summarization
Dataset Description
BBC news articles with single-sentence summaries. Focused on extreme summarization where the summary is
a single sentence capturing the essence of the article.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and processed through a frozen 380M parameter visual encoder… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-en-xsum.multimodal-privacy
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
Recent advances in multi-modal Large Language Models (M-LLMs) have demonstrated a powerful ability to synthesize implicit information from disparate sources, including images and text. These resourceful data from social media also introduce a significant and underexplored privacy risk: the inference of sensitive personal attributes from seemingly daily media content. However, the lack of benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/xaddh/multimodal-privacy.CC-HARD
🧠 CC-HARD: A Challenging Dataset for Design-to-Code Generation
📄 Paper on arXiv
📄 Paper on ACM
CC-HARD is a challenging benchmark dataset introduced in the KDD 2025 paper LaTCoder: Converting Webpage Design to Code with Layout-as-Thought.
It was specifically designed to evaluate layout fidelity in webpage design-to-code generation.
The dataset consists of 128 webpage screenshots and their corresponding HTML/CSS code, manually curated from the Common Crawl corpus. Unlike prior… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/CC-HARD.MedATLAS-BenchThe public subset of MedATLAS-Bench.
Unzip public_data.zip into data/ folder.
It contains a list of folders with names matching sample_id in test.json.
We include public_data_samples.zip as a smaller sample for people to quickly review the quality.
Data distributions of the public subset:
Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/Xalitobeirut/Nemotron-Personas-Korea.Flux-Anime-x-Realistic-MixDog-XrayPlease provide your name, institution, and a brief explanation of how you intend to use the data to our email: anonymous_ab6@hotmail.com, alongside the automatic verification steps on the website.
Access will be granted to researchers with a valid use case.
