datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.unified-agent-trajectories
Unified Benchmark Agent Trajectories
Dataset release: v2.1.1 (2026-09-18)Record format: unified-agent-sft-v1
A growing collection of benchmark agent execution trajectories converted into one
transparent, multimodal, tool-aware representation. These are complete recorded benchmark
runs—not ordinary chat transcripts—including benchmark tasks, model reasoning and answers,
tool calls, tool observations, runtime status, and benchmark scores when available. The
directory layout is… See the full description on the dataset page: https://huggingface.co/datasets/ChrisDing1105/unified-agent-trajectories.Draw-and-Understand
🎨 Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level comprehension and limit interaction to textual instructions, thereby constraining their flexibility in usage and depth of response. Therefore, we introduce the… See the full description on the dataset page: https://huggingface.co/datasets/Afeng-x/Draw-and-Understand.UniMM-Chat
Dataset Card for UniMM-Chat
Dataset Summary
UniMM-Chat dataset is an open-source, knowledge-intensive, and multi-round multimodal dialogue data powered by GPT-3.5, which consists of 1.1M diverse instructions.
UniMM-Chat leverages complementary annotations from different VL datasets and employs GPT-3.5 to generate multi-turn dialogues corresponding to each image, resulting in 117,238 dialogues, with an average of 9.89 turns per dialogue.
A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/Yirany/UniMM-Chat.una-fraza-al-diya
Una fraza al diya
Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total.
Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/
Citation
If you use this dataset, please cite:
Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.paleo-hebrew-seals-unambiguous
PaleoHebrew-Seals Real Benchmark (Unambiguous Subset)
This repository hosts the real benchmark part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions from photographs.
Why this dataset is needed
Paleo-Hebrew seal inscriptions are difficult for standard OCR systems: the signs are sparse, shallow, frequently worn, and embedded in irregular seal impressions captured under uncontrolled lighting and viewpoint changes.… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-unambiguous.ogiri-bokete-unsloth-vlm
Japanese Bokete Ogiri — Unsloth VLM format
YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。
各JSONLレコードは「1画像 + 1回答」です。
{
"messages": [
{"role": "user", "content": [
{"type": "image", "image": "images/124469.jpg"},
{"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"}
]},
{"role": "assistant", "content": [
{"type": "text", "text": "..."}
]}
]
}
Files
train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/untitled84/Nemotron-Personas-Korea.Uni-SafeBench
Uni-SafeBench
Content Warning
Uni-SafeBench contains harmful, and unsafe textual and visual content. The dataset is intended only for academic and non-commercial research on the safety evaluation of unified multimodal large models.
The dataset should not be used to generate harmful content, improve harmful instruction following, bypass safety mechanisms, or deploy unsafe generation systems.
Dataset Description
Uni-SafeBench is a safety evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Hades2002/Uni-SafeBench.
