datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sbsfigures
SBSFigures
The official dataset repository for the following paper:
SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images
Risa Shionoda, Kuniaki Saito, Shohei Tanaka, Tosho Hirasawa, Yoshitaka Ushiku, The AAAI-25 Workshop on Document Understanding and Intelligence
Abstract
Building a large-scale figure QA dataset requires a considerable amount of work, from gathering and selecting figures to extracting attributes like text, numbers, and colors… See the full description on the dataset page: https://huggingface.co/datasets/omron-sinicx/sbsfigures.gnss-jamming-spoofing-detection
GNSS Jamming & Spoofing Detection Dataset
A physics-informed synthetic dataset for detecting GPS/GNSS cyber-attacks
(Jamming and Spoofing) from satellite-signal features. Built for the
GNSS Guardian project — Introduction to Data Science final project.
Overview
14,850 samples across 450 scenarios × 33 time-steps each
3 balanced classes: Normal / Jamming / Spoofing (4,950 each)
26 columns: multi-constellation signal features + attack metadata + text descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Omrilevi123/gnss-jamming-spoofing-detection.omr_benchmark
Muse OMR Benchmark
What this is
A small, clean benchmark dataset for OMR (Optical Music Recognition — recognizing music notation from images/PDFs).
It contains 1077 pairs:
a symbolic music score (the “ground truth”, see dataset fields below)
a corresponding PDF rendering with data augmentation applied
All underlying works are Public Domain.
Why it exists
OMR is often evaluated on private or inconsistent datasets. This dataset aims to provide the community… See the full description on the dataset page: https://huggingface.co/datasets/musegroup/omr_benchmark.sets_lego_omr_full
Dataset Card for sets_lego_omr_full
This dataset combines official LEGO sets from LDRAW OMR with metadata from Rebrickable. Each entry contains the full MPD file as a string plus associated metadata such as set number, name, theme, year, and parts count. It is intended for building LLM fine-tuning datasets for LEGO model generation tasks.
Dataset Details
Dataset Sources
OMR files: LDraw OMR Library
Metadata: Rebrickable
MPD file format: LDraw File Format… See the full description on the dataset page: https://huggingface.co/datasets/DylanRiden/sets_lego_omr_full.omr-olimpic
OLiMPiC (OLYMPIAD of Music PrInted-to-Code) — synthetic + scanned
OLiMPiC contains piano accompaniments from the OpenScore Lieder corpus: end-to-end ground truth (Linearized MusicXML) paired with MuseScore-rendered synthetic images (synthetic/) and real IMSLP flatbed scans (scanned/).
scanned/: grand-staff image crops from IMSLP scans with LMX transcriptions — the realistic evaluation split.
synthetic/: rendered pages/systems with LMX transcriptions.
Cite: Torras, Mayer et al.… See the full description on the dataset page: https://huggingface.co/datasets/e7mac/omr-olimpic.DQA_Template_datasetnvidia_math_generated_solution_750Kmfa-vs-sae-2026-webapp-datanvidia-math-vectorizedcar_parts_datasetdl-trm-phase2-codebook-v32
DL-TRM Phase 2 Codebook V32
This standalone dataset contains the Phase 2 discrete Z traces for V32.
The Hugging Face dataset viewer reads data/train.jsonl.
Each row contains:
raw_puzzle_id
route_id
route_local_id
medoid_old_id
selected_stage_a_example_index
cluster_size
vocab_size
z_trace: a length-16 list of discrete Z token IDs
The original PyTorch artifacts remain in the repo:
z_traces.pt
codebook.pt
transition_vq_model.pt
diagnostics.json
z_trace_manifest.json
checkpoints/
lidar_imu_odometryverovio-synth-omr
Verovio Synthetic OMR Benchmark
A page-level Optical Music Recognition (OMR) evaluation benchmark of 1,998
synthetic score images rendered with Verovio,
paired with both **kern (Humdrum) and MusicXML transcriptions.
Released as the synthetic half of the Transcoda evaluation suite. See
btrkeks/transcoda-59M-zeroshot-v1
for the matching model and the project repository for the benchmark runner
(scripts/benchmark/).
Intended Use
Evaluation only. This dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/btrkeks/verovio-synth-omr.dl-trm-phase2-codebook-v16
DL-TRM Phase 2 Codebook V16
This standalone dataset contains the Phase 2 discrete Z traces for V16.
The Hugging Face dataset viewer reads data/train.jsonl.
Each row contains:
raw_puzzle_id
route_id
route_local_id
medoid_old_id
selected_stage_a_example_index
cluster_size
vocab_size
z_trace: a length-16 list of discrete Z token IDs
The original PyTorch artifacts remain in the repo:
z_traces.pt
codebook.pt
transition_vq_model.pt
diagnostics.json
z_trace_manifest.json
checkpoints/
omr-grpo-train
OMR GRPO Train — Math-Image RL Training Set
74,971 rows · 8 sources · images embedded as bytes (self-contained)
This is the RL training parquet used for Group Relative Policy Optimization (GRPO) on math and STEM image-reasoning tasks. It was built by merging and filtering eight public visual-math datasets into a single verl-compatible parquet. Images are embedded directly as PNG bytes — no external downloads required.
The reasoning system prompt is baked into every row's prompt… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/omr-grpo-train.dl-trm-phase2-codebook-v256
DL-TRM Phase 2 Codebook V256
This standalone dataset contains the Phase 2 discrete Z traces for V256.
The Hugging Face dataset viewer reads data/train.jsonl.
Each row contains:
raw_puzzle_id
route_id
route_local_id
medoid_old_id
selected_stage_a_example_index
cluster_size
vocab_size
z_trace: a length-16 list of discrete Z token IDs
The original PyTorch artifacts remain in the repo:
z_traces.pt
codebook.pt
transition_vq_model.pt
diagnostics.json
z_trace_manifest.json
checkpoints/
wiki2023_plus
Overview
The dataset includes the description from Wikipedia and categories of films published in 2023.
This dataset is used to evaluate the ability of LLM to memorize and extract information described in the document.
See "Where is the Answer? An Empirical Study of Positional Bias for Parametric Knowledge Extraction in Language Model (NAACL2025 Long paper)" for how we use this dataset for training and evaluation.
Data Split
film_doc_all.jsonl includes lines of… See the full description on the dataset page: https://huggingface.co/datasets/omron-sinicx/wiki2023_plus.HotelReservationsDataset
Hotel Reservation EDA & Prediction
Overview
This project explores a hotel reservation dataset from Kaggle with the goal of understanding booking behavior
and determining whether it is possible to predict reservation status (canceled or not canceled) based on the available features.
The project focuses on Exploratory Data Analysis (EDA) only — no machine learning model was trained.
Dataset
This dataset includes 36,275 hotel reservations with 19 features… See the full description on the dataset page: https://huggingface.co/datasets/Omrihahami/HotelReservationsDataset.Website_Traffic_and_Engagementdl-trm-phase2-codebook-v128
DL-TRM Phase 2 Codebook V128
This standalone dataset contains the Phase 2 discrete Z traces for V128.
The Hugging Face dataset viewer reads data/train.jsonl.
Each row contains:
raw_puzzle_id
route_id
route_local_id
medoid_old_id
selected_stage_a_example_index
cluster_size
vocab_size
z_trace: a length-16 list of discrete Z token IDs
The original PyTorch artifacts remain in the repo:
z_traces.pt
codebook.pt
transition_vq_model.pt
diagnostics.json
z_trace_manifest.json
checkpoints/
open_math_not_in_trainnvidia_math_512_47KphaseZAlignBench
HalCap-Bench
HalCap-Bench dataset.
Columns
model
image_source
image_name
image_type
sentence_index
caption
annotation
error_type
error_words
agreement_ratio
fleiss_Pi
n_correct
n_incorrect
n_unknown
image_url
image_path_in_repo
Notes
Notes
For COCO/CC12M items, the image is referenced by image_url.
For SD/Imagen/data_generation items, the image file is stored under images/ and referenced by image_path_in_repo.
CharBench
CharBench - Character-level benchmark and analysis suite for LLMs.
CharBench is a large-scale benchmark for studying tokenization and character-level behavior in modern language models.
For complete details on data curation and evaluation, see the paper.
If you have ideas and suggestions to improve charbench feel free to reach out!
uzan dot omri at gmail.com
Usage
from datasets import load_dataset
ds = load_dataset("omriuz/CharBench")
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/omriuz/CharBench.arc-agi-1-zloopvit-traces
ARC-AGI-1 ZLoopViT Traces
Canonical intermediate grid trajectories for all 400 training tasks in
ARC-AGI-1. Each row corresponds to one original official train or test pair;
generated augmentations are not included.
Dataset contents
400 tasks
1,718 trajectories: 1,302 demonstration/train pairs and 416 test pairs
5,059 visible intermediate transitions
Exact final-output validation on every official pair
Important columns:
task_id: official ARC task identifier… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/arc-agi-1-zloopvit-traces.arc-agi-1-ruleloopvit-rules
ARC-AGI-1 RuleLoopViT Rules
This dataset contains one canonical, task-specific English rule for each of the
400 official ARC-AGI-1 training tasks. Rules were inferred only from official
demonstration input/output pairs. Official test inputs, test outputs, and test
traces were excluded from rule authoring.
Each row includes:
a concise standalone core_rule_text;
a five-section full_rule_text;
the corresponding structured sections;
augmentation-aware references for colors… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/arc-agi-1-ruleloopvit-rules.debussy-omr-fullpage-lvl
Debussy OMR – Full-Page Level Dataset
A dataset of 101 paired samples of full-page images from handwritten music scores written by Claude Debussy (French Composer, 1862 - 1918) and their corresponding MusicXML transcriptions.
The manuscript images originate from the collections of the Bibliothèque nationale de France (BnF) and were accessed through the IIIF protocol.
Dataset Description
Property
Value
Samples
101
Documents
37
Image format
JPEG… See the full description on the dataset page: https://huggingface.co/datasets/HugoSchtr/debussy-omr-fullpage-lvl.nvidia_math_47KOMR-forms
Dataset Card for "OMR-forms"
More Information needed
