CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01omron-sinicx /sbsfigures SBSFigures The official dataset repository for the following paper: SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images Risa Shionoda, Kuniaki Saito, Shohei Tanaka, Tosho Hirasawa, Yoshitaka Ushiku, The AAAI-25 Workshop on Document Understanding and Intelligence Abstract Building a large-scale figure QA dataset requires a considerable amount of work, from gathering and selecting figures to extracting attributes like text, numbers, and colors… See the full description on the dataset page: https://huggingface.co/datasets/omron-sinicx/sbsfigures.imagevisual-question-answering1M<n<10M5 likes814 downloads8mo agoHugging Face02Omrilevi123 /gnss-jamming-spoofing-detection GNSS Jamming & Spoofing Detection Dataset A physics-informed synthetic dataset for detecting GPS/GNSS cyber-attacks (Jamming and Spoofing) from satellite-signal features. Built for the GNSS Guardian project — Introduction to Data Science final project. Overview 14,850 samples across 450 scenarios × 33 time-steps each 3 balanced classes: Normal / Jamming / Spoofing (4,950 each) 26 columns: multi-constellation signal features + attack metadata + text descriptions… See the full description on the dataset page: https://huggingface.co/datasets/Omrilevi123/gnss-jamming-spoofing-detection.tabulartabular-classification10K<n<100K0 likes634 downloads2mo agoHugging Face03musegroup /omr_benchmark Muse OMR Benchmark What this is A small, clean benchmark dataset for OMR (Optical Music Recognition — recognizing music notation from images/PDFs). It contains 1077 pairs: a symbolic music score (the “ground truth”, see dataset fields below) a corresponding PDF rendering with data augmentation applied All underlying works are Public Domain. Why it exists OMR is often evaluated on private or inconsistent datasets. This dataset aims to provide the community… See the full description on the dataset page: https://huggingface.co/datasets/musegroup/omr_benchmark.documentimage-feature-extractionn<1K1 likes552 downloads9mo agoHugging Face04DylanRiden /sets_lego_omr_full Dataset Card for sets_lego_omr_full This dataset combines official LEGO sets from LDRAW OMR with metadata from Rebrickable. Each entry contains the full MPD file as a string plus associated metadata such as set number, name, theme, year, and parts count. It is intended for building LLM fine-tuning datasets for LEGO model generation tasks. Dataset Details Dataset Sources OMR files: LDraw OMR Library Metadata: Rebrickable MPD file format: LDraw File Format… See the full description on the dataset page: https://huggingface.co/datasets/DylanRiden/sets_lego_omr_full.texttext-to-3d1K<n<10K2 likes263 downloads8mo agoHugging Face05e7mac /omr-olimpic OLiMPiC (OLYMPIAD of Music PrInted-to-Code) — synthetic + scanned OLiMPiC contains piano accompaniments from the OpenScore Lieder corpus: end-to-end ground truth (Linearized MusicXML) paired with MuseScore-rendered synthetic images (synthetic/) and real IMSLP flatbed scans (scanned/). scanned/: grand-staff image crops from IMSLP scans with LMX transcriptions — the realistic evaluation split. synthetic/: rendered pages/systems with LMX transcriptions. Cite: Torras, Mayer et al.… See the full description on the dataset page: https://huggingface.co/datasets/e7mac/omr-olimpic.imageimage-to-text10K<n<100K0 likes243 downloads29d agoHugging Face06omribenh /DQA_Template_datasetimagen<1K0 likes200 downloads2y agoHugging Face07omrisap /nvidia_math_generated_solution_750Ktext100K<n<1M0 likes148 downloads6mo agoHugging Face08omrifahn /mfa-vs-sae-2026-webapp-datatabularn<1K0 likes145 downloads8mo agoHugging Face09omrisap /nvidia-math-vectorizedtabular10K<n<100K0 likes132 downloads6mo agoHugging Face10omrastogi /car_parts_datasetimage1K<n<10K2 likes86 downloads2y agoHugging Face11omrisap /dl-trm-phase2-codebook-v32 DL-TRM Phase 2 Codebook V32 This standalone dataset contains the Phase 2 discrete Z traces for V32. The Hugging Face dataset viewer reads data/train.jsonl. Each row contains: raw_puzzle_id route_id route_local_id medoid_old_id selected_stage_a_example_index cluster_size vocab_size z_trace: a length-16 list of discrete Z token IDs The original PyTorch artifacts remain in the repo: z_traces.pt codebook.pt transition_vq_model.pt diagnostics.json z_trace_manifest.json checkpoints/ tabular10K<n<100K0 likes85 downloads4mo agoHugging Face12omrastogi /lidar_imu_odometrytext10K<n<100K0 likes65 downloads5mo agoHugging Face13btrkeks /verovio-synth-omr Verovio Synthetic OMR Benchmark A page-level Optical Music Recognition (OMR) evaluation benchmark of 1,998 synthetic score images rendered with Verovio, paired with both **kern (Humdrum) and MusicXML transcriptions. Released as the synthetic half of the Transcoda evaluation suite. See btrkeks/transcoda-59M-zeroshot-v1 for the matching model and the project repository for the benchmark runner (scripts/benchmark/). Intended Use Evaluation only. This dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/btrkeks/verovio-synth-omr.imageimage-to-text1K<n<10K1 likes65 downloads4mo agoHugging Face14omrisap /dl-trm-phase2-codebook-v16 DL-TRM Phase 2 Codebook V16 This standalone dataset contains the Phase 2 discrete Z traces for V16. The Hugging Face dataset viewer reads data/train.jsonl. Each row contains: raw_puzzle_id route_id route_local_id medoid_old_id selected_stage_a_example_index cluster_size vocab_size z_trace: a length-16 list of discrete Z token IDs The original PyTorch artifacts remain in the repo: z_traces.pt codebook.pt transition_vq_model.pt diagnostics.json z_trace_manifest.json checkpoints/ tabular10K<n<100K0 likes61 downloads4mo agoHugging Face15ngqtrung /omr-grpo-train OMR GRPO Train — Math-Image RL Training Set 74,971 rows · 8 sources · images embedded as bytes (self-contained) This is the RL training parquet used for Group Relative Policy Optimization (GRPO) on math and STEM image-reasoning tasks. It was built by merging and filtering eight public visual-math datasets into a single verl-compatible parquet. Images are embedded directly as PNG bytes — no external downloads required. The reasoning system prompt is baked into every row's prompt… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/omr-grpo-train.textimage-text-to-text10K<n<100K0 likes54 downloads3mo agoHugging Face16omrisap /dl-trm-phase2-codebook-v256 DL-TRM Phase 2 Codebook V256 This standalone dataset contains the Phase 2 discrete Z traces for V256. The Hugging Face dataset viewer reads data/train.jsonl. Each row contains: raw_puzzle_id route_id route_local_id medoid_old_id selected_stage_a_example_index cluster_size vocab_size z_trace: a length-16 list of discrete Z token IDs The original PyTorch artifacts remain in the repo: z_traces.pt codebook.pt transition_vq_model.pt diagnostics.json z_trace_manifest.json checkpoints/ tabular10K<n<100K0 likes53 downloads4mo agoHugging Face17omron-sinicx /wiki2023_plus Overview The dataset includes the description from Wikipedia and categories of films published in 2023. This dataset is used to evaluate the ability of LLM to memorize and extract information described in the document. See "Where is the Answer? An Empirical Study of Positional Bias for Parametric Knowledge Extraction in Language Model (NAACL2025 Long paper)" for how we use this dataset for training and evaluation. Data Split film_doc_all.jsonl includes lines of… See the full description on the dataset page: https://huggingface.co/datasets/omron-sinicx/wiki2023_plus.tabular1K<n<10K1 likes52 downloads1y agoHugging Face18Omrihahami /HotelReservationsDataset Hotel Reservation EDA & Prediction Overview This project explores a hotel reservation dataset from Kaggle with the goal of understanding booking behavior and determining whether it is possible to predict reservation status (canceled or not canceled) based on the available features. The project focuses on Exploratory Data Analysis (EDA) only — no machine learning model was trained. Dataset This dataset includes 36,275 hotel reservations with 19 features… See the full description on the dataset page: https://huggingface.co/datasets/Omrihahami/HotelReservationsDataset.tabular10K<n<100K0 likes50 downloads10mo agoHugging Face19OmriShtayer /Website_Traffic_and_Engagementtabulartable-question-answeringn<1K0 likes36 downloads1y agoHugging Face20omrisap /dl-trm-phase2-codebook-v128 DL-TRM Phase 2 Codebook V128 This standalone dataset contains the Phase 2 discrete Z traces for V128. The Hugging Face dataset viewer reads data/train.jsonl. Each row contains: raw_puzzle_id route_id route_local_id medoid_old_id selected_stage_a_example_index cluster_size vocab_size z_trace: a length-16 list of discrete Z token IDs The original PyTorch artifacts remain in the repo: z_traces.pt codebook.pt transition_vq_model.pt diagnostics.json z_trace_manifest.json checkpoints/ tabular10K<n<100K0 likes36 downloads4mo agoHugging Face21omrisap /open_math_not_in_traintext10K<n<100K0 likes34 downloads6mo agoHugging Face22omrisap /nvidia_math_512_47Ktabular10K<n<100K0 likes32 downloads6mo agoHugging Face23omrisap /phaseZtext10K<n<100K0 likes31 downloads8mo agoHugging Face24omron-sinicx /AlignBench HalCap-Bench HalCap-Bench dataset. Columns model image_source image_name image_type sentence_index caption annotation error_type error_words agreement_ratio fleiss_Pi n_correct n_incorrect n_unknown image_url image_path_in_repo Notes Notes For COCO/CC12M items, the image is referenced by image_url. For SD/Imagen/data_generation items, the image file is stored under images/ and referenced by image_path_in_repo. image100K<n<1M2 likes28 downloads7mo agoHugging Face25omriuz /CharBench CharBench - Character-level benchmark and analysis suite for LLMs. CharBench is a large-scale benchmark for studying tokenization and character-level behavior in modern language models. For complete details on data curation and evaluation, see the paper. If you have ideas and suggestions to improve charbench feel free to reach out! uzan dot omri at gmail.com Usage from datasets import load_dataset ds = load_dataset("omriuz/CharBench") Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/omriuz/CharBench.tabular100K<n<1M0 likes27 downloads8mo agoHugging Face26omrisap /arc-agi-1-zloopvit-traces ARC-AGI-1 ZLoopViT Traces Canonical intermediate grid trajectories for all 400 training tasks in ARC-AGI-1. Each row corresponds to one original official train or test pair; generated augmentations are not included. Dataset contents 400 tasks 1,718 trajectories: 1,302 demonstration/train pairs and 416 test pairs 5,059 visible intermediate transitions Exact final-output validation on every official pair Important columns: task_id: official ARC task identifier… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/arc-agi-1-zloopvit-traces.textimage-to-image1K<n<10K0 likes27 downloads1mo agoHugging Face27omrisap /arc-agi-1-ruleloopvit-rules ARC-AGI-1 RuleLoopViT Rules This dataset contains one canonical, task-specific English rule for each of the 400 official ARC-AGI-1 training tasks. Rules were inferred only from official demonstration input/output pairs. Official test inputs, test outputs, and test traces were excluded from rule authoring. Each row includes: a concise standalone core_rule_text; a five-section full_rule_text; the corresponding structured sections; augmentation-aware references for colors… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/arc-agi-1-ruleloopvit-rules.tabulartext-classificationn<1K0 likes27 downloads1mo agoHugging Face28HugoSchtr /debussy-omr-fullpage-lvl Debussy OMR – Full-Page Level Dataset A dataset of 101 paired samples of full-page images from handwritten music scores written by Claude Debussy (French Composer, 1862 - 1918) and their corresponding MusicXML transcriptions. The manuscript images originate from the collections of the Bibliothèque nationale de France (BnF) and were accessed through the IIIF protocol. Dataset Description Property Value Samples 101 Documents 37 Image format JPEG… See the full description on the dataset page: https://huggingface.co/datasets/HugoSchtr/debussy-omr-fullpage-lvl.imageimage-to-textn<1K0 likes25 downloads7mo agoHugging Face29omrisap /nvidia_math_47Ktext10K<n<100K0 likes25 downloads7mo agoHugging Face30saurabh1896 /OMR-forms Dataset Card for "OMR-forms" More Information needed imagen<1K1 likes24 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.