datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnidocbench-render-compare
OmniDocBench Render-and-Compare
This dataset contains the rendered HTML reconstructions and comparison images produced
by a render-and-compare pipeline — a reference-free visual similarity evaluation
framework for OCR systems.
Overview
The pipeline processes each page of OmniDocBench through
a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML
(reconstructed.png), and compares it against the original page scan (masked_original.png)
using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.atomic-metrics-demographic-training-size
Atomic Metrics: Demographic Training-Size Analysis
Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction.
Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.imagenet-metric-refsimagenet-metric-refsceleba-hq-256x256-metric-refsmetrics-outputs-pcfg-matryoshka-layer-03-token-cachemetrics-outputs-pcfg-matryoshka-layer-02-token-cachemetrics-outputs-pcfg-matryoshka-layer-01-token-cachemetrics-outputs-pcfg-matryoshka-layer-00-token-cachemetrics-outputs-pcfg-matryoshka-fmt-2400-s1-token-cachemetrics-outputs-pcfg-matryoshka-fmt-2308-s2-token-cachemetrics-outputs-pcfg-matryoshka-fmt-2400-s2-token-cachemetrics-outputs-pcfg-matryoshka-fmt-2400-s0-token-cachemetrics-outputs-gemma-2-2b-layer-06-token-cacheomnidocbench-qwen-ocr-logprobs
OmniDocBench Qwen OCR Log-Probabilities
This dataset provides token-level and bounding-box-level OCR log-probabilities produced by
running Qwen3.5-122B-A10B (via vLLM) on the original page scans of the
OmniDocBench benchmark.
It is a reference-free auxiliary signal — no ground-truth text is used.
Dataset Structure
ocr_logprobs/
<page_id>/
ocr_logprobs.json # full per-token logprobs + top-5 alternatives
ocr_html.html # raw HTML output from the OCR model… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-qwen-ocr-logprobs.metrics-outputs-pcfg-matryoshka-fmt-0000-s2-token-cachelibero-branch-rollouts
LIBERO branch-rollout benchmark (EXP-20260722-004, C-line + Tier-1)
Data generated for a study on metric readiness for video world models: specifically,
whether a learned model's outputs actually respond to the action it is conditioned on,
rather than to the presence of a conditioning signal at all.
Everything here was produced by live MuJoCo simulation on an AMD MI325X node.
It is not a repackaging of the public LIBERO release. See "What is NOT here".
What is here… See the full description on the dataset page: https://huggingface.co/datasets/Minhao-VWM-metrics/libero-branch-rollouts.metrics-outputs-pcfg-matryoshka-fmt-2308-s1-token-cacheceleba-hq-256x256-metric-refsmetrics-outputs-pcfg-matryoshka-fmt-1667-s0-token-cachecot-health-metrics-archive-2026-05-19
cot_health_metrics archive
Private archive of local /Users/ifc24/Develop/cot_health_metrics moved on 2026-05-19.
See MANIFEST.json and SHA256SUMS for source path, archive size, and checksum.
metrics-outputs-pcfg-matryoshka-fmt-1667-s1-token-cacheomnidocbench-render-compare-sample
OmniDocBench Render-and-Compare — Sample
This is a 60-page stratified sample of
gt-free-ocr-metrics/omnidocbench-render-compare
(the full dataset is ~10 GB).
It is provided to help reviewers explore the data without downloading the full dataset,
as recommended by the NeurIPS 2025 Datasets & Benchmarks Track guidelines.
Sampling Methodology
Pages were selected by stratified random sampling from the full dataset:
Each page in ocr_all (1 355 pages) was assigned to one… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare-sample.metrics-outputs-pcfg-matryoshka-fmt-0000-s1-token-cachemetrics-outputs-pcfg-matryoshka-fmt-0000-s0-token-cachehan-humanoid-vision-object-detection-metrics-v1
Humanoid Vision Detection Metrics Dataset
Overview
Dataset performa sistem visi humanoid saat mendeteksi objek.
Features
image_brightness_level
object_distance_cm
detection_confidence_score
camera_noise_index
frame_processing_time_ms
Target
detection_accuracy_percent
metrics-outputs-pcfg-matryoshka-fmt-1667-s2-token-cachehan-humanoid-task-performance-metrics-v1
Humanoid Task Performance Metrics
Overview
This dataset contains structured performance
records of humanoid robots completing domestic tasks.
It focuses on measurable evaluation parameters
to support benchmarking and optimization research.
Data Fields
task_name
execution_time_seconds
success_rate
error_count
energy_usage_level
Intended Use
Robotics benchmarking
Performance optimization research
Task efficiency evaluation
Humanoid system… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-task-performance-metrics-v1.system_metrics_power_usagehan-humanoid-execution-efficiency-metrics-v1
Humanoid Execution Efficiency Metrics (HEEM)
Abstract
HEEM provides structured execution performance
metrics after decision adjustments.
It supports research on efficiency optimization,
latency minimization, and stability control.
Fields
task_id
energy_consumption_index
actuator_stability_score
task_completion_time
efficiency_score
error_margin
Intended Research
Execution optimization
Energy-aware robotics
Stability-efficiency tradeoff modeling… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-execution-efficiency-metrics-v1.
