datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnidocbench-render-compare
OmniDocBench Render-and-Compare
This dataset contains the rendered HTML reconstructions and comparison images produced
by a render-and-compare pipeline — a reference-free visual similarity evaluation
framework for OCR systems.
Overview
The pipeline processes each page of OmniDocBench through
a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML
(reconstructed.png), and compares it against the original page scan (masked_original.png)
using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.compare_oraclegspc-compare
GSPC compare
Compare door. Prints the living board only. No second scoreboard.
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-compare.ai-tools-pricing-2026
AI Tools Pricing & Features Dataset 2026
A structured dataset of 104 AI tools across 9 categories — pricing plans, user ratings, and feature lists. Built for market analysis, recommendation systems, and pricing research.
Dataset Description
This dataset covers the AI software landscape in 2026, including LLMs, coding assistants, image generators, and more. Each entry contains real pricing data, user ratings, and feature sets.
Source
Curated from the live… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/ai-tools-pricing-2026.omnidocbench-render-compare-sample
OmniDocBench Render-and-Compare — Sample
This is a 60-page stratified sample of
gt-free-ocr-metrics/omnidocbench-render-compare
(the full dataset is ~10 GB).
It is provided to help reviewers explore the data without downloading the full dataset,
as recommended by the NeurIPS 2025 Datasets & Benchmarks Track guidelines.
Sampling Methodology
Pages were selected by stratified random sampling from the full dataset:
Each page in ocr_all (1 355 pages) was assigned to one… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare-sample.pairwise_compare_all_dpo_llama_3_70b_datapairwise_compare_top1_and_SPINS_llama_3_70b_dpo_datapairwise_only_compare_rank_1_to_all_llama_3_70b_dpo_dataFINETUNE4_compare8k2FINETUNE4_compare8kdata-compareFINETUNE4_compare15kDocStream_Ground_Truth_Compare
Feature
Type
Description
event_idx
int
Event index within the session (aligns with HF and local JSONL).
prompt
string
The full prompt used for the model prediction (from local JSONL).
pred_thinking
string
Model-generated chain-of-thought from the local JSONL.
thinking
string
Gold/reference chain-of-thought from HF.
pred_depth
int
Predicted depth value.
pred_annotation
string
Predicted one-sentence event annotation.
gt_depth
int
Gold/reference depth label.
gt_annotation… See the full description on the dataset page: https://huggingface.co/datasets/VictorShea/DocStream_Ground_Truth_Compare.
