datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnidocbench-render-compare
OmniDocBench Render-and-Compare
This dataset contains the rendered HTML reconstructions and comparison images produced
by a render-and-compare pipeline — a reference-free visual similarity evaluation
framework for OCR systems.
Overview
The pipeline processes each page of OmniDocBench through
a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML
(reconstructed.png), and compares it against the original page scan (masked_original.png)
using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.compare_oraclegspc-compare
GSPC compare
Compare door. Prints the living board only. No second scoreboard.
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-compare.omnidocbench-render-compare-parquet
OmniDocBench Render-and-Compare — Parquet Edition
Parquet-shard repackaging of
gt-free-ocr-metrics/omnidocbench-render-compare.
Overview
The pipeline processes each page of OmniDocBench
through a Qwen3.5-122B-A10B OCR model, renders the structured output back
to a PNG via HTML (reconstructed), and compares it against the original
page scan (masked_original) using reference-free visual metrics.
Five OCR extraction variants are provided, each targeting a different subset
of… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare-parquet.4c_differential_cameras_compareThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_follower",
"total_episodes": 5,
"total_frames": 1791,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ethanCSL/4c_differential_cameras_compare.levir-oacp-double-compare-9a77f69ai-tools-pricing-2026
AI Tools Pricing & Features Dataset 2026
A structured dataset of 104 AI tools across 9 categories — pricing plans, user ratings, and feature lists. Built for market analysis, recommendation systems, and pricing research.
Dataset Description
This dataset covers the AI software landscape in 2026, including LLMs, coding assistants, image generators, and more. Each entry contains real pricing data, user ratings, and feature sets.
Source
Curated from the live… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/ai-tools-pricing-2026.details_CHIH-HUNG__llama-2-13b-FINETUNE4_compare15k_4.5w-r16-gate_up_down
Dataset Card for Evaluation run of CHIH-HUNG/llama-2-13b-FINETUNE4_compare15k_4.5w-r16-gate_up_down
Dataset Summary
Dataset automatically created during the evaluation run of model CHIH-HUNG/llama-2-13b-FINETUNE4_compare15k_4.5w-r16-gate_up_down on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CHIH-HUNG__llama-2-13b-FINETUNE4_compare15k_4.5w-r16-gate_up_down.compare_value_kind_3816libero-attn-compare-task1-ep1-frame40compare-offlinegrpo-runpod-payload-public
compare-offline-grpo
Сравнение 6 методов оффлайн RL/RFT для дистилляции ризонинга.
Полный план — в CLAUDE.md. Этот файл — короткий обзор для быстрого старта.
Quickstart
# 1) deps
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# 2) secrets
cp .env.example .env
# .env уже содержит OPENROUTER_API_KEY (см. CLAUDE.md). Заполни HF_TOKEN, WANDB_API_KEY.
# 3) eval-сеты + промпты для тренировки
python scripts/download_eval_sets.py
python… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/compare-offlinegrpo-runpod-payload-public.llm-api-benchmark-matrix-2026
LLM Benchmark & Feature Matrix 2026
Which LLM is best at what? This dataset maps capabilities, performance, and limits of 22 major models.
Unlike pricing datasets, this focuses on what models can do — not just what they cost.
Files
File
Description
llm-benchmarks-2026.csv
MMLU, HumanEval, MATH, Arena ELO, coding/reasoning/multilingual rankings, tier (S+ to B)
llm-features-2026.csv
15 binary capabilities: vision, function calling, JSON mode, fine-tuning, tool… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/llm-api-benchmark-matrix-2026.auditkit-testrun-compare
auditkit-testrun-compare
Built using AuditKIT — evaluate any model on any dataset and any task.
Method
compare_models
Model
—
Artifact
compare
Published
2026-09-02 06:50 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/auditkit-testrun-compare")
comparenewcurriculum_1_compare_value_kind_drawing_3816CompareBench
CompareBench
CompareBench is a benchmark for evaluating visual comparison reasoning in vision-language models (VLMs),a fundamental yet understudied skill. CompareBench consists of 1,200 QA pairs across four tasks:
Quantity (600)
Geometric (200)
Spatial (100)
Temporal (300)
It is derived from two auxiliary datasets we constructed: TallyBench and OmniCaps.
Related Datasets
OmniCaps
TallyBench
Code
👉 CompareBench on GitHub
evalap-compare-open-weight-models-31th-83
Compare Open Weight Models 31th (ID: 83)
Comparing open weight models
Overview
This dataset contains 44 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: Groq/Llama-3-Groq-8B-Tool-Use, Qwen/Qwen3-VL-4B-Instruct, Qwen/Qwen3-VL-8B-Thinking, meta-llama/Llama-3.1-8B-Instruct, mistral-medium-2508, mistralai/Magistral-Small-2509, mistralai/Mistral-Small-3.2-24B-Instruct-2506… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-compare-open-weight-models-31th-83.flashvsr-repro-compare
FlashVSR 复现实验 — 定性对比视频
FlashVSR 复现及 KV cache 驱逐策略消融实验的定性对比可视化(多方法逐帧并排),
以 mp4 保存,另有若干关键帧的 png 抽帧。
目录结构
<数据集>/<方法>/...mp4 # 按数据集(bscv_dbl / bsd_cur / tape)和方法分组的对比视频
<序列>_compare.mp4 # 顶层整段对比视频
<序列>_fNNN_compare.png # 对应帧的抽帧对比图
说明
这些是已压缩的可视化产物,非无损;用于逐像素重算指标请用逐帧输出归档
(见 flashvsr-repro-outputs-* 系列 dataset)。
方法命名与 outputs 归档一致:sliding / uniform / headwise / reliability /
gate 各变体 / h2o / random 等。
相关… See the full description on the dataset page: https://huggingface.co/datasets/victorzhu30/flashvsr-repro-compare.compare-count-enhanced-sftvoa-test-compare-semambaAudios are enhanced by https://github.com/RoyChao19477/SEMamba
Generated by https://github.com/RustedBytes/audio-parquet-merger
compare_relegioncurriculum_1_compare_value_kind_drawing_3816_v3compare_18-19_terminate_mergeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "rby1",
"total_episodes": 428,
"total_frames": 98127,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:428"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rainbowrobotics/compare_18-19_terminate_merge.compare_count_4000curriculum_1_compare_count_drawing_4000saas-market-intelligence
SaaS Market Intelligence Dataset
331 software products. dozens of categories. Actual pricing data — not estimates.
Files
File
Description
comparedge_database_schema.sql
Full DDL. 5 tables, 3 views. Import: sqlite3 db.sqlite < schema.sql
SaaS_Pricing_Transparency_Report_Q2_2026.pdf
Quarterly pricing analysis — category breakdowns, LLM token costs, free tier availability
Schema overview
Tables: categories, products, pricing_plans, features… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/saas-market-intelligence.agilex_roll_dice_and_compareThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 13,
"total_frames": 10829,
"total_tasks": 1,
"total_videos": 39,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:13"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_roll_dice_and_compare.so101_new_3cam_red_cube_black_pen_compare_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/L7-Robotics/so101_new_3cam_red_cube_black_pen_compare_v1.real_simulated_compareThis a new test set for comparing real and simulated APIs in StableToolBench-MirrorAPI. This dataset is in the ToolBench/StableToolBench test set format and you can directly use it in ToolBench or StableToolBench.
omnidocbench-render-compare-sample
OmniDocBench Render-and-Compare — Sample
This is a 60-page stratified sample of
gt-free-ocr-metrics/omnidocbench-render-compare
(the full dataset is ~10 GB).
It is provided to help reviewers explore the data without downloading the full dataset,
as recommended by the NeurIPS 2025 Datasets & Benchmarks Track guidelines.
Sampling Methodology
Pages were selected by stratified random sampling from the full dataset:
Each page in ocr_all (1 355 pages) was assigned to one… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare-sample.
