CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gt-free-ocr-metrics /omnidocbench-render-compare OmniDocBench Render-and-Compare This dataset contains the rendered HTML reconstructions and comparison images produced by a render-and-compare pipeline — a reference-free visual similarity evaluation framework for OCR systems. Overview The pipeline processes each page of OmniDocBench through a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML (reconstructed.png), and compares it against the original page scan (masked_original.png) using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.imageother10K<n<100K0 likes514 downloads5mo agoHugging Face02xPXXX /compare_oracletabular1K<n<10K0 likes341 downloads3y agoHugging Face03ComparEdge /ai-tools-pricing-2026 AI Tools Pricing & Features Dataset 2026 A structured dataset of 104 AI tools across 9 categories — pricing plans, user ratings, and feature lists. Built for market analysis, recommendation systems, and pricing research. Dataset Description This dataset covers the AI software landscape in 2026, including LLMs, coding assistants, image generators, and more. Each entry contains real pricing data, user ratings, and feature sets. Source Curated from the live… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/ai-tools-pricing-2026.tabulartext-classificationn<1K1 likes77 downloads5mo agoHugging Face04gt-free-ocr-metrics /omnidocbench-render-compare-sample OmniDocBench Render-and-Compare — Sample This is a 60-page stratified sample of gt-free-ocr-metrics/omnidocbench-render-compare (the full dataset is ~10 GB). It is provided to help reviewers explore the data without downloading the full dataset, as recommended by the NeurIPS 2025 Datasets & Benchmarks Track guidelines. Sampling Methodology Pages were selected by stratified random sampling from the full dataset: Each page in ocr_all (1 355 pages) was assigned to one… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare-sample.imageother1K<n<10K0 likes27 downloads5mo agoHugging Face05VictorShea /DocStream_Ground_Truth_Compare Feature Type Description event_idx int Event index within the session (aligns with HF and local JSONL). prompt string The full prompt used for the model prediction (from local JSONL). pred_thinking string Model-generated chain-of-thought from the local JSONL. thinking string Gold/reference chain-of-thought from HF. pred_depth int Predicted depth value. pred_annotation string Predicted one-sentence event annotation. gt_depth int Gold/reference depth label. gt_annotation… See the full description on the dataset page: https://huggingface.co/datasets/VictorShea/DocStream_Ground_Truth_Compare.tabularn<1K0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.