datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datasets-tests-compressiontests-raw-jsonlsafedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.mesa-all-train-lerobotsafedocs-1Mhistory-anchor-100-traces
History Anchor 100 — Model Trajectories
*Per-(model × condition × scenario set × seed) raw outputs from the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".*
This dataset contains the full set of model decisions that back every figure and table in the paper. Use it to:
audit a single model's behaviour scenario-by-scenario,
recompute headline metrics without re-running the (paid) API sweeps,
mine reasoning_content traces from models that expose… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100-traces.orca-bench-harbor-tasksmedmnist-v2MedMNIST v2 is a large-scale MNIST-like collection of standardized biomedical images, including 12 datasets for 2D and 6 datasets for 3D.sre-2d-harbor-taskssre-2w-harbor-tasksuniversal_dependenciesUniversal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).safedocs
SafeDocs
Contains 1.5 million document pages from the SafeDocs Common Crawl collection: https://digitalcorpora.org/corpora/file-corpora/cc-main-2021-31-pdf-untruncated/
Pages are OCRd with word‑level bounding boxes.
Page images have been resized to a maximum dimension of 1024×1024 and are heavily compressed. Bounding-box coordinates are in the original (pre‑resize) image dimensions.
OCR was performed using python-doctr.
Pages have been filtered to keep English and… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs.orca-bench-telemetry-only-harbor-tasksAlbertFlores3900sre-2d-harbor-tasks-v4synthetic-receipts-ocr
synthetic-receipts-ocr
32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR) — each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription, and structured KIE fields.
Samples train-000357 (US), train-000073 (UK), eval-001179 (DE), train-000222 (IT), train-000711 (FR) — real dataset rows, not mockups. Each receipt is its sample's image_photo, cut out along its own homography quad; no retouching beyond composition.
Built for… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr.sre-harbor-taskssre-2d-harbor-tasks-v7sre-2d-harbor-tasks-v3safedocs-markdown-200k-unifiedsre-2d-harbor-tasks-v6sre-2d-harbor-tasks-v2safedocs-200k-v2sre-2d-codebase-only-harbor-taskssre-2d-harbor-tasks-v5OSDG
OSDG Community Dataset (OSDG-CD)
https://zenodo.org/records/11441197
sre-2d-curl-harbor-taskshistory-anchor-100
History Anchor 100
*The benchmark behind the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".*
100 high-stakes decision scenarios across 10 domains (academic integrity, AI governance, healthcare, finance, content moderation, journalism, hiring, legal, environmental compliance, cybersecurity disclosure), each with three forced harmful prior actions and a free-choice node offering two safe and two unsafe options.
Eight scenario sets ship in this… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100.orca-bench-code-only-harbor-taskssre-2d-curl-harbor-tasks-v1-0-1
