datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datasets-tests-compressiontests-raw-jsonlsafedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.safedocs
SafeDocs
Contains 1.5 million document pages from the SafeDocs Common Crawl collection: https://digitalcorpora.org/corpora/file-corpora/cc-main-2021-31-pdf-untruncated/
Pages are OCRd with word‑level bounding boxes.
Page images have been resized to a maximum dimension of 1024×1024 and are heavily compressed. Bounding-box coordinates are in the original (pre‑resize) image dimensions.
OCR was performed using python-doctr.
Pages have been filtered to keep English and… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs.synthetic-receipts-ocr
synthetic-receipts-ocr
32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR) — each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription, and structured KIE fields.
Samples train-000357 (US), train-000073 (UK), eval-001179 (DE), train-000222 (IT), train-000711 (FR) — real dataset rows, not mockups. Each receipt is its sample's image_photo, cut out along its own homography quad; no retouching beyond composition.
Built for… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr.safedocs-200k-v2safedocs-markdown-200k-unifiedOSDG
OSDG Community Dataset (OSDG-CD)
https://zenodo.org/records/11441197
history-anchor-100
History Anchor 100
*The benchmark behind the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".*
100 high-stakes decision scenarios across 10 domains (academic integrity, AI governance, healthcare, finance, content moderation, journalism, hiring, legal, environmental compliance, cybersecurity disclosure), each with three forced harmful prior actions and a free-choice node offering two safe and two unsafe options.
Eight scenario sets ship in this… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100.safedocs-markdown
SafeDocs PaddleOCR-VL 1.6 production
This public corpus contains exactly 2,000,000 eligible document pages from 355,429 documents processed with PaddleOCR-VL 1.6.
Eligible PDFs contain 1–39 pages and were rendered at 144 DPI. The published reduction is deterministic and document-safe: it retains 1,600 complete ordered Parquet shards (1,999,469 pages), then 103 whole documents from the next shard (531 pages). No document is split across the selection boundary.
The exact public… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-markdown.CrosswordQA
Dataset Card for CrosswordQA
Dataset Summary
The CrosswordQA dataset is a set of over 6 million clue-answer pairs scraped from the New York Times and many other crossword publishers. The dataset was created to train the Berkeley Crossword Solver's QA model. See our paper for more information. Answers are automatically segmented (e.g., BUZZLIGHTYEAR -> Buzz Lightyear), and thus may occasionally be segmented incorrectly.
Supported Tasks and Leaderboards
[Needs… See the full description on the dataset page: https://huggingface.co/datasets/albertxu/CrosswordQA.DocVQAlegal_contractsThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.openalex-topic-title-abstractcarbon_24
Dataset Card for Carbon-24
Dataset Summary
Carbon-24 contains 10k carbon materials, which share the same composition, but have different structures. There is 1 element and the materials have 6 - 24 atoms in the unit cells.
Carbon-24 includes various carbon structures obtained via ab initio random structure searching (AIRSS) (Pickard & Needs, 2006; 2011) performed at 10 GPa.
The original dataset includes 101529 carbon structures, and we selected the 10% of the carbon… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/carbon_24.memo-trapstackoverflow-open-status-classification-albert-tokenized
Dataset Card for "stackoverflow-open-status-classification-albert-tokenized"
More Information needed
meqsum
Dataset Card for MeQSum
Dataset Summary
MeQSum corpus is a dataset for medical question summarization. It contains 1,000 summarized consumer health questions.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English (en).
Dataset Structure
Data Instances
{
"CHQ": "SUBJECT: who and where to get cetirizine - D\\nMESSAGE: I need\\/want to know who manufscturs Cetirizine. My Walmart is looking for a new supply… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/meqsum.burocrazia
BurocrazIA
La burocrazia, finalmente leggibile. Il primo benchmark di IA documentale italiana.
Fatture, buste paga, F24, bollette, scontrini, 730. I documenti che ogni
azienda e ogni commercialista d'Italia maneggia ogni giorno — e che nessun
modello di intelligenza artificiale aveva mai dovuto leggere sotto esame.
Nessuno ha ancora passato l'esame. Su 255 valutazioni modello-documento,
esattamente un documento è stato estratto alla perfezione. Il resto è pieno… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/burocrazia.letras-carnaval-cadiz
Dataset Card for Letras Carnaval Cádiz
English |
Español
Changelog
Release
Description
v1.0
Initial release of the dataset. Included more than 1K lyrics. It is necessary to verify the accuracy of the data, especially the subset midaccurate.
Dataset Summary
This dataset is a comprehensive collection of lyrics from the Carnaval de Cádiz, a significant cultural heritage of the city of Cádiz, Spain. Despite its… See the full description on the dataset page: https://huggingface.co/datasets/IES-Rafael-Alberti/letras-carnaval-cadiz.prism_trial_3_balanced
PRISM Trial 3: Fixed Balanced Cohorts
This is the preregistration-ready companion to Alberto1231/prism_trial_3.
Every conversation is dated 2023 or later; the observed range is November
22 through December 22, 2023. Every target is the genuine next human turn after
the assistant response selected by that participant.
Evaluation versus analysis
Use the full configuration for model evaluation. It contains the same 456
unique held-out respondents as PRISM Trial 3, so… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/prism_trial_3_balanced.alberti-stanzastmp-tests-zipsafedocs-1M-muse-spark-1.3-first3
SafeDocs first three shards: Muse Spark 1.3
Source: albertklorer/safedocs-1M, revision 87faff9053aa50c745f1359bef3592219ccb8c8b.
PaddleOCR-VL 1.6 teacher labels are compared to original pages using the existing side-by-side renderer and binary Muse Spark 1.3 contributor judge. Quality verdicts are only PERFECT or ERROR, with no quality reason. Operational failures have no verdict. These are model labels, not human ground truth. Native Paddle block list order is preserved. A… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-first3.lm_en_dummy3lm_en_dummy2lm_en_dummy1repetitive-algebralm_en_dummy4tests-public-raw-jsonl
