CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertvillanova /datasets-tests-compressiontextn<1K0 likes60k downloads5y agoHugging Face02albertvillanova /tests-raw-jsonltext10K<n<100K1 likes38k downloads5y agoHugging Face03albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads3d agoHugging Face04albertklorer /safedocs SafeDocs Contains 1.5 million document pages from the SafeDocs Common Crawl collection: https://digitalcorpora.org/corpora/file-corpora/cc-main-2021-31-pdf-untruncated/ Pages are OCRd with word‑level bounding boxes. Page images have been resized to a maximum dimension of 1024×1024 and are heavily compressed. Bounding-box coordinates are in the original (pre‑resize) image dimensions. OCR was performed using python-doctr. Pages have been filtered to keep English and… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs.image1M<n<10M0 likes769 downloads10mo agoHugging Face05albertobarnabo /synthetic-receipts-ocr synthetic-receipts-ocr 32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR) — each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription, and structured KIE fields. Samples train-000357 (US), train-000073 (UK), eval-001179 (DE), train-000222 (IT), train-000711 (FR) — real dataset rows, not mockups. Each receipt is its sample's image_photo, cut out along its own homography quad; no retouching beyond composition. Built for… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr.imageimage-to-text10K<n<100K0 likes669 downloads2mo agoHugging Face06albertklorer /safedocs-200k-v2text10K<n<100K0 likes629 downloads19d agoHugging Face07albertklorer /safedocs-markdown-200k-unifiedtext10K<n<100K0 likes607 downloads1mo agoHugging Face08albertmartinez /OSDG OSDG Community Dataset (OSDG-CD) https://zenodo.org/records/11441197 texttext-classification100K<n<1M1 likes577 downloads2y agoHugging Face09albertoRodriguez97 /history-anchor-100 History Anchor 100 *The benchmark behind the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".* 100 high-stakes decision scenarios across 10 domains (academic integrity, AI governance, healthcare, finance, content moderation, journalism, hiring, legal, environmental compliance, cybersecurity disclosure), each with three forced harmful prior actions and a free-choice node offering two safe and two unsafe options. Eight scenario sets ship in this… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100.texttext-generationn<1K0 likes514 downloads4mo agoHugging Face10albertklorer /safedocs-markdown SafeDocs PaddleOCR-VL 1.6 production This public corpus contains exactly 2,000,000 eligible document pages from 355,429 documents processed with PaddleOCR-VL 1.6. Eligible PDFs contain 1–39 pages and were rendered at 144 DPI. The published reduction is deterministic and document-safe: it retains 1,600 complete ordered Parquet shards (1,999,469 pages), then 103 whole documents from the next shard (531 pages). No document is split across the selection boundary. The exact public… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-markdown.image100K<n<1M0 likes333 downloads1mo agoHugging Face11albertxu /CrosswordQA Dataset Card for CrosswordQA Dataset Summary The CrosswordQA dataset is a set of over 6 million clue-answer pairs scraped from the New York Times and many other crossword publishers. The dataset was created to train the Berkeley Crossword Solver's QA model. See our paper for more information. Answers are automatically segmented (e.g., BUZZLIGHTYEAR -> Buzz Lightyear), and thus may occasionally be segmented incorrectly. Supported Tasks and Leaderboards [Needs… See the full description on the dataset page: https://huggingface.co/datasets/albertxu/CrosswordQA.textquestion-answering1M<n<10M6 likes318 downloads4y agoHugging Face12albertklorer /DocVQAimagequestion-answering10K<n<100K0 likes316 downloads8mo agoHugging Face13albertvillanova /legal_contractsThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.text100K<n<1M54 likes313 downloads5y agoHugging Face14albertmartinez /openalex-topic-title-abstracttext1M<n<10M1 likes301 downloads2y agoHugging Face15albertvillanova /carbon_24 Dataset Card for Carbon-24 Dataset Summary Carbon-24 contains 10k carbon materials, which share the same composition, but have different structures. There is 1 element and the materials have 6 - 24 atoms in the unit cells. Carbon-24 includes various carbon structures obtained via ab initio random structure searching (AIRSS) (Pickard & Needs, 2006; 2011) performed at 10 GPa. The original dataset includes 101529 carbon structures, and we selected the 10% of the carbon… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/carbon_24.tabularother10K<n<100K1 likes221 downloads4y agoHugging Face16Albertmade /memo-traptabularn<1K0 likes180 downloads2y agoHugging Face17reubenjohn /stackoverflow-open-status-classification-albert-tokenized Dataset Card for "stackoverflow-open-status-classification-albert-tokenized" More Information needed text1M<n<10M0 likes178 downloads4y agoHugging Face18albertvillanova /meqsum Dataset Card for MeQSum Dataset Summary MeQSum corpus is a dataset for medical question summarization. It contains 1,000 summarized consumer health questions. Supported Tasks and Leaderboards [More Information Needed] Languages English (en). Dataset Structure Data Instances { "CHQ": "SUBJECT: who and where to get cetirizine - D\\nMESSAGE: I need\\/want to know who manufscturs Cetirizine. My Walmart is looking for a new supply… See the full description on the dataset page: https://huggingface.co/datasets/albertvillanova/meqsum.textsummarization1K<n<10K11 likes153 downloads3y agoHugging Face19albertobarnabo /burocrazia BurocrazIA La burocrazia, finalmente leggibile. Il primo benchmark di IA documentale italiana. Fatture, buste paga, F24, bollette, scontrini, 730. I documenti che ogni azienda e ogni commercialista d'Italia maneggia ogni giorno — e che nessun modello di intelligenza artificiale aveva mai dovuto leggere sotto esame. Nessuno ha ancora passato l'esame. Su 255 valutazioni modello-documento, esattamente un documento è stato estratto alla perfezione. Il resto è pieno… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/burocrazia.imageimage-to-textn<1K0 likes148 downloads15d agoHugging Face20IES-Rafael-Alberti /letras-carnaval-cadiz Dataset Card for Letras Carnaval Cádiz English | Español Changelog Release Description v1.0 Initial release of the dataset. Included more than 1K lyrics. It is necessary to verify the accuracy of the data, especially the subset midaccurate. Dataset Summary This dataset is a comprehensive collection of lyrics from the Carnaval de Cádiz, a significant cultural heritage of the city of Cádiz, Spain. Despite its… See the full description on the dataset page: https://huggingface.co/datasets/IES-Rafael-Alberti/letras-carnaval-cadiz.tabular1K<n<10K3 likes142 downloads2y agoHugging Face21Alberto1231 /prism_trial_3_balanced PRISM Trial 3: Fixed Balanced Cohorts This is the preregistration-ready companion to Alberto1231/prism_trial_3. Every conversation is dated 2023 or later; the observed range is November 22 through December 22, 2023. Every target is the genuine next human turn after the assistant response selected by that participant. Evaluation versus analysis Use the full configuration for model evaluation. It contains the same 456 unique held-out respondents as PRISM Trial 3, so… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/prism_trial_3_balanced.texttext-generation1K<n<10K1 likes140 downloads1mo agoHugging Face22alvp /alberti-stanzastabular1K<n<10K0 likes136 downloads2y agoHugging Face23albertvillanova /tmp-tests-ziptabularn<1K0 likes127 downloads5y agoHugging Face24albertklorer /safedocs-1M-muse-spark-1.3-first3 SafeDocs first three shards: Muse Spark 1.3 Source: albertklorer/safedocs-1M, revision 87faff9053aa50c745f1359bef3592219ccb8c8b. PaddleOCR-VL 1.6 teacher labels are compared to original pages using the existing side-by-side renderer and binary Muse Spark 1.3 contributor judge. Quality verdicts are only PERFECT or ERROR, with no quality reason. Operational failures have no verdict. These are model labels, not human ground truth. Native Paddle block list order is preserved. A… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-first3.tabularn<1K0 likes125 downloads6d agoHugging Face25albertvillanova /lm_en_dummy3textn<1K0 likes124 downloads5y agoHugging Face26albertvillanova /lm_en_dummy2textn<1K0 likes123 downloads5y agoHugging Face27albertvillanova /lm_en_dummy1textn<1K0 likes122 downloads5y agoHugging Face28Albertmade /repetitive-algebratabular1K<n<10K0 likes122 downloads2y agoHugging Face29albertvillanova /lm_en_dummy4textn<1K0 likes121 downloads5y agoHugging Face30albertvillanova /tests-public-raw-jsonltextn<1K0 likes121 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.