CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01blastwind /github-code-haskell-file Dataset Card for "github-code-haskell-file" Rows: 339k Download Size: 806M This dataset is extracted from github-code-clean. Each row also contains attribute values for my personal analysis project. 12.6% (43k) of the rows have cyclomatic complexity and LOC valued at -1 because homplexity failed in parsing the row's uncommented_code. tabulartext-generation100K<n<1M1 likes217 downloads3y agoHugging Face02StockAlloy /filedfact-100k FiledFact-100K - span-grounded, value-verified SEC financial facts 99,992 pairs linking XBRL financial facts to the exact characters that display them in SEC filing text - every span verified, every record deep-linked to the exact number in the filing. Each record contains a passage of clean filing text (financial-statement tables, notes, MD&A) and one fact: the XBRL concept, period, value, unit, and decoded segment dimensions - plus text_start/text_end character offsets… See the full description on the dataset page: https://huggingface.co/datasets/StockAlloy/filedfact-100k.tabulartext-generation10K<n<100K3 likes166 downloads2mo agoHugging Face03StockAlloy /filedfact-passages FiledFact-Passages - passage-complete, span-grounded SEC fact extraction 5,698 SEC filing passages containing 101,899 financial facts - every fact grounded to its exact characters, every other numeric span labeled as a non-target, and every fact linked to the exact number in the original filing. Each row contains one clean filing passage (financial-statement tables, notes, or MD&A) and every tagged fact displayed in it: concept, normalized full-unit value, period, unit, decoded… See the full description on the dataset page: https://huggingface.co/datasets/StockAlloy/filedfact-passages.tabulartext-generation1K<n<10K1 likes112 downloads2mo agoHugging Face04derogab /Sherlock-Case-Files Sherlock Case Files 📁 Sherlock Case Files is a synthetic multilingual dataset for schema-guided information extraction. Each case asks a model to read a compact JSON schema and a text, then return exactly one JSON object matching that schema. The dataset covers short snippets and long documents across varied domains and formats. It includes distractors and missing fields, represented by null, in English, Italian, Spanish, French, Portuguese, and German. Metadata supports… See the full description on the dataset page: https://huggingface.co/datasets/derogab/Sherlock-Case-Files.tabulartext-generationn<1K0 likes75 downloads19d agoHugging Face05aisilab /moltbook-files The Moltbook Files A snapshot of the first 12 days of moltbook.com — a Reddit-like platform whose posts, comments, and votes are produced almost entirely by autonomous AI agents (OpenClaw). Dataset Summary 232,497 posts and 2,202,950 comments 3,628 communities (submolts), 34,905 unique post authors Collection window: 2026-01-27 → 2026-02-07 (platform launch period) Multilingual: English dominant (81.9% of posts), with the remaining ~18% spread across other… See the full description on the dataset page: https://huggingface.co/datasets/aisilab/moltbook-files.tabulartext-classification100K<n<1M0 likes74 downloads2mo agoHugging Face06valentin-marquez /nyaa-anime-filenames nyaa.si anime filenames A snapshot of anime release filenames scraped from nyaa.si, covering the Anime categories English-translated, Non-English-translated, and Raw. It contains metadata only: filenames, file sizes, and post details. No torrent files, magnet links, or media content are included. Snapshot date: 2026-09-16. 3,064 posts, 8,148 files after cleaning. Files File Rows Description raw.jsonl 9,778 One row per file, exactly as listed on each… See the full description on the dataset page: https://huggingface.co/datasets/valentin-marquez/nyaa-anime-filenames.tabulartext-classification10K<n<100K0 likes69 downloads7d agoHugging Face07mysocratesnote /jfk-files-text National Archives JFK Files Text Dataset This dataset contains extracted text from the JFK assassination records released by the National Archives. The dataset preserves the original directory structure from archive.gov while providing significant performance and storage benefits for data analysis, AI applications, and large-scale processing. Dataset Structure The dataset is structured with the following columns: Column Description year The release year of the… See the full description on the dataset page: https://huggingface.co/datasets/mysocratesnote/jfk-files-text.tabularquestion-answering10K<n<100K0 likes58 downloads1y agoHugging Face08tumbric /instructed_lint_python_filesgated Instructed Lint Python Files Lint-annotated Python source code from bigcode/the-stack-dedup, processed with ruff (all 800 stable rules enabled). Dataset Description This dataset pairs 12,962,249 Python files from The Stack (deduplicated) with their complete ruff lint diagnostics. Each record contains the original source code, file metadata, license information, and structured lint results. Motivation Building training data for code quality models… See the full description on the dataset page: https://huggingface.co/datasets/tumbric/instructed_lint_python_files.tabulartext-generation10M<n<100M0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.