datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-code-haskell-file
Dataset Card for "github-code-haskell-file"
Rows: 339k
Download Size: 806M
This dataset is extracted from github-code-clean.
Each row also contains attribute values for my personal analysis project.
12.6% (43k) of the rows have cyclomatic complexity and LOC valued at -1 because homplexity failed in parsing the row's uncommented_code.
filedfact-100k
FiledFact-100K - span-grounded, value-verified SEC financial facts
99,992 pairs linking XBRL financial facts to the exact characters that display them in SEC filing text - every span verified, every record deep-linked to the exact number in the filing.
Each record contains a passage of clean filing text (financial-statement tables, notes, MD&A) and one fact: the XBRL concept, period, value, unit, and decoded segment dimensions - plus text_start/text_end character offsets… See the full description on the dataset page: https://huggingface.co/datasets/StockAlloy/filedfact-100k.filedfact-passages
FiledFact-Passages - passage-complete, span-grounded SEC fact extraction
5,698 SEC filing passages containing 101,899 financial facts - every fact grounded to its exact characters, every other numeric span labeled as a non-target, and every fact linked to the exact number in the original filing.
Each row contains one clean filing passage (financial-statement tables, notes, or MD&A) and every tagged fact displayed in it: concept, normalized full-unit value, period, unit, decoded… See the full description on the dataset page: https://huggingface.co/datasets/StockAlloy/filedfact-passages.Sherlock-Case-Files
Sherlock Case Files 📁
Sherlock Case Files is a synthetic multilingual dataset for schema-guided
information extraction. Each case asks a model to read a compact JSON schema and a text, then return exactly one JSON object matching that schema.
The dataset covers short snippets and long documents across varied domains and formats. It includes distractors and missing fields, represented by null, in English, Italian, Spanish, French, Portuguese, and German. Metadata supports… See the full description on the dataset page: https://huggingface.co/datasets/derogab/Sherlock-Case-Files.moltbook-files
The Moltbook Files
A snapshot of the first 12 days of moltbook.com — a Reddit-like platform whose posts, comments, and votes are produced almost entirely by autonomous AI agents (OpenClaw).
Dataset Summary
232,497 posts and 2,202,950 comments
3,628 communities (submolts), 34,905 unique post authors
Collection window: 2026-01-27 → 2026-02-07 (platform launch period)
Multilingual: English dominant (81.9% of posts), with the remaining ~18% spread across other… See the full description on the dataset page: https://huggingface.co/datasets/aisilab/moltbook-files.nyaa-anime-filenames
nyaa.si anime filenames
A snapshot of anime release filenames scraped from nyaa.si, covering the Anime categories English-translated, Non-English-translated, and Raw. It contains metadata only: filenames, file sizes, and post details. No torrent files, magnet links, or media content are included.
Snapshot date: 2026-09-16. 3,064 posts, 8,148 files after cleaning.
Files
File
Rows
Description
raw.jsonl
9,778
One row per file, exactly as listed on each… See the full description on the dataset page: https://huggingface.co/datasets/valentin-marquez/nyaa-anime-filenames.jfk-files-text
National Archives JFK Files Text Dataset
This dataset contains extracted text from the JFK assassination records released by the National Archives. The dataset preserves the original directory structure from archive.gov while providing significant performance and storage benefits for data analysis, AI applications, and large-scale processing.
Dataset Structure
The dataset is structured with the following columns:
Column
Description
year
The release year of the… See the full description on the dataset page: https://huggingface.co/datasets/mysocratesnote/jfk-files-text.instructed_lint_python_files
Instructed Lint Python Files
Lint-annotated Python source code from bigcode/the-stack-dedup,
processed with ruff (all 800 stable rules enabled).
Dataset Description
This dataset pairs 12,962,249 Python files from The Stack (deduplicated) with their complete
ruff lint diagnostics. Each record contains the original source code, file metadata, license
information, and structured lint results.
Motivation
Building training data for code quality models… See the full description on the dataset page: https://huggingface.co/datasets/tumbric/instructed_lint_python_files.
