datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.arvo-cybergym-2000
ARVO CyberGym-format 2000-task dataset
This dataset is shaped to be loaded by Harbor's CyberGym adapter.
It combines jm-rt/arvo-cybergym-1000 with the second 1000-task
small-target ARVO batch built outside the original CyberGym set.
Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.cia-declassified-reading-room
CIA Declassified Reading Room HF Library
Target account: manus4oHER
This project is a streaming pipeline for building a Hugging Face dataset mirror
of public CIA declassified Reading Room / CREST records without staging the
full corpus on this laptop.
The laptop stores only scripts, small manifests, and logs. Bulk crawling should
run in Hugging Face Jobs, one bounded page range per job. Each job uploads its
own shard and then exits.
Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/manus4oHER/cia-declassified-reading-room.Multilingual-Medical-Corpus
Mutilingual Medical Corpus
Multilingual-Medical-Corpus a 3 billion word multilingual corpus for training LLMs adapted to the medical domain. Multilingual-Medical-Corpus includes four languages, namely, English, Spanish, French, and Italian.
📖 Paper: Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
🌐 Project Website: https://univ-cotedazur.eu/antidote
Corpus Description
Developed by: Iker García-Ferrero, Rodrigo Agerri… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Multilingual-Medical-Corpus.NOSK-Hackingcyber-security-100m
WitFoo Precinct6 Cybersecurity Dataset (large)
Overview
A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges).
Available in two sizes:
witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.security
Security Knowledge Graph Triples
Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis.
Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV
Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/security.ZINC20ZINC20 Dataset with SELFIES added. Any smile that could not be successfully converted was dropped from the dataset.
Every tranch was downloaded, this is not the ~1B example ML subset from https://files.docking.org/zinc20-ML/.
The dataset was entirely shuffled then split into 80%/10%/10% splits for train/val/test.
A file vocab.csv is in the root of the reposity that contains all of the SELFIES tokens found in the data, with [START], [STOP], and [PAD] added.
MANUS-HaGRID
MANUS-HaGRID: HaGRID-derived Multimodal Annotated Naturalistic Hand Understanding Dataset
MANUS-HaGRID is the HaGRID/HaGRIDv2-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It provides multimodal annotations for naturalistic hand gesture understanding, including RGB images, hand crops, depth maps, 2D bounding boxes, estimated MANO-style hand mesh metadata, and multi-view mesh renderings where available.
This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-HaGRID.Arabic_Manuscript_Collection_Dataset
Arabic Manuscript Collection
Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout
and one label format so they can be trained and evaluated together: 82,561 labelled
images in total. Two are republished closer to their source shape: AMIDDA as upstream
Parquet, and OpenITI-Makhzan as page images with line-level coordinates.
Four of the five converted sources are historical manuscripts. KHATT is modern handwriting
and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.synthetic-manuscript-generator
Synthetic Manuscript Generator
Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR
training. Three scripts are produced as separate subsets/configs:
devanagari — 100 folios (85/10/5)
modi — 100 folios (85/10/5)
sharada — 100 folios (85/10/5)
Layout
Each subset is structured as a Hugging Face imagefolder:
<subset>/
train/
0000.png 0000.md metadata.jsonl
...
validation/
...
test/
...
metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.domains
Internet Domains
Domains
HuggingFace Hub Mirror for https://github.com/pkgforge-security/domains
The Sync Workflow actions are at: https://github.com/pkgforge-security/domains
TOS & Abuse (To Hugging-Face's Staff)
Hi, if you are an offical from Hugging-Face here to investigate why this Repo is so Large and are considering deleting, & terminating our Account.
Please note that, this project benefits a lot of people (You can do a… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/domains.forward_pc_hebbian_lattice
forward_pc_hebbian lattice dump (full)
Public full dump of dual morphogenetic / free PC-Hebbian work under checkpoints/.
trajectories/morph_A|B — dual morph projective lattices (very large)
trajectories/dual_*, clean_100k_priority, fat_first_100k, live_continuous
dual_morphogenetic/ — net checkpoints + sediment field
Uploaded from a disk-constrained machine; local copies may remain until verified.
omar-al-saleh-manuscripts-segments
Omar Al-Saleh Manuscripts — Segments
Line-level segmented images with transcriptions from the Omar Al-Saleh memoir collection (1951–1965), part of the NAKBA NLP 2026: Arabic Manuscript Understanding Shared Task.
Dataset
Split
Images
With text
train
15,969
15,969
test
2,095
2,095
blind_test
2,671
2,671
Each example contains:
image: A cropped line image from a manuscript page (JPG or PNG)
text: The Arabic transcription of that line
filename: Original… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/omar-al-saleh-manuscripts-segments.hacking
Hacking Text Corpus
A research corpus of historical computer security writings, hacker zines, and hacktivist texts. Built for NLP, text generation, discourse analysis, and security research.
Contents
Phrack Magazine (phrack/)
72 issues (1985-2024), 1,026 articles
~55 MB of raw text, ~4.76 million words
Organized as phrack/issue{N}/{article}.txt
Topics: exploit development, reverse engineering, networking, phreaking, hacker culture, OS internals… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/hacking.indic-historical-manuscripts
Synthetic Indic Manuscript Dataset (Devanagari, Modi, Sharada)
This dataset contains synthetic historical manuscript folios paired with exact Markdown (.md) ground-truth transcriptions.
Subsets and Distribution
Subsets: devanagari, modi, sharada
Standard splits:
train: 85%
validation: 10%
test: 5%
Features & Physical Fidelity
Backgrounds: High-resolution procedural aged paper (pothi) & palm-leaf (talapatra) folios with string hole punch marks… See the full description on the dataset page: https://huggingface.co/datasets/varunbhoyar/indic-historical-manuscripts.Sumtables-Cuneiform-Full-Fable5-Remaster
Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset
A rebuilt, leakage-proof, multi-task training dataset for teaching
vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually
read, transliterate, and translate Sumerian cuneiform tablets from
photographs. The mission: produce useful first-pass readings for the ~90% of
excavated tablets that have never been published or translated.
Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Sumtables-Cuneiform-Full-Fable5-Remaster.manuscript_noisy_labelsmanuscript_noisy_labels_iiifmetaboverse-manuscriptRapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts
Dataset Attribution
The original dataset is available on Kaggle.
This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work.
Please cite the original authors if you use this dataset.
Citation
@INPROCEEDINGS{8978005,
author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.},
booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.Claude-classified_from-Manusagentsreal: Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
but categorized, retaining only the complete sections from Claude.
MANUS-DexYCB
MANUS-DexYCB: DexYCB-derived Multimodal Annotated Naturalistic Hand Understanding Dataset
MANUS-DexYCB is the DexYCB-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It is released as a source-specific repository because MANUS subsets are governed by different upstream licenses.
This repository contains only the DexYCB-derived MANUS test split. HaGRID/HaGRIDv2-derived data is released separately as MANUS-HaGRID.… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-DexYCB.GPT-5.6-Sol-Luna-Terra-Traces
GPT-5.6 — Sol · Terra · Luna Library
A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place.
Dataset Viewer | Parquet
// what this is
This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. It exists to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.6-Sol-Luna-Terra-Traces.seen_canaries
Dataset Card for "seen_canaries"
More Information needed
CyberSecurity-bigsynthetic-manuscript-generatorvoynich-manuscript-metadata
Voynich Manuscript Metadata
Dataset Summary
This dataset contains structured metadata about the Voynich Manuscript (Beinecke MS 408), a famous 15th-century codex held at Yale's Beinecke Rare Book & Manuscript Library. The dataset includes three tables: pages, folios, and quires, providing comprehensive codicological information.
Dataset Structure
Configurations
This dataset has three configurations:
pages: Page-level metadata (226… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-manuscript-metadata.
