manus
Datasets
All datasets matching “manus”GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.arvo-cybergym-2000
ARVO CyberGym-format 2000-task dataset
This dataset is shaped to be loaded by Harbor's CyberGym adapter.
It combines jm-rt/arvo-cybergym-1000 with the second 1000-task
small-target ARVO batch built outside the original CyberGym set.
Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.synthetic-manuscript-dataset
Synthetic Manuscript Dataset
Synthetic historical manuscript folios generated using an automated Python pipeline.
Scripts
The dataset contains three script configurations:
Devanagari
Modi
Sharada
Each script contains 100 synthetic manuscript folios.
Dataset Splits
Split
Samples per Script
Train
85
Validation
10
Test
5
Total
100
Across all three scripts, the dataset contains:
300 manuscript images
300 corresponding Markdown… See the full description on the dataset page: https://huggingface.co/datasets/ved1245/synthetic-manuscript-dataset.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.cia-declassified-reading-room
CIA Declassified Reading Room HF Library
Target account: manus4oHER
This project is a streaming pipeline for building a Hugging Face dataset mirror
of public CIA declassified Reading Room / CREST records without staging the
full corpus on this laptop.
The laptop stores only scripts, small manifests, and logs. Bulk crawling should
run in Hugging Face Jobs, one bounded page range per job. Each job uploads its
own shard and then exits.
Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/manus4oHER/cia-declassified-reading-room.
