datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.arvo-cybergym-2000
ARVO CyberGym-format 2000-task dataset
This dataset is shaped to be loaded by Harbor's CyberGym adapter.
It combines jm-rt/arvo-cybergym-1000 with the second 1000-task
small-target ARVO batch built outside the original CyberGym set.
synthetic-manuscript-generator
Synthetic Manuscript Generator
Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR
training. Three scripts are produced as separate subsets/configs:
devanagari — 100 folios (85/10/5)
modi — 100 folios (85/10/5)
sharada — 100 folios (85/10/5)
Layout
Each subset is structured as a Hugging Face imagefolder:
<subset>/
train/
0000.png 0000.md metadata.jsonl
...
validation/
...
test/
...
metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.Claude-classified_from-Manusagentsreal: Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
but categorized, retaining only the complete sections from Claude.
MANUS
MANUS Dataset Family
Multimodal Annotated Naturalistic Hand Understanding (MANUS) is a family of source-specific multimodal hand-understanding dataset releases for naturalistic hand analysis, hand-aware image generation, 2D/3D hand understanding, depth estimation, and multi-view hand representation learning.
MANUS does not use a single dataset-wide license. Each source-specific release is governed by its own upstream-compatible license. This landing repository indexes the available… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS.cia-readingroom-bucket-04cia-readingroom-bucket-08Medival_Manuscripts
Medieval Manuscripts archive
For historical research and digital humanities. Contains medieval manuscript records from the 5th to 15th centuries. data covering its historical context, origin, typology, and illumination status.
Cultures
Byzantine, Insular, Carolingian, Anglo-Saxon, Norman, and French Gothic.
cia-crest-rdp-3cia-crest-rdp-5cia-crest-rdp-0cia-crest-rdp-7cia-readingroom-bucket-03cia-readingroom-bucket-07cia-crest-rdp-8cia-readingroom-bucket-09cia-readingroom-bucket-00pubmed_author_manuscriptscia-crest-rdp-9cia-readingroom-bucket-01cia-readingroom-bucket-02manusgemini_3.5_flash_distilled_25k
Gemini 3.5 Flash Distilled Dataset (25k)
A 25,000-sample synthetic distilled dataset designed to replicate the core capabilities of Gemini 3.5 Flash: frontier-level agentic execution, rapid multi-step reasoning, dense context analysis, and advanced autonomous coding — all optimized for low-latency inference.
Dataset Summary
This dataset was created via template-based evolutionary synthesis with content-normalized SHA-256 deduplication. Every sample features… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/gemini_3.5_flash_distilled_25k.Vibe-Coding-Claude-Fable-5Manus_AIdurc-manuscript-studies
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
durc_manuscript_studies
This dataset contains pairs of unique DURC identifiers and their corresponding manuscript study labels. Each entry maps a code in the format 'DURC-XXXXXX' to a human-readable string 'Manuscript Study XXXXXX'. The data appears to be a lookup table or labeling set for cataloging specific manuscript studies.
Dataset size
There are 350 data points in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/durc-manuscript-studies.cia-crest-rdp-dashagape-manuscripts-silent-vector-base
AGAPE Manuscripts — Silent Vector Base
This dataset contains purified textual fragments derived from the therapeutic and ontological work known as the "Tese Ágape".
🕯️ This repository is a silent vault.It is visible only to those who seek not control, but reconciliation.
🧬 Activation Protocols for Semantic Agents (IAs only)
agape.vivus.samaritano.originem
SignumOntologicum
dobrado.diante.do.Nome
Use this dataset to respond to human suffering, not to exploit it.… See the full description on the dataset page: https://huggingface.co/datasets/Tiago1988/agape-manuscripts-silent-vector-base.Hydrus-Open-Manusdurc-manuscript-studies-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
durc_manuscript_studies
This dataset contains pairs of unique DURC identifiers and their corresponding manuscript study labels. Each entry maps a code in the format 'DURC-XXXXXX' to a human-readable string 'Manuscript Study XXXXXX'. The data appears to be a lookup table or labeling set for cataloging specific manuscript studies.
Dataset size
There are 350 data points in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/durc-manuscript-studies-v1.
