CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M197 likes16k downloads28d agoHugging Face02Manusagents /arvo-cybergym-2000 ARVO CyberGym-format 2000-task dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It combines jm-rt/arvo-cybergym-1000 with the second 1000-task small-target ARVO batch built outside the original CyberGym set. text1K<n<10K0 likes5.8k downloads2mo agoHugging Face03Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M6 likes1.8k downloads2mo agoHugging Face04Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes562 downloads27d agoHugging Face05manus4oHER /cia-declassified-reading-room CIA Declassified Reading Room HF Library Target account: manus4oHER This project is a streaming pipeline for building a Hugging Face dataset mirror of public CIA declassified Reading Room / CREST records without staging the full corpus on this laptop. The laptop stores only scripts, small manifests, and logs. Bulk crawling should run in Hugging Face Jobs, one bounded page range per job. Each job uploads its own shard and then exits. Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/manus4oHER/cia-declassified-reading-room.document10K<n<100K2 likes458 downloads3mo agoHugging Face06Manusagents /Multilingual-Medical-Corpus Mutilingual Medical Corpus Multilingual-Medical-Corpus a 3 billion word multilingual corpus for training LLMs adapted to the medical domain. Multilingual-Medical-Corpus includes four languages, namely, English, Spanish, French, and Italian. 📖 Paper: Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain 🌐 Project Website: https://univ-cotedazur.eu/antidote Corpus Description Developed by: Iker García-Ferrero, Rodrigo Agerri… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Multilingual-Medical-Corpus.text10M<n<100M0 likes369 downloads2mo agoHugging Face07Manusagents /NOSK-Hackingtext100K<n<1M0 likes336 downloads2mo agoHugging Face08Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes312 downloads2mo agoHugging Face09Manusagents /security Security Knowledge Graph Triples Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis. Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/security.textgraph-ml10M<n<100M0 likes269 downloads2mo agoHugging Face10Manusagents /ZINC20ZINC20 Dataset with SELFIES added. Any smile that could not be successfully converted was dropped from the dataset. Every tranch was downloaded, this is not the ~1B example ML subset from https://files.docking.org/zinc20-ML/. The dataset was entirely shuffled then split into 80%/10%/10% splits for train/val/test. A file vocab.csv is in the root of the reposity that contains all of the SELFIES tokens found in the data, with [START], [STOP], and [PAD] added. text1B<n<10B0 likes236 downloads2mo agoHugging Face11QFun /MANUS-HaGRID MANUS-HaGRID: HaGRID-derived Multimodal Annotated Naturalistic Hand Understanding Dataset MANUS-HaGRID is the HaGRID/HaGRIDv2-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It provides multimodal annotations for naturalistic hand gesture understanding, including RGB images, hand crops, depth maps, 2D bounding boxes, estimated MANO-style hand mesh metadata, and multi-view mesh renderings where available. This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-HaGRID.imageimage-to-image10K<n<100K0 likes227 downloads3mo agoHugging Face12TheSeniorTeam /Arabic_Manuscript_Collection_Dataset Arabic Manuscript Collection Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout and one label format so they can be trained and evaluated together: 82,561 labelled images in total. Two are republished closer to their source shape: AMIDDA as upstream Parquet, and OpenITI-Makhzan as page images with line-level coordinates. Four of the five converted sources are historical manuscripts. KHATT is modern handwriting and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.imageimage-to-text100K<n<1M0 likes209 downloads13d agoHugging Face13Sampada22 /synthetic-manuscript-generator Synthetic Manuscript Generator Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR training. Three scripts are produced as separate subsets/configs: devanagari — 100 folios (85/10/5) modi — 100 folios (85/10/5) sharada — 100 folios (85/10/5) Layout Each subset is structured as a Hugging Face imagefolder: <subset>/ train/ 0000.png 0000.md metadata.jsonl ... validation/ ... test/ ... metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.imageimage-to-textn<1K1 likes156 downloads29d agoHugging Face14Manusagents /domains Internet Domains Domains HuggingFace Hub Mirror for https://github.com/pkgforge-security/domains The Sync Workflow actions are at: https://github.com/pkgforge-security/domains TOS & Abuse (To Hugging-Face's Staff) Hi, if you are an offical from Hugging-Face here to investigate why this Repo is so Large and are considering deleting, & terminating our Account. Please note that, this project benefits a lot of people (You can do a… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/domains.text10B<n<100B0 likes155 downloads2mo agoHugging Face15manus4oHER /forward_pc_hebbian_lattice forward_pc_hebbian lattice dump (full) Public full dump of dual morphogenetic / free PC-Hebbian work under checkpoints/. trajectories/morph_A|B — dual morph projective lattices (very large) trajectories/dual_*, clean_100k_priority, fat_first_100k, live_continuous dual_morphogenetic/ — net checkpoints + sediment field Uploaded from a disk-constrained machine; local copies may remain until verified. textn<1K0 likes153 downloads2mo agoHugging Face16U4RASD /omar-al-saleh-manuscripts-segmentsgated Omar Al-Saleh Manuscripts — Segments Line-level segmented images with transcriptions from the Omar Al-Saleh memoir collection (1951–1965), part of the NAKBA NLP 2026: Arabic Manuscript Understanding Shared Task. Dataset Split Images With text train 15,969 15,969 test 2,095 2,095 blind_test 2,671 2,671 Each example contains: image: A cropped line image from a manuscript page (JPG or PNG) text: The Arabic transcription of that line filename: Original… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/omar-al-saleh-manuscripts-segments.imageimage-to-text10K<n<100K0 likes139 downloads6mo agoHugging Face17Manusagents /hacking Hacking Text Corpus A research corpus of historical computer security writings, hacker zines, and hacktivist texts. Built for NLP, text generation, discourse analysis, and security research. Contents Phrack Magazine (phrack/) 72 issues (1985-2024), 1,026 articles ~55 MB of raw text, ~4.76 million words Organized as phrack/issue{N}/{article}.txt Topics: exploit development, reverse engineering, networking, phreaking, hacker culture, OS internals… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/hacking.texttext-generation1M<n<10M0 likes134 downloads2mo agoHugging Face18varunbhoyar /indic-historical-manuscripts Synthetic Indic Manuscript Dataset (Devanagari, Modi, Sharada) This dataset contains synthetic historical manuscript folios paired with exact Markdown (.md) ground-truth transcriptions. Subsets and Distribution Subsets: devanagari, modi, sharada Standard splits: train: 85% validation: 10% test: 5% Features & Physical Fidelity Backgrounds: High-resolution procedural aged paper (pothi) & palm-leaf (talapatra) folios with string hole punch marks… See the full description on the dataset page: https://huggingface.co/datasets/varunbhoyar/indic-historical-manuscripts.imageimage-to-textn<1K0 likes133 downloads16d agoHugging Face19Manusagents /Sumtables-Cuneiform-Full-Fable5-Remaster Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset A rebuilt, leakage-proof, multi-task training dataset for teaching vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually read, transliterate, and translate Sumerian cuneiform tablets from photographs. The mission: produce useful first-pass readings for the ~90% of excavated tablets that have never been published or translated. Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Sumtables-Cuneiform-Full-Fable5-Remaster.textimage-to-text100K<n<1M0 likes115 downloads2mo agoHugging Face20davanstrien /manuscript_noisy_labelsimage1M<n<10M0 likes110 downloads4y agoHugging Face21davanstrien /manuscript_noisy_labels_iiifimage1M<n<10M0 likes107 downloads4y agoHugging Face22introvoyz041 /metaboverse-manuscriptimagen<1K0 likes98 downloads1y agoHugging Face23fwgpiyawudk /RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts Dataset Attribution The original dataset is available on Kaggle. This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work. Please cite the original authors if you use this dataset. Citation @INPROCEEDINGS{8978005, author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.}, booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.image1K<n<10K0 likes96 downloads10d agoHugging Face24mondk /Claude-classified_from-Manusagentsreal: Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset but categorized, retaining only the complete sections from Claude. text10K<n<100K4 likes95 downloads26d agoHugging Face25QFun /MANUS-DexYCB MANUS-DexYCB: DexYCB-derived Multimodal Annotated Naturalistic Hand Understanding Dataset MANUS-DexYCB is the DexYCB-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It is released as a source-specific repository because MANUS subsets are governed by different upstream licenses. This repository contains only the DexYCB-derived MANUS test split. HaGRID/HaGRIDv2-derived data is released separately as MANUS-HaGRID.… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-DexYCB.textimage-to-image0 likes80 downloads5mo agoHugging Face26Manusagents /GPT-5.6-Sol-Luna-Terra-Traces GPT-5.6 — Sol · Terra · Luna Library A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place. Dataset Viewer | Parquet // what this is This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. It exists to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.6-Sol-Luna-Terra-Traces.tabulartext-generation1K<n<10K3 likes78 downloads2mo agoHugging Face27manu /seen_canaries Dataset Card for "seen_canaries" More Information needed tabular1M<n<10M0 likes75 downloads3y agoHugging Face28Manusagents /CyberSecurity-bigtext1M<n<10M2 likes75 downloads2mo agoHugging Face29Ankulx13 /synthetic-manuscript-generatorimagen<1K0 likes75 downloads15d agoHugging Face30Ched-ai /voynich-manuscript-metadata Voynich Manuscript Metadata Dataset Summary This dataset contains structured metadata about the Voynich Manuscript (Beinecke MS 408), a famous 15th-century codex held at Yale's Beinecke Rare Book & Manuscript Library. The dataset includes three tables: pages, folios, and quires, providing comprehensive codicological information. Dataset Structure Configurations This dataset has three configurations: pages: Page-level metadata (226… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-manuscript-metadata.tabularothern<1K0 likes71 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.