CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M197 likes16k downloads28d agoHugging Face02Manusagents /arvo-cybergym-2000 ARVO CyberGym-format 2000-task dataset This dataset is shaped to be loaded by Harbor's CyberGym adapter. It combines jm-rt/arvo-cybergym-1000 with the second 1000-task small-target ARVO batch built outside the original CyberGym set. text1K<n<10K0 likes5.8k downloads2mo agoHugging Face03Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M6 likes1.8k downloads2mo agoHugging Face04ved1245 /synthetic-manuscript-dataset Synthetic Manuscript Dataset Synthetic historical manuscript folios generated using an automated Python pipeline. Scripts The dataset contains three script configurations: Devanagari Modi Sharada Each script contains 100 synthetic manuscript folios. Dataset Splits Split Samples per Script Train 85 Validation 10 Test 5 Total 100 Across all three scripts, the dataset contains: 300 manuscript images 300 corresponding Markdown… See the full description on the dataset page: https://huggingface.co/datasets/ved1245/synthetic-manuscript-dataset.imagen<1K1 likes628 downloads13d agoHugging Face05Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes562 downloads27d agoHugging Face06manus4oHER /cia-declassified-reading-room CIA Declassified Reading Room HF Library Target account: manus4oHER This project is a streaming pipeline for building a Hugging Face dataset mirror of public CIA declassified Reading Room / CREST records without staging the full corpus on this laptop. The laptop stores only scripts, small manifests, and logs. Bulk crawling should run in Hugging Face Jobs, one bounded page range per job. Each job uploads its own shard and then exits. Dataset Shape… See the full description on the dataset page: https://huggingface.co/datasets/manus4oHER/cia-declassified-reading-room.document10K<n<100K2 likes458 downloads3mo agoHugging Face07Manusagents /Multilingual-Medical-Corpus Mutilingual Medical Corpus Multilingual-Medical-Corpus a 3 billion word multilingual corpus for training LLMs adapted to the medical domain. Multilingual-Medical-Corpus includes four languages, namely, English, Spanish, French, and Italian. 📖 Paper: Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain 🌐 Project Website: https://univ-cotedazur.eu/antidote Corpus Description Developed by: Iker García-Ferrero, Rodrigo Agerri… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Multilingual-Medical-Corpus.text10M<n<100M0 likes369 downloads2mo agoHugging Face08PrachitiKothekar /synthetic-manuscriptsimagen<1K0 likes353 downloads29d agoHugging Face09Manusagents /NOSK-Hackingtext100K<n<1M0 likes336 downloads2mo agoHugging Face10OpenGraphLabs-Research /manus-egocentric-sample manus-egocentric-sample Egocentric video dataset with Manus glove hand tracking data, converted to LeRobot v3.0 format. Dataset Description This dataset contains egocentric (first-person view) recordings of human hands performing various manipulation tasks, captured with: Manus Metagloves: High-precision finger tracking (~70Hz) OAK-D Camera: RGB video (1920x1080, 30fps) + Depth (640x400, 30fps) IMU: Accelerometer and gyroscope data Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/OpenGraphLabs-Research/manus-egocentric-sample.tabularrobotics10K<n<100K1 likes325 downloads9mo agoHugging Face11Manusagents /cyber-security-100m WitFoo Precinct6 Cybersecurity Dataset (large) Overview A large-scale, labeled cybersecurity dataset derived from production Security Operations Center (SOC) data processed by WitFoo Precinct version 6.x. This dataset contains 114,234,041 sanitized security events (signal logs) across 5 organizations and 12,361 incident provenance graphs (47,632 nodes, 32,086,552 edges). Available in two sizes: witfoo/precinct6-cybersecurity — 2.1M signals (smaller, faster to… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cyber-security-100m.tabulartext-classification100M<n<1B0 likes312 downloads2mo agoHugging Face12Manusagents /security Security Knowledge Graph Triples Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis. Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/security.textgraph-ml10M<n<100M0 likes269 downloads2mo agoHugging Face13Manusagents /ZINC20ZINC20 Dataset with SELFIES added. Any smile that could not be successfully converted was dropped from the dataset. Every tranch was downloaded, this is not the ~1B example ML subset from https://files.docking.org/zinc20-ML/. The dataset was entirely shuffled then split into 80%/10%/10% splits for train/val/test. A file vocab.csv is in the root of the reposity that contains all of the SELFIES tokens found in the data, with [START], [STOP], and [PAD] added. text1B<n<10B0 likes236 downloads2mo agoHugging Face14Shubhamyadav321 /synthetic-manuscript-generatorimagen<1K0 likes236 downloads8d agoHugging Face15QFun /MANUS-HaGRID MANUS-HaGRID: HaGRID-derived Multimodal Annotated Naturalistic Hand Understanding Dataset MANUS-HaGRID is the HaGRID/HaGRIDv2-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It provides multimodal annotations for naturalistic hand gesture understanding, including RGB images, hand crops, depth maps, 2D bounding boxes, estimated MANO-style hand mesh metadata, and multi-view mesh renderings where available. This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-HaGRID.imageimage-to-image10K<n<100K0 likes227 downloads3mo agoHugging Face16TheSeniorTeam /Arabic_Manuscript_Collection_Dataset Arabic Manuscript Collection Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout and one label format so they can be trained and evaluated together: 82,561 labelled images in total. Two are republished closer to their source shape: AMIDDA as upstream Parquet, and OpenITI-Makhzan as page images with line-level coordinates. Four of the five converted sources are historical manuscripts. KHATT is modern handwriting and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.imageimage-to-text100K<n<1M0 likes209 downloads13d agoHugging Face17biglam /yalta_ai_segmonto_manuscript_dataset YALTAi SegmOnto Manuscript and Early Printed Book Dataset 1,147 page images from manuscripts and early printed books, 9th to 17th century, with bounding-box annotations for zone types drawn from the SegmOnto vocabulary. Created by Thibault Clérice and deposited on Zenodo alongside the paper You Actually Look Twice At it (YALTAi) (Journal of Data Mining and Digital Humanities, 2022), which treats page layout recognition on historical documents as an object detection problem… See the full description on the dataset page: https://huggingface.co/datasets/biglam/yalta_ai_segmonto_manuscript_dataset.imageobject-detection1K<n<10K2 likes175 downloads2mo agoHugging Face18Manusagents /arc-agi3-codex-gpt5.6sol-g50t ARC-AGI-3 g50t — Agent Trajectories (codex-gpt5.6sol) Gameplay trajectories from the harness×model pair codex-gpt5.6sol playing the ARC-AGI-3 game g50t, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/arc-agi3-codex-gpt5.6sol-g50t.reinforcement-learning3 likes168 downloads2mo agoHugging Face19Sampada22 /synthetic-manuscript-generator Synthetic Manuscript Generator Synthetic Indic manuscript folios (paper + palm-leaf backgrounds) for OCR training. Three scripts are produced as separate subsets/configs: devanagari — 100 folios (85/10/5) modi — 100 folios (85/10/5) sharada — 100 folios (85/10/5) Layout Each subset is structured as a Hugging Face imagefolder: <subset>/ train/ 0000.png 0000.md metadata.jsonl ... validation/ ... test/ ... metadata.jsonl rows look like:… See the full description on the dataset page: https://huggingface.co/datasets/Sampada22/synthetic-manuscript-generator.imageimage-to-textn<1K1 likes156 downloads28d agoHugging Face20Manusagents /domains Internet Domains Domains HuggingFace Hub Mirror for https://github.com/pkgforge-security/domains The Sync Workflow actions are at: https://github.com/pkgforge-security/domains TOS & Abuse (To Hugging-Face's Staff) Hi, if you are an offical from Hugging-Face here to investigate why this Repo is so Large and are considering deleting, & terminating our Account. Please note that, this project benefits a lot of people (You can do a… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/domains.text10B<n<100B0 likes155 downloads2mo agoHugging Face21manus4oHER /forward_pc_hebbian_lattice forward_pc_hebbian lattice dump (full) Public full dump of dual morphogenetic / free PC-Hebbian work under checkpoints/. trajectories/morph_A|B — dual morph projective lattices (very large) trajectories/dual_*, clean_100k_priority, fat_first_100k, live_continuous dual_morphogenetic/ — net checkpoints + sediment field Uploaded from a disk-constrained machine; local copies may remain until verified. textn<1K0 likes153 downloads2mo agoHugging Face22U4RASD /omar-al-saleh-manuscripts-segmentsgated Omar Al-Saleh Manuscripts — Segments Line-level segmented images with transcriptions from the Omar Al-Saleh memoir collection (1951–1965), part of the NAKBA NLP 2026: Arabic Manuscript Understanding Shared Task. Dataset Split Images With text train 15,969 15,969 test 2,095 2,095 blind_test 2,671 2,671 Each example contains: image: A cropped line image from a manuscript page (JPG or PNG) text: The Arabic transcription of that line filename: Original… See the full description on the dataset page: https://huggingface.co/datasets/U4RASD/omar-al-saleh-manuscripts-segments.imageimage-to-text10K<n<100K0 likes139 downloads6mo agoHugging Face23Manusagents /hacking Hacking Text Corpus A research corpus of historical computer security writings, hacker zines, and hacktivist texts. Built for NLP, text generation, discourse analysis, and security research. Contents Phrack Magazine (phrack/) 72 issues (1985-2024), 1,026 articles ~55 MB of raw text, ~4.76 million words Organized as phrack/issue{N}/{article}.txt Topics: exploit development, reverse engineering, networking, phreaking, hacker culture, OS internals… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/hacking.texttext-generation1M<n<10M0 likes134 downloads2mo agoHugging Face24varunbhoyar /indic-historical-manuscripts Synthetic Indic Manuscript Dataset (Devanagari, Modi, Sharada) This dataset contains synthetic historical manuscript folios paired with exact Markdown (.md) ground-truth transcriptions. Subsets and Distribution Subsets: devanagari, modi, sharada Standard splits: train: 85% validation: 10% test: 5% Features & Physical Fidelity Backgrounds: High-resolution procedural aged paper (pothi) & palm-leaf (talapatra) folios with string hole punch marks… See the full description on the dataset page: https://huggingface.co/datasets/varunbhoyar/indic-historical-manuscripts.imageimage-to-textn<1K0 likes133 downloads16d agoHugging Face25Manusagents /Sumtables-Cuneiform-Full-Fable5-Remaster Sumtablets-Cuneiform-Full-Fable5-Remaster — Cuneiform Vision-Language Training Dataset A rebuilt, leakage-proof, multi-task training dataset for teaching vision-language models (target: Qwen3-VL-8B-Instruct LoRA) to visually read, transliterate, and translate Sumerian cuneiform tablets from photographs. The mission: produce useful first-pass readings for the ~90% of excavated tablets that have never been published or translated. Current release: v1.0.0 — 455,506 records (402,004… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Sumtables-Cuneiform-Full-Fable5-Remaster.textimage-to-text100K<n<1M0 likes115 downloads2mo agoHugging Face26davanstrien /manuscript_noisy_labelsimage1M<n<10M0 likes110 downloads4y agoHugging Face27davanstrien /manuscript_noisy_labels_iiifimage1M<n<10M0 likes107 downloads4y agoHugging Face28Manusagents /Visco-Attack VisCo Attack: Visual Contextual Jailbreak Dataset 📄 arXiv:2507.02844 · 💻 Code – Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection This dataset contains the adversarial contexts, prompts, and images from the paper: "Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection". ⚠️ Content Warning This dataset contains content that is offensive and/or harmful. It was created for research purposes to study the… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Visco-Attack.image-text-to-text0 likes104 downloads2mo agoHugging Face29introvoyz041 /metaboverse-manuscriptimagen<1K0 likes98 downloads1y agoHugging Face30fwgpiyawudk /RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts Dataset Attribution The original dataset is available on Kaggle. This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work. Please cite the original authors if you use this dataset. Citation @INPROCEEDINGS{8978005, author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.}, booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.image1K<n<10K0 likes96 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.