datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pile-val-backupThis is a backup for the pile val dataset downloaded from here: https://the-eye.eu/public/AI/pile/val.jsonl.zst
Please respect the original license of the dataset.
freshwatr-map-packs
FRESHWATR offline map packs
Vector tile packs (PMTiles, OpenMapTiles schema, z0–14) for the FRESHWATR iOS
fishing app's automatic offline maps — one pack per US state plus DC.
Map data © OpenStreetMap
contributors, licensed under the Open Database License (ODbL).
Tiles generated with Planetiler
using the OpenMapTiles schema.
Extracts courtesy of Geofabrik.
catalog.json lists every pack with byte size and SHA-256; the app verifies
each download against it before install.
mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony
MITRE+NVD+ExploitDB Dataset (Alpaca/ChatML/Harmony)
A dataset for training AI assistants/agents on vulnerability analysis and pentesting Q&A. It is built by the pentestds pipeline, which fetches and merges data from MITRE CVE, NVD (CVSS enrichment), ExploitDB, and a small set of HuggingFace datasets. Provenance is recorded for every entry, and the pipeline emits Alpaca, ChatML, and Harmony JSONL files.
Dataset Summary
This dataset is designed for training AI agents to… See the full description on the dataset page: https://huggingface.co/datasets/jason-oneal/mitre-stix-cve-exploitdb-dataset-alpaca-chatml-harmony.Fable-5-traces
Glint Research Dataset Card
Fable 5 Pi Agent Traces
A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation.
Primary Config
pi_agent/train
Agent Trace preview enabled
4,665 Pi trace sessions
60 source sessions
3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/MithralJ/Fable-5-traces.mitre-attack-synthetic-scenarios
MITRE ATT&CK Synthetic Scenario Logs v3.0
Expanded Dataset: 30 scenarios × 8 events = 240 synthetic events
Axis
Coverage
Environment
endpoint, cloud, SaaS, identity, CI/CD, OT/IoT
Actor Type
external_apt, ransomware, insider, compromised_vendor, careless_admin, automated_threat
Intent
exfiltration, impact, fraud, persistence, reconnaissance, cryptomining, espionage
Detection Source
EDR, IAM, SIEM, DLP, DNS, proxy, cloud_audit, email_gateway, CASB, NDR, PAM, firewall… See the full description on the dataset page: https://huggingface.co/datasets/koushikcs09/mitre-attack-synthetic-scenarios.reclaim
RECLAIM
RECLAIM is a benchmark of 100 NeurIPS 2025 papers for measuring whether an AI agent can
reproduce a published machine learning result. For each paper the benchmark fixes, before
any agent starts, the central claim, the single result a run has to recover, what counts as
a successful reproduction, the artifacts the authors released, a difficulty tier, and an
audited GPU-hour estimate.
This repository holds the benchmark rows. The agent harness, the grading rubric, the… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/reclaim.mit67_resnet50_places365-oof-featuresmit67_swin_tiny-oof-featuresprotein-secondary-structure-netsurfp
NetSurfP-3.0 Secondary-Structure Splits
This dataset repo contains NetSurfP-derived protein secondary-structure labels
converted for Protein-I-JEPA probe training and evaluation.
Source page: https://services.healthtech.dtu.dk/services/NetSurfP-3.0/5-Dataset.php
Profile: hhblits
Labels are Q3 per-residue labels:
H: helix
E: beta strand
C: coil/other
.: ignored residue for loss and accuracy
Splits
Split
Rows
JSONL
TSV
train
10348
train.jsonl
tsv/train.tsv… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/protein-secondary-structure-netsurfp.mit67_resnet50_places365_oof_featuresgemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.security-attacks-MITREprompt-injection-dataset
Prompt Injection Dataset
A labeled dataset of benign prompts and prompt-injection attempts for training, evaluating, and experimenting with first-line prompt-injection detection for LLM, RAG, and agentic AI applications.
This dataset supports the ai-mitra/prompt-injection-detector model.
Source code and training pipeline:
https://github.com/tg-mitra/prompt-injection-detector
📊 Dataset Summary
Property
Value
Version
1.0.0
Training examples
1,130… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-injection-dataset.mit67_swin_tiny_oof_featuresMIT-10M
MIT-10M
Paper: https://aclanthology.org/2025.coling-main.346/
Introduction:
Image Translation (IT) holds immense potential across diverse domains, enabling the translation of textual content within images into various languages.
However, existing datasets often suffer from limitations in scale, diversity, and quality, hindering the development and evaluation of IT models.
To address this issue, we introduce MIT-10M, a large-scale parallel corpus of multilingual image translation… See the full description on the dataset page: https://huggingface.co/datasets/liboaccn/MIT-10M.mit67_gnn-oof-featuresllm-router-dataset
llm-router dataset
Training data for ai-mitra/llm-router,
a prompt task-classifier used by the
llm-router Python
package to route prompts to the best-fit LLM in agentic AI systems.
Each row is a prompt labeled with the task category it belongs to.
Labels
simple, coding, reasoning, security, summarization
Files
File
Rows
Purpose
training_data.jsonl
1700 (340/label)
Used to train the classifier: 35 hand-written examples per category plus… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/llm-router-dataset.prompt-slimmer-slm
Prompt Slimmer SLM — Demo Dataset
Synthetic examples for experimenting with prompt rewriting and sentence selection. Exported without changing the examples or their original splits from the shared GitHub codebase.
Model · Project page
Configuration
Train
Validation
Test
Purpose
rewrites-expanded (default)
41
2
2
Expanded rewriting dataset: 45 examples
rewrites
9
2
2
Original dataset used by the first adapter
selector
256
64
64
KEEP/DROP labels for source spans… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-slimmer-slm.mit67_gnn_oof_featuresMitre_Attacks_Framework_Dataset
MITRE ATT&CK Enterprise Dataset
Overview
This dataset provides a comprehensive collection of MITRE ATT&CK Enterprise techniques (v14.1) in JSONL format, designed for cybersecurity professionals, red teams, and threat hunters.
Each entry maps to a specific ATT&CK technique, including its ID, name, description, real-world example, and source.
The dataset is structured for seamless integration into security tools such as SIEMs, threat intelligence platforms, or custom red… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Mitre_Attacks_Framework_Dataset.mit67_vit_base-oof-featuresBehaviouralLoC-Mitigation
BehaviouralLoC-Mitigation
BehaviouralLoC-Mitigation contains the supervised fine-tuning corpora used for
misaligned-motive mitigation in A Behavioural Framework for Predicting and
Understanding Loss of Control in Frontier Artificial Intelligence Systems.
The corpus covers five motive aspects. Following the paper, examples were
generated in distribution with Qwen3.5-27B, and the prompts were augmented by
safety experts.
Configurations
The three paper configurations… See the full description on the dataset page: https://huggingface.co/datasets/T-STAR-Lab/BehaviouralLoC-Mitigation.formulae__mita-v1.1-7b-2-24-2025-details
Dataset Card for Evaluation run of formulae/mita-v1.1-7b-2-24-2025
Dataset automatically created during the evaluation run of model formulae/mita-v1.1-7b-2-24-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-v1.1-7b-2-24-2025-details.egotouch-annotations-v1-leftfix
egotouch-annotations-v1-leftfix
Annotations only. This repository does not contain images or video.
This release corrects the left-hand MANO rotation convention in
egotouch-annotations-v1.
It keeps the original episode structure, instructions, tactile data, and sample index.
Item
Count
Episodes
111,159
Frames and training samples
3,687,389
Repaired left-hand episodes
54,373
Repaired left-hand frames
1,781,845
Unchanged right-only episodes
56,786… See the full description on the dataset page: https://huggingface.co/datasets/MIT-Media-Lab/egotouch-annotations-v1-leftfix.vila-datasetwikidata-parallel-descriptions-en-ja
Wikidata parallel descriptions en-ja
Parallel corpus for machine translation generated from wikidata dump (2024-05-06).
Currently we processed only English/Japanese pair.
The jsonl file is ready-to-train by Hugging Face transformers trainer for translation tasks.
Dataset Details
https://www.wikidata.org/wiki/Wikidata:Database_download
Dataset Creation
As Wikidata description field does not represent exact direct translation, filtering is required for… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/wikidata-parallel-descriptions-en-ja.Mitre-ATTACK-reasoning-datasetformulae__mita-elite-v1.1-7b-2-25-2025-details
Dataset Card for Evaluation run of formulae/mita-elite-v1.1-7b-2-25-2025
Dataset automatically created during the evaluation run of model formulae/mita-elite-v1.1-7b-2-25-2025
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/formulae__mita-elite-v1.1-7b-2-25-2025-details.neurips-2025-paperswithcode-artifactsgraph-preflexor-grpo-benchmark
Citation
@misc{pal2026graphnativereinforcementlearningenables,
title={Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination},
author={Subhadeep Pal and Shashwat Sourav and Tirthankar Ghosal and Markus J. Buehler},
year={2026},
eprint={2607.00924},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2607.00924},
}
