datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sqa3d-community-mirrorscanqa-community-mirrorovos-wake-word-bench-picovoice-smart-mirror
OVOS wake_word bench — picovoice-smart-mirror
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.MirrorAPI-Bench
MirrorAPI-Bench dataset
This dataset contains the following resources:
The testing data for MirrorAPI and MirrorAPI-Cache:
test_sft/: The testing data for MirrorAPI with SFT prompts.
test_cot/: The testing data for MirrorAPI with CoT prompts.
reference/: The references for testing data without JSON format. These references are used for evaluation. Note that cache.jsonl has only 200 random samples from test_cache.json.
test_cache.json: The testing data for MirrorAPI-Cache.… See the full description on the dataset page: https://huggingface.co/datasets/stabletoolbench/MirrorAPI-Bench.backln-guest-post-quality-public-mirror
Backln Guest Post Quality Public Mirror
Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus.
Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content.
Schema
label: one of published, manual_review, rejected.
source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.MirrorAPI-Training
MirrorAPI training dataset
This dataset contains the training data for MirrorAPI and MirrorAPI-Cache:
train_sft.json, train_cot.json, train_augment.json: The training data for MirrorAPI .
train_cache.json: The training data for MirrorAPI-Cache.
alfred-json-mirror
ALFRED JSON Mirror
This mirror contains the official ALFRED lite trajectory JSONs repackaged for em-eval.
Source:
Official lite archive: https://ai2-vision-alfred.s3-us-west-2.amazonaws.com/json_2.1.0.7z
Upstream repository: askforalfred/alfred
Files:
tests_seen.jsonl: 483 trajectories
tests_unseen.jsonl: 488 trajectories
train.jsonl: 6574 trajectories
valid_seen.jsonl: 251 trajectories
valid_unseen.jsonl: 255 trajectories
Each row is the original traj_data.json payload with one… See the full description on the dataset page: https://huggingface.co/datasets/thomas-yanxin/alfred-json-mirror.mirror-sql
MIRROR-SQL
Provenance-Controlled Database Environments for Text-to-SQL Agents.
13 PostgreSQL environments · 176 tables · 2762 columns · 390 annotated question/SQL pairs.
MIRROR-SQL takes the opposite approach to contamination from every other text-to-SQL corpus.
Spider and BIRD sample public databases. BEAVER uses real private warehouses that cannot be
redistributed. LiveSQLBench out-runs leakage temporally by rebuilding from changing sources.
MIRROR-SQL instead purpose-builds… See the full description on the dataset page: https://huggingface.co/datasets/1digitaldesign/mirror-sql.black_mirror_scripts_S1-5Black Mirror Scripts Dataset (Seasons 1-5)
This dataset, titled 'black_mirror_scripts_S1-5.csv', contains the meticulously compiled transcripts of the critically acclaimed anthology series Black Mirror, covering Seasons 1 through 5. Each entry in this dataset is categorized by unique identifiers including Script ID, Title, Scene, Dialogue, and Timestamp, making it an ideal resource for natural language processing tasks, script analysis, sentiment analysis, and more.
Dataset Composition
Our… See the full description on the dataset page: https://huggingface.co/datasets/tmobley96/black_mirror_scripts_S1-5.beacon3d-qa-mirror
Beacon3D QA Mirror
Community mirror generated from the public beacon-3d/Beacon3D GitHub repository.
Contents:
train.json: merged QA annotations across scannet, 3rscan, and multiscan
system_prompt.json: official Beacon3D judge prompt
This mirror intentionally contains benchmark annotations only. It does not redistribute raw ScanNet / 3RScan / MultiScan scene assets.
mirror-recogna-nlp__UltrachatBR
UltrachatBR: Um Dataset em Português baseado no Ultrachat
O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa.
Processo de Tradução
O processo de tradução foi realizado utilizando a API do… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-recogna-nlp__UltrachatBR.mirror-meta-math__MetaMathQAView the project page:
https://meta-math.github.io/
see our paper at https://arxiv.org/abs/2309.12284
Note
All MetaMathQA data are augmented from the training sets of GSM8K and MATH.
None of the augmented data is from the testing set.
You can check the original_question in meta-math/MetaMathQA, each item is from the GSM8K or MATH train set.
Model Details
MetaMath-Mistral-7B is fully fine-tuned on the MetaMathQA datasets and based on the powerful Mistral-7B model.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-meta-math__MetaMathQA.mirror-threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-threat-intelligence-dataset.mirror-Code_Vulnerability_Security_DPO
Cybernative.ai Code Vulnerability and Security Dataset
Dataset Description
The Cybernative.ai Code Vulnerability and Security Dataset is a dataset of synthetic Data Programming by Demonstration (DPO) pairs, focusing on the intricate relationship between secure and insecure code across a variety of programming languages. This dataset is meticulously crafted to serve as a pivotal resource for researchers, cybersecurity professionals, and AI developers who are keen… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Code_Vulnerability_Security_DPO.mirror-SWE-Next-SFT-Trajectories
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
SWE-Next SFT Trajectories
SWE-Next SFT Trajectories is the supervised fine-tuning dataset released with SWE-Next: Scalable Real-World Software Engineering Tasks for Agents. It contains 3,693 ShareGPT-style multi-turn training examples collected from expert agent rollouts on 2,308 execution-grounded SWE tasks synthesized from real merged pull requests.
The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-SWE-Next-SFT-Trajectories.mirror-tech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-tech-docs.mirror-aware-inference
mirror-aware-inference
A framework to measure how much of an output originates from user input (prompt), training data biases, inductive biases from model architecture, or novel composition of retrieved information.
1. Introduction
This project implements a Mirror-Aware Inference that performs "bias-tracking" by analyzing the model's internal state during generation.
The scripts perform a series of backpropagation passes to measure the influence of different components… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/mirror-aware-inference.mirror-APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-APIGen-MT-5k.mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset
🤖 Agentic Coding CoT Dataset
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant Data… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset.mirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1
Dataset Description:
This is a RL dataset for general function-calling by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Nemotron-RL-Agentic-Function-Calling-Pivot-v1.mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1
🤖 Agentic Coding CoT Dataset v1.1
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.magic-mirror-flux-datamirror-terraform_sec
Terraform Security Dataset
A comprehensive dataset of 62,406 Terraform projects analyzed for security vulnerabilities using tfsec. This dataset is designed for training Large Language Models (LLMs) to understand, identify, and fix security issues in Terraform infrastructure-as-code.
📊 Dataset Overview
Total Examples: 62,406 Terraform projects
Secure Projects: 43,575 (69.8%)
Insecure Projects: 18,831 (30.2%)
Format: JSONL (JSON Lines)
Task: Security analysis and… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-terraform_sec.mirror-Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Trendyol-Cybersecurity-Instruction-Tuning-Dataset.mirror-hermes-agent-traces-filtered
Hermes Agent Reasoning Traces - Quality Filtered
A structurally filtered subset of lambda/hermes-agent-reasoning-traces, pruned from 7,646 to 3,679 rows using automated quality analysis targeting reasoning depth, structural integrity, and tool-call validity.
Why This Matters for Agent Training
Most agentic datasets teach models what tool to call but not how to reason about tool selection. The difference matters in production: an agent that dispatches tools without… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-hermes-agent-traces-filtered.mirror-witness-atlas
Mirror-Witness Atlas
The Kreuzer–Skarke landscape carries an involution: every reflexive polytope has a polar
dual, Batyrev's construction makes dual pairs into mirror pairs, and the exchange swaps
(h¹¹, h¹²) ↔ (h¹², h¹¹). Every geometry has a partner that returns it whole. This atlas
reads that structure through a relational lens — understanding and recognition on the
same ordered pair — and ships, for every claim it makes, either an exact recomputation
or an explicit citation… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/mirror-witness-atlas.beacon3d-grounding-mirror
Beacon3D Grounding Mirror
Community mirror generated from the public beacon-3d/Beacon3D GitHub repository.
Contents:
train.json: merged grounding annotations across scannet, 3rscan, and multiscan
The official Beacon3D grounding task evaluates object-id prediction and chain/object accuracy. This mirror does not redistribute raw ScanNet / 3RScan / MultiScan scene assets.
syndata-rrd-mirrormirrorqwen2.5-0.5B-gsm8k-PRM-data-ST-2mirror-glaive-function-calling-v2
