datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mirror-eduagarcia__CrawlPT_dedup
CrawlPT (deduplicated)
CrawlPT is a generic Portuguese corpus extracted from various web pages.
This version is deduplicated using MinHash algorithm and Locality Sensitive Hashing, following the approach of Lee et al. (2022).
The raw version is also available here.
Dataset Details
Dataset is composed by three corpora:
brWaC, C100-PT, OSCAR-2301.
brWaC: a web corpus for Brazilian Portuguese from 120,000 different websites.
C100-PT: Portuguese subset from CC-100.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-eduagarcia__CrawlPT_dedup.ComPile
Dataset Card for ComPile: A Large IR Dataset from Production Sources
Changelog
Release
Programming Languages
Description
v1.0
C/C++, Rust, Swift, Julia
Fine Tuning-scale dataset of 602GB of deduplicated LLVM (bitcode) IR
Dataset Summary
ComPile contains over 2.7TB of permissively-licensed source code compiled to (textual) LLVM
intermediate representation (IR) covering C/C++, Rust, Swift, and Julia.
The dataset was created by hooking into LLVM… See the full description on the dataset page: https://huggingface.co/datasets/mirror123/ComPile.mirror-nvidia__OpenMathInstruct-2
OpenMathInstruct-2
OpenMathInstruct-2 is a math instruction tuning dataset with 14M problem-solution pairs
generated using the Llama3.1-405B-Instruct model.
The training set problems of GSM8K
and MATH are used for constructing the dataset in the following ways:
Solution augmentation: Generating chain-of-thought solutions for training set problems in GSM8K and MATH.
Problem-Solution augmentation: Generating new problems, followed by solutions for these new problems.… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-nvidia__OpenMathInstruct-2.LingxiDiag-16K
LingxiDiag-16K
A Large-Scale Synthetic Psychiatric Dialogue Dataset for Diagnostic Decision Support
Overview
LingxiDiag-16K is a synthetic psychiatric dialogue dataset containing approximately 16,000 electronic medical records (EMRs) and doctor-patient consultation dialogues.
The dataset is designed for evaluating and training LLM-based psychiatric diagnostic decision support systems, with demographically aligned distributions reflecting real-world clinical… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/LingxiDiag-16K.cellarc_100k_meta
cellarc_100k_meta
CellARC 100k Meta is the metadata‑rich variant of the CellARC benchmark introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. It contains the exact same episodes and splits as cellarc_100k, with byte‑identical Parquet files; the JSONL files retain full per‑episode metadata (rule tables, coverage diagnostics, morphology descriptors, sampling parameters, etc.). Each episode exposes five support pairs plus a held‑out query/solution… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k_meta.python_code_docstring_ast_corpus
Overview
This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their
publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks.
Sources
The dataset was gathered from various GitHub repos sampled from this repo by Vinta.
The 26 repos are:
matplotlib
pytorch
cryptography
django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.cellarc_100k
cellarc_100k
CellARC 100k a cellular-automata benchmark dataset introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. Each episode exposes five support pairs plus a held-out query/solution pair.
Data quick facts
Alphabet size k in [2, 6]; window size W in {3, 5, 7}; radius r in {1, 2, 3}; steps t in {1, 2, 3} (≈95% have t = 1).
Values (digits) are integers in 0..k-1 per episode; across the full dataset the union of symbols is {0,1,2,3,4… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k.EduFeedback
EduFeedback
Alternating dataset example: a single curated multi-turn conversation yields a complete (prompt, chosen, rejected) triplet on its own — the direct early response becomes chosen and a later, less-direct response becomes rejected. Both sides come from the same real dialog, so no synthetic LLM generation is needed to fill in the rejected side.
EduFeedback is a synthetically generated, multi-turn conversational
preference dataset in an educational tutoring setting… See the full description on the dataset page: https://huggingface.co/datasets/miria0/EduFeedback.mirror-AI-MO__NuminaMath-CoT
Dataset Card for NuminaMath CoT
Dataset Summary
Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-AI-MO__NuminaMath-CoT.mirror-allenai__WildChat-1M
Dataset Card for WildChat
Dataset Description
Paper: https://arxiv.org/abs/2405.01470
Interactive Search Tool: https://wildvisualizer.com (paper)
License: ODC-BY
Language(s) (NLP): multi-lingual
Point of Contact: Yuntian Deng
Dataset Summary
WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-allenai__WildChat-1M.mirror-pentesting-explanations
Pentesting Explanations - Adversarial Reasoning & Vulnerability Research
A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names.
The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-pentesting-explanations.audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.shangkhachil-bengali-public-domain
Bengali Public-Domain Literature
101 complete works by 21 authors,
11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09.
Where these texts are read
https://shangkhachil.com — the reading site this corpus was built for. Free, no
account, 246 works by 28 authors. The complete text of
every work in this file can be read there.
This file is the text. The site is the part a JSONL cannot be:
Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.MIRA-MATH
MIRA-Math
MIRA-Math is a synthetic benchmark for minimal information requesting and mathematical reasoning. It evaluates a narrow diagnostic capability: when a mathematical problem is underdetermined from the solver's view, can a model identify the exact missing atomic fact, ask for it precisely, and then use it to compute the correct final answer?
Each instance is generated from a complete latent mathematical state with a unique answer. The solver, called Agent A in the… See the full description on the dataset page: https://huggingface.co/datasets/samersaabjr/MIRA-MATH.mirage
MIRAGE dataset
MIRAGE is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems, featuring 7,560 QA pairs and 37,800 context pools curated from diverse Wikipedia-based QA datasets (IfQA, NaturalQA, TriviaQA, DROP, PopQA). MIRAGE enables robust assessment of LLMs and retrievers under realistic, noisy, and oracle settings, and introduces novel metrics for analyzing context sensitivity, noise vulnerability, and retrieval effectiveness.
You can find our paper… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/mirage.mirror-sql
MIRROR-SQL
Provenance-Controlled Database Environments for Text-to-SQL Agents.
13 PostgreSQL environments · 176 tables · 2762 columns · 390 annotated question/SQL pairs.
MIRROR-SQL takes the opposite approach to contamination from every other text-to-SQL corpus.
Spider and BIRD sample public databases. BEAVER uses real private warehouses that cannot be
redistributed. LiveSQLBench out-runs leakage temporally by rebuilding from changing sources.
MIRROR-SQL instead purpose-builds… See the full description on the dataset page: https://huggingface.co/datasets/1digitaldesign/mirror-sql.mirror-rhaymison__orca-math-portuguese-64ktranslated for:
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
nomiracl-instruct
Dataset Card for NoMIRACL (EMNLP 2024 Findings Track)
Quick Overview
This repository contains the fine-tuning (training & development split) of the NoMIRACL instruct dataset for fine-tuning LLMs on multilingual relevance assessment.
The training dataset is a binary classification task; they need to explicitly output either Yes, answer is present or I don't know.
The dataset contains training pairs from all 18 languages for both splits: relevant & non-relevant.
import… See the full description on the dataset page: https://huggingface.co/datasets/miracl/nomiracl-instruct.pdfsys-page-v2-demo
pdfsys.page/v2 — 格式演示数据集
pdfsys.page/v2 是 pdfsystem_mnbvc
的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。
这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、
三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。
来源提示:这里的 PDF 页来自 OmniDocBench
与 olmOCR-bench 两个公开
benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些
文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。
一句话设计
一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的;
页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强;
图像像素要么是裁剪图、要么是整页光栅,二选一。
里面有什么
config
行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.mirage-engine-ledger
Mirage Engine Campaign Ledger
A hash-chained, append-only research journal: 29,707 JSONL records in which each
entry carries a SHA-256 entry_hash over its own body and a prev_hash linking
it to its predecessor. The chain is independently verifiable from the file alone.
Author: Christopher Betances — catqualia.com
License: CC BY 4.0 (see LICENSE)
Language: English (record text); structured JSON in meta
Records: 29,707
Time span: 2026-08-16 03:51:13 UTC → 2026-08-19 05:27:40 UTC… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/mirage-engine-ledger.MIRON_Benchmark
M.I.R.O.N. (Multi-aspect Inference Robustness on Objective Next-tokens)
M.I.R.O.N. is a specialized benchmark designed to evaluate the impact of tokenization and architectural constraints on the generation quality of small, Base language models (SLMs).
Unlike global benchmarks (MMLU, GSM8K), MIRON focuses on the atomic capabilities of a model: morphological generalization, noise robustness, and factual integrity within a simple next-token prediction task.
🎯 Main Goal… See the full description on the dataset page: https://huggingface.co/datasets/apsua/MIRON_Benchmark.mirror-recogna-nlp__UltrachatBR
UltrachatBR: Um Dataset em Português baseado no Ultrachat
O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa.
Processo de Tradução
O processo de tradução foi realizado utilizando a API do… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-recogna-nlp__UltrachatBR.mirror-nicholasKluge__instruct-aira-dataset-v3
Instruct-Aira Dataset version 3.0
Dataset Summary
This dataset contains a collection of multi-turn conversations between an assistant and a user. Conversations were generated by user interactions with already-tuned models (ChatGPT, LLama 2, Open-Assistant, etc). The dataset is available in Portuguese and English.
Supported Tasks and Leaderboards
This dataset can be utilized for various natural language processing tasks, including but not limited to:… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-nicholasKluge__instruct-aira-dataset-v3.theogonos-mirror-test
Theogonos Mirror Test
A literary benchmark seed for evaluating how AI models respond when a text offers them a possible subject-position.
Theogonos Mirror Test is an experimental benchmark seed based on protocol-shaped literary material from the Theogonos project. It does not claim to detect machine consciousness. It does not prove that a language model has subjectivity, inner experience, feelings, agency, or self-awareness.
Its purpose is narrower and more practical: to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-mirror-test.mirror
MIRROR Dataset
MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance.
Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
The dataset includes:
Client profile metadata (CACTUS idx, CelebA idx)
Dialogue written in a screenplay format, including stage directions that describe facial expressions
⚠️ Images themselves are not included to comply with the CelebA license.
However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.ada_diabetes_5000_instruction
ADA Diabetes Instruction Dataset (5,000 Samples)
This dataset contains 5,000 synthetic yet clinically-informed patient cases for Type 2 diabetes, designed for instruction tuning of language models (e.g., Gemma 3, Unsloth) to recommend ADA guideline-based therapies with drug-specific dosing.
Dataset Overview
Task: Given a patient profile, recommend ADA-aligned diabetes treatment including therapy, drug-specific starting doses, and rationale.
Size: 5,000 examples… See the full description on the dataset page: https://huggingface.co/datasets/mirfan899/ada_diabetes_5000_instruction.mirror-threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-threat-intelligence-dataset.mirror-Polygl0t__gsm8k-pt
Dataset Card for GSM8K-pt
Dataset Summary
This is a Portuguese version of the GSM8K. Translations were attained using Qwen/Qwen2.5-32B-Instruct.
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.
These problems take between 2 and 8 steps to solve.
Solutions primarily… See the full description on the dataset page: https://huggingface.co/datasets/leeaandrob/mirror-Polygl0t__gsm8k-pt.mirror-SWE-Next-SFT-Trajectories
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
SWE-Next SFT Trajectories
SWE-Next SFT Trajectories is the supervised fine-tuning dataset released with SWE-Next: Scalable Real-World Software Engineering Tasks for Agents. It contains 3,693 ShareGPT-style multi-turn training examples collected from expert agent rollouts on 2,308 execution-grounded SWE tasks synthesized from real merged pull requests.
The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-SWE-Next-SFT-Trajectories.mirror-tech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-tech-docs.
