datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k
DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT
This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for
reinforcement-learning experiments. Difficulty is defined by how often the
pretrained google/gemma-4-26B-A4B teacher solved each question across eight
temperature-1 samples under the same rule-based grader used by the RL training
pipeline.
The Hub dataset has three configurations—easy, medium, and hard—and each
configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.arxiv-tex-corpus-mediumarxiv-tex-corpus-medium (15GB)
Medium-scale LaTeX corpus from arXiv (math, CS, physics, statistics)
📄 Paper: https://arxiv.org/abs/2602.17288
📚 Overview
arxiv-tex-corpus-medium (15GB) is a medium-sized version of the arXiv LaTeX corpus, containing structured LaTeX source content extracted from selected arXiv categories.
This dataset is restricted to the following categories:
math
cs
hep-th
hep-ph
quant-ph
stat.ML
stat.TH
This version (~15GB) is intended for:
Research… See the full description on the dataset page: https://huggingface.co/datasets/KiteFishAI/arxiv-tex-corpus-medium.HelixLM-medium-1500.0Mt-2988750pt-20260528
david-thrower/HelixLM-medium-1500.0Mt-2988750pt-20260528
A 1.5 billion token corpus of high quality educational pretraining data
Composition:
A sample from randomly sampled shards of https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus (99%)
Randomly chosen rows from: https://huggingface.co/datasets/open-web-math/open-web-math (1%)
Composition based on Kye's (https://huggingface.co/kye) recommendations at… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/HelixLM-medium-1500.0Mt-2988750pt-20260528.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.medium-web-pentesting
Medium Web Pentesting Articles
Dataset Description
A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body.
This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.Olympiads_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 13284
Filtered size: 13240
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_medium.Numina_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 37133
Filtered size: 37133
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina_medium.medium-articles-en
Dataset Card for "medium-articles-en"
fabiochiu/medium-articles filtered for en only and 100 GPT-4 tiktoken tokens or more.
polaris_filtered_nemotron_medium_math_verifiable
Polaris Filtered Nemotron Medium Sympy Verifiable (v2)
This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using math verify-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset.
Dataset Summary
Total Original Samples: 2,424,392
Final Kept Samples: 357,790 (14.8%)
Target Reasoning Length: 4k-8k… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_math_verifiable.Nemotron-Math-v2-Medium-10k
Nemotron-Math-v2-Medium-10k
A lightweight 10,500-problem subset of
nvidia/Nemotron-Math-v2
for long-horizon Python-TIR reinforcement learning. It contains 1,500 problems
from each metadata.reason_high_with_tool.pass bucket 1 through 7. A
deterministic seed-42 shuffle assigns 500 examples to validation and 10,000
to train.
This Hugging Face release intentionally contains no teacher traces. The
full messages/tools aggregation is retained as a separate local artifact.… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Nemotron-Math-v2-Medium-10k.MediumSetPT
📚 Dataset de Perguntas e Respostas por Tópico
Este repositório contém um dataset com 40.000 amostras estruturadas para tarefas de Processamento de Linguagem Natural (PLN), com foco em perguntas temáticas e respostas desenvolvidas.
📁 Estrutura dos Dados
Cada amostra é representada em formato JSON com os seguintes campos:
id (string): Identificador único da amostra (UUID).
topic (lista de strings): Lista com os tópicos abordados.
prompts (lista de strings):… See the full description on the dataset page: https://huggingface.co/datasets/AxeML/MediumSetPT.polaris_filtered_nemotron_medium_sympy_verifiable
Polaris Filtered Nemotron Medium Sympy Verifiable
This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using sympy-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset.
Dataset Summary
Total Original Samples: 2,500,820
Final Kept Samples: 263,123 (10.5%)
Target Reasoning Length: Optimized for… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_sympy_verifiable.muat-pca-10-medium
Subject Models for Interpretability Training
These examples are intended for training an interpreter to:
Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification.
Signature Extraction
Neuron Profile Methods
pca
Prompt Format
separate
Signature Dataset
configs/dataset_gen/signature_dataset.json
Model Architecture
Number of Layers
8 to 10
Neurons per… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-pca-10-medium.medium_512_1k_tokens_prompts
Medium 512-1K Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ By using this dataset you agree to our Terms of Use.
Overview
703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer.
Statistics
Rows
Token range
File size
Format
703
512 – 1 000
2.9 MB
CSV
Use-cases
Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.unpredictable_rated-mediumThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.mediumish-small-agent-sft-preview-v0.1
mediumish-small-agent-sft-preview-v0.1
Procedurally generated ChatML SFT data for medium/small models, covering agentic tool use,
anti-hallucination habits, grounded refusal, multi-step reasoning, and related epistemic
behaviors.
This build replaces the earlier 5,000-row preview. 6,976 rows, stratified across 20 domains
(reliability/tool-use, bible study, hidden-assumption reasoning, advanced math, code repair
against a documented spec, rulebook/policy simulation, ARC-style grid… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/mediumish-small-agent-sft-preview-v0.1.collabllm-medium-rl-grpo
CollabLLM medium — RL (GRPO) train/validation split
Inputs for GRPO training on the CollabLLM medium document-writing task, as used to produce
the verl-grpo-medium-qwen3-4b-step{50,100,129} checkpoints.
file
rows
size
rl_train.parquet
2072
16M
rl_validation.parquet
—
2.2M
Derived from the CollabLLM medium task (source articles: Kamaljp/medium_articles). Each row is
a single-turn prompt that the training loop expands into a multi-turn conversation with a… See the full description on the dataset page: https://huggingface.co/datasets/yuhan-nlp/collabllm-medium-rl-grpo.tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumfalcon-refinedweb-1M_en_medium
BEE-spoke-data/falcon-refinedweb-1M_en_medium
A sample from falcon-refinedweb:
more than 512 & less than 8192 gpt4 tiktoken tokens
en only (via fasttext-langdetect)
1M samples
GPT-4 tiktoken token count:
token_count
count 1000000.000000
mean 1197.179246
std 964.177338
min 513.000000
25% 653.000000
50% 871.000000
75% 1315.000000
max 8191.000000
Total count: 1197.18 M tokens
mediumish-small-agent-sft-v3
mediumish-small-agent-sft-v3
50000 ChatML SFT rows assembled for a same-day training run.
Field
Value
Release channel
deadline_candidate
production_sft
false (Phase F / council not claimed)
Format
single column chatml (full multi-turn + tool traces)
Tool-bearing rows
6545
Generated
2026-07-19
Source mix
Track
Rows
curriculum
36864
reasoning_policy
5000
truth_seeker
4997
habit_lock
2500
agent_gym_live
639… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/mediumish-small-agent-sft-v3.ping-technical-assistant-mediumNow 3x the size of Ping Technical Assitant Small!
NOTE: A new LoRA will be trained on this data soon!
Ping Technical Assistant Dataset Small
This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-medium.gemma-270m-medium-qa本資料集包含由 ** gemini-2.0-flash ** 生成的對話資料,採用 OpenAI Chat Messages 格式(.jsonl)。資料來源結合:
Reference-free:由 seed 派生的單輪問答。
Reference-based:依據參考文本生成單輪問答。
檔案路徑:data/train.jsonl(選配:data/train.parquet)
結構說明
每列為一筆樣本:{"id": "...", "type": "...", "seed": "...", "context": "...", "messages": [{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]}
type 欄位標示資料來源:reference_free 或 reference_based。
seed 欄位儲存 Reference-free 的原始 seed 指令,或 Reference-based 的參考文本片段。
context 欄位僅在… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/gemma-270m-medium-qa.muat-mean-std-fourier-5-pca-10-medium
Subject Models for Interpretability Training
These examples are intended for training an interpreter to:
Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification.
Signature Extraction
Neuron Profile Methods
mean, std, pca, fourier
Prompt Format
separate
Signature Dataset
configs/dataset_gen/signature_dataset.json
Model Architecture
Number of Layers
6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-fourier-5-pca-10-medium.muat-fourier-5-medium
Subject Models for Interpretability Training
These examples are intended for training an interpreter to:
Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification.
Signature Extraction
Neuron Profile Methods
fourier
Prompt Format
separate
Signature Dataset
configs/dataset_gen/signature_dataset.json
Model Architecture
Number of Layers
6 to 8
Neurons… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-fourier-5-medium.VeriReason-reasoning-reproduced-1513_luna-medium
VeriReason reasoning reproduced (Luna medium)
This dataset is a deterministic sample of the locally updated reproduced VeriReason files.
Source files: train (1).jsonl, validation (1).jsonl
Sampling: independent shuffle then prefix slice, fixed seed 42
Splits: train 1513, validation 189
Original target dataset: Jongbin-kr/VeriReason-RTL-Coder_7b_reasoning_tb (train 1513 / validation 189)
Schema: id, instruction, output, tb, tb_result
Quality checks: valid JSON, unique IDs… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/VeriReason-reasoning-reproduced-1513_luna-medium.muat-mean-std-medium
Subject Models for Interpretability Training
These examples are intended for training an interpreter to:
Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification.
Signature Extraction
Neuron Profile Methods
mean, std
Prompt Format
separate
Signature Dataset
configs/dataset_gen/signature_dataset.json
Model Architecture
Number of Layers
6 to 8… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-medium.ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513
GPT-5.5 Medium Reannotation of Qwen3.5-Positive OCR2 Coding Steps
This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and contains GPT-5.5 medium-reasoning reannotations for the 1,536 Qwen3.5-positive error steps.
Summary
Source dataset: JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513
Source rows: 500 K2-Think Codeforces traces
Source manifest-selected Qwen3.5 labels: 10,000 steps
GPT-5.5 reannotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513.nemotron-math-v2-medium-mini
Nemotron Math V2 Medium Mini
A compact textual reasoning dataset derived from the medium split of nvidia/Nemotron-Math-v2 at immutable revision 8e793210e175b6406c752a870f585f62de98c0d3.
Selection boundary
This extraction intentionally excludes tool-use semantics:
Accept exactly two messages with roles user then assistant.
Reject records declaring tools, containing assistant tool calls, containing tool-result messages, or containing any additional turns.
Require… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/nemotron-math-v2-medium-mini.mn_business_benchmark_dataset_medium
mn_business_benchmark_dataset_10000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics, стратегийн 10000 мөртэй синтетик benchmark dataset.
Schema
id: 1-ээс 10000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: бизнесийн сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG card дүрслэл
Generated deterministically by… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_medium.
