datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolma3_mix-6T-1025-7B
⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️
For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T.
Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B.
For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.fineweb-edu-2013-qwen2-7b
FineWeb-Edu 2013 with Qwen2-7B token counts
Every 2013 FineWeb-Edu document, prepared for continued pretraining, with token
counts computed by a pinned Qwen2-7B tokenizer.
The pipeline is year-agnostic: the year, source revision, tokenizer contract,
and selection rule all come from a config file. 2013 uses
processing_config.json. The 2017 companion dataset, which is large enough to
require shuffling and a token budget rather than retaining everything, is at… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2013-qwen2-7b.fineweb-edu-2017-qwen2-7b
FineWeb-Edu 2017 (~100B-token subset) with Qwen2-7B token counts
A ~100B-token subset of FineWeb-Edu 2017, prepared for continued pretraining,
with token counts computed by a pinned Qwen2-7B tokenizer.
This dataset is a selected subset, not the complete 2017 crawl year. 2017
contains about 168B Qwen2-7B tokens, above the 100B target, so it was shuffled
and subsetted: data/train/ holds 101,840,059 documents and 100,000,020,347
tokens, which is 59.29% of the 171,755,787 documents… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2017-qwen2-7b.tis-subset-datasets-Llama-2-7b-hf
Targeted Instruction Selection Subsets (Llama-2-7b-hf)
This repository contains pre-computed instruction training subsets selected from a large candidate pool for targeted instruction fine-tuning, as presented in the paper A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't).
Paper: https://huggingface.co/papers/2602.14696
GitHub Repository: https://github.com/dcml-lab/targeted-instruction-selection
Description
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-DCML/tis-subset-datasets-Llama-2-7b-hf.Qwen2.5-7B-Instruct-Self-Calibration
Efficient Test-Time Scaling via Self-Calibration
This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.multireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.RTL-Coder_7b_reasoning_tb_combined
Verireason-RTL-Coder_7b_reasoning_tb_combined
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
This is the combined version of VeriReason-RTL-Coder_7b_reasoning_tb and VeriReason-RTL-Coder_7b_reasoning_tb_simple.
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_combined
Project… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning_tb_combined.VeriReason-RTL-Coder_7b_reasoning_tb_simple
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.VeriReason-RTL-Coder_7b_reasoning_tb
Verireason-RTL-Coder_7b_reasoning_tb
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb.qwen2.5-7b-instruct-nla-L20-finefineweb-100k
Qwen2.5-7B-Instruct NLA training data — residual stream, block 20
Training data for a Natural Language Autoencoder on Qwen/Qwen2.5-7B-Instruct:
residual-stream activations paired with natural-language explanations of the text
they were taken from.
Unlike the dataset this is derived from, the activation_vector column is
included — every EasyNLA/nanoNLA trainer requires it.
Trained models: https://huggingface.co/Yooniel/qwen2.5-7b-instruct-nla-L20
(AV val ppl 4.07, AR held-out FVE… See the full description on the dataset page: https://huggingface.co/datasets/Yooniel/qwen2.5-7b-instruct-nla-L20-finefineweb-100k.Convergent-7B-data
Convergent-7B Training Data
Training data for the bigcompute.science research companion model.
Early Preview — This dataset is a work in progress. It is expressly designed to train a research assistant for the bigcompute.science MCP server as part of the Convergent conjecture-driven GPU research project. The dataset will be updated frequently as new experiments, findings, and tool definitions are added. Expect changes to schema, tool names, and content until we reach a GA… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/Convergent-7B-data.RTL-Coder_7b_reasoning
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_7b_reasoning.polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps
Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507
Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's
actual continuation for each, and the token accounting behind it.
The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is
decoded to text, the teacher is shown it under its own chat template, and the teacher's
reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.InverseCoder-CL-7B-Evol-Instruct-90K
InverseCoder: Unleashing the Power of Instruction-Tuned Code LLMs with Inverse-Instruct
InverseCoder is a series of code LLMs instruction-tuned by generating data from itself through Inverse-Instruct.
Models and Datasets
Base Model
InverseCoder
Dataset
6.7B
deepseek-ai/deepseek-coder-6.7b-base
wyt2000/InverseCoder-DS-6.7B
wyt2000/InverseCoder-DS-6.7B-Evol-Instruct-90K
7B
codellama/CodeLlama-7b-Python-hf
wyt2000/InverseCoder-CL-7B… See the full description on the dataset page: https://huggingface.co/datasets/wyt2000/InverseCoder-CL-7B-Evol-Instruct-90K.QFT-OLMo-3-7B-Think
QFT-OLMo-3-7B-Think
Two query-aligned mathematical reasoning SFT datasets generated from the same
13,094 DAPO-Math-17K prompts. The datasets differ only in the assistant
trajectory distribution:
think: allenai/Olmo-3-7B-Think
dexpert: Think + 1.0 × (Think − Think-SFT), using
allenai/Olmo-3-7B-Think-SFT as the reference model
Each source query was sampled exactly once from each distribution. A query is
included only when both trajectories are verifier-correct and both pass the… See the full description on the dataset page: https://huggingface.co/datasets/qingyangzhang/QFT-OLMo-3-7B-Think.olmo-3-7b-think_ifeval
allenai/OLMo-3-7B-Think — ifeval
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Think
Dataset: ifeval (541 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-think_ifeval.das-dpo-data-searchr1-7b
DAS dpo data-searchr1-7b
This dataset provides DPO preference data for post-training SearchR1 7B search agents with DAS.
The file is provided in LLaMA-Factory compatible DPO format with prompt, chosen, rejected, and optional system fields.
VeriReason-RTL-Coder_7b_no_reasoning_tb
VeriReason-RTL-Coder_7b_no_reasoning_tb
This dataset is an answer-only derivative of
Jongbin-kr/VeriReason-RTL-Coder_7b_reasoning_tb, created for
an SFT ablation on the effect of explicit reasoning paths.
Transformation
The output field is changed as follows:
<think>reasoning path</think><answer>solution</answer>
becomes:
<answer>solution</answer>
The complete <answer>...</answer> block, including the Verilog code, is retained.
For instructions containing the… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/VeriReason-RTL-Coder_7b_no_reasoning_tb.Medical-QA-Mistral7B-FinetuningSFT_DATA-openthoughts-1k_rows-main-Qwen2.5-7B-Instruct-SkillFactoryYou can train using these datasets with LLaMA-Factory if you add this to your data/datasets.json files.
"example_dataset": {
"hf_hub_url": "SkillFactory/SFT_DATA-openthoughts-1k_rows-main-Qwen2.5-7B-Instruct-SkillFactory",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"},
"tags": {
"user_tag": "user",
"assistant_tag": "assistant",
"role_tag": "role",
"content_tag": "content"
},
"subset": "sft_train"
}
Llama-2-7b-chat-finetune
plagas y enfermedades en el cultivo del tomate Dataset 1000
Dataset de 1000 instrucciones sobre la plagas y enfermedades en el cultivo del tomate.
Uso
from datasets import load_dataset
dataset = load_dataset("anyerg21/plagas-enfermedades-tomate-1000")
Estructura
instruction: Pregunta sobre el cultivo del tomate
input: Campo vacio
output: Respuesta
category: Categoria tematica
question_type: Tipo de pregunta
difficulty: Nivel de dificultad
Ejemplo… See the full description on the dataset page: https://huggingface.co/datasets/anyerg21/Llama-2-7b-chat-finetune.EVAL-OT-Qwen2.5-7B-Instruct-QwQ-1k_rows-RL
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt we will feed into the model to solve the question
model_responses
An array of strings that the model generated to answer the prompt (usually size of 4 or 34 depending on the evaluation task)
model_responses__eval_is_correct
An array aligned with model_responses containing booleans: True when… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/EVAL-OT-Qwen2.5-7B-Instruct-QwQ-1k_rows-RL.olmo-3-7b-instruct_alpaca-text-generation-384
allenai/OLMo-3-7B-Instruct — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Instruct
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-instruct_alpaca-text-generation-384.creative-qwen2.5-7b-stories
Ethahtz/creative-qwen2.5-7b-stories
LLM creative generations from the creative_tasks pipeline (generate_infinite_chats.py) or any compatible generations.jsonl.
Reproducibility and full per-run parameters are in generation_config.json in this dataset repository (one entry per config / run).
Configs and loading
Qwen-Qwen2.5-7B-Instruct_dsshort_story_prompts_bkvllm_seed42_top
Model: Qwen/Qwen2.5-7B-Instruct
Prompt source (dataset):… See the full description on the dataset page: https://huggingface.co/datasets/Ethahtz/creative-qwen2.5-7b-stories.code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.qwen2.5-7b-ultrafeedback-armorm-binarizedThis repository contains the data for the paper Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model.
Github: https://github.com/DtYXs/Pre-DPO
ThinkSafe-R1-Distill-7B
ThinkSafe Dataset
This dataset is associated with the paper THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.
Paper: https://arxiv.org/abs/2601.23143GitHub: https://github.com/seanie12/ThinkSafe.git
Citation
If you use this dataset, please cite:
@article{lee2025thinksafe,
title={THINKSAFE: Self-Generated Safety Alignment for Reasoning Models},
author={Lee, Seanie and others},
journal={arXiv preprint arXiv:2601.23143},
year={2025}
}
EVAL-cd3args-Qwen2.5-7B-Instruct-RL
Column Details
Column
Description
question
The question we want the model to answer
answer
The string answer
task
The name of the task the row belongs to
prompt
The prompt we will feed into the model to solve the question
model_responses
An array of strings that the model generated to answer the prompt (usually size of 4 or 34 depending on the evaluation task)
model_responses__eval_is_correct
An array aligned with model_responses containing booleans: True when… See the full description on the dataset page: https://huggingface.co/datasets/SkillFactory/EVAL-cd3args-Qwen2.5-7B-Instruct-RL.olmo-3-7b-think_bookmia-label0-5pct-raw
allenai/OLMo-3-7B-Think — bookmia-label0-5pct-raw
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Think
Dataset: bookmia-label0-5pct-raw (247 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-think_bookmia-label0-5pct-raw.zephyr-7b-beta-invoices
Zephyr-7B-Beta Customer Support Chatbot
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Introduction
Welcome to the zephyr-7b-beta-invoices repository! This project leverages the Zephyr-7B-Beta model trained on the "Bitext-Customer-Support-LLM-Chatbot-Training-Dataset" to create a state-of-the-art customer support chatbot. Our goal is to provide an efficient and accurate chatbot for handling invoice-related… See the full description on the dataset page: https://huggingface.co/datasets/erfanvaredi/zephyr-7b-beta-invoices.
