datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ReflectionEvoGithub Repo for ReflectEvo: https://github.com/bigai-nlco/ReflectEvo
Arxiv Paper for ReflectEvo: https://arxiv.org/abs/2505.16475
Reflection-Dataset-ShareGPT-v2
Simple "Reflection" method dataset inspired by mattshumer
This is the ShareGPT version. Find prompt and response pair dataset here
This dataset was synthetically generated using Glaive AI. There have been structure improvements and added more rows.
2026-08-04-qwen36-self-reflection-20-80-train
⚠️ SUPERSEDED — do not train from this bundle
Built 2026-08-04 under the old rendering policy, where Qwen3.6 emitted a <think> block on the
final assistant turn only. The repository has since moved to preserve-thinking rendering, in
which every assistant turn carries a think block (real trace, or the empty marker). Both files here
are stale as a result:
mixture.jsonl — rendered under the old policy, so it trains different strings than the current
pipeline produces. It also… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-self-reflection-20-80-train.ReflectionSeq-DS
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
📄 Paper •
🏠 Repo •
🤖 Models •
📚 Datasets
Introduction
ReflectionCoder is a novel approach that effectively leverages reflection sequences constructed by integrating compiler feedback to improve one-off code generation performance. Please refer to our paper and repo for more details!
Models
Model
Checkpoint
Size
HumanEval (+)
MBPP (+)… See the full description on the dataset page: https://huggingface.co/datasets/SenseLLM/ReflectionSeq-DS.unnaturalhermes-reflections-100kReflectionSeq-GPT
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
📄 Paper •
🏠 Repo •
🤖 Models •
📚 Datasets
Introduction
ReflectionCoder is a novel approach that effectively leverages reflection sequences constructed by integrating compiler feedback to improve one-off code generation performance. Please refer to our paper and repo for more details!
Models
Model
Checkpoint
Size
HumanEval (+)
MBPP (+)… See the full description on the dataset page: https://huggingface.co/datasets/SenseLLM/ReflectionSeq-GPT.DARS_synthethsis_reflection
DARS: Dual-Model Verbal Reflection Datasets
This repository contains the training datasets for the DARS (Dual-model Reflective Scoring) framework, a novel approach for automated student answer scoring that uses verbal reflection at inference time.
Overview
The DARS framework employs two specialized models working in tandem:
Reasoner: Generates initial assessments and refines them based on feedback
Critic: Provides targeted verbal reflections and determines when reasoning… See the full description on the dataset page: https://huggingface.co/datasets/jiazhengli/DARS_synthethsis_reflection.Reflection-Dataset-v2
Second version of a simple "Reflection" method dataset inspired by mattshumer
This is the prompt and response version. Find ShareGPT version here
This dataset was synthetically generated using Glaive AI. There have been structure improvements and added more rows.
EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Llama-3.1-8B-Philos-Reflection
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Llama-3.1-8B-Philos-Reflection-details.Overthinking-FCS-ReflectionReflection-Dataset-ShareGPT-v1
V2 is out!!! V2
Simple "Reflection" method dataset inspired by mattshumer
This is the ShareGPT version. Find prompt and response pair dataset here
This dataset was synthetically generated using Glaive AI.
EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-details.Reflection-Dataset-v1
V2 is out!!! V2
Simple "Reflection" method dataset inspired by mattshumer
This is the prompt and response version. Find ShareGPT version here
This dataset was synthetically generated using Glaive AI.
Reflection-AlpacaA mix of multiple other Datasets and also some Selfmade and GPT-4o Prompts.
Reflection-Chinese-Dataset
Reflection-Chinese-Dataset·Reflection中文数据集
Based on mahiatlinux/Reflection-Dataset-v2, translated using RA Translation Tool
unnaturalhermes-reflections-30kapi_graph_reflectiongrok-reflection-cot-ru
march228/grok-reflection-cot-ru
Russian synthetic dataset with question, internal thought text, and final answer.
What is inside
Rows: 4190
Split: train
Main fields:
question
thought_text
answer
thought1..thought5
model
task_type
reflection_count
Format
The dataset is stored as train.jsonl.
thought_text is the joined internal monologue with blank lines between thought blocks.thought1..thought5 preserve the original segmented form from the SQLite source.… See the full description on the dataset page: https://huggingface.co/datasets/march228/grok-reflection-cot-ru.reflection-dataDataset from https://huggingface.co/datasets/glaiveai/reflection-v1
reflection-v1the dataset was generated using synthetic data from llama-3.1-70b following the format of the popular reflection model to improve reasoning on small language models
to ensure diversity, the model can be teached to learn when to use shorter or longer text to reflect on the actual task
column '0' offers shorter responses while column '1' offer longer ones
mattshumer__Reflection-Llama-3.1-70B-details
Dataset Card for Evaluation run of mattshumer/Reflection-Llama-3.1-70B
Dataset automatically created during the evaluation run of model mattshumer/Reflection-Llama-3.1-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mattshumer__Reflection-Llama-3.1-70B-details.Weyaxi_HelpSteer-filtered-Reflection-Gemini-1.5-Flash-ShareGPTSystem prompt taken from here and slightly modified. 4 examples were generated with GPT-4o and then slightly modified.
The examples were sent to Gemini-1.5-Flash followed by the real user turn. Any samples which had responses which were stopped with finish_reason: "SAFETY" were skipped, and responses which did not contain one of each start/stop tag were regenerated. If it failed to generate within 3 tries the sample was skipped.
model = genai.GenerativeModel(
"models/gemini-1.5-flash"… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Weyaxi_HelpSteer-filtered-Reflection-Gemini-1.5-Flash-ShareGPT.Catholic_ReflectionsTutorObservations-Reflectionsalpaca-reflection2glaiveai__Reflection-Llama-3.1-70B-details
Dataset Card for Evaluation run of glaiveai/Reflection-Llama-3.1-70B
Dataset automatically created during the evaluation run of model glaiveai/Reflection-Llama-3.1-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/glaiveai__Reflection-Llama-3.1-70B-details.Vietnamese-mahiatlinux-Reflection-Dataset-ShareGPT-v2-gg-translatedreflection-40k-sharegpt
Datasets
mahiatlinux/Reflection-Dataset-ShareGPT-v2: 9171
isaiahbjork/reflection-scienceqa-sharegpt: 12726
Harshkmr/orca-math-word-reflection: 2435
isaiahbjork/cot-logic-reasoning: 10500
isaiahbjork/chain-of-thought-sharegpt: 7143
isaiahbjork/reflection-spelling-puzzles-sharegpt: 2756
olabs-ai__reflection_model-details
Dataset Card for Evaluation run of olabs-ai/reflection_model
Dataset automatically created during the evaluation run of model olabs-ai/reflection_model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/olabs-ai__reflection_model-details.ReflectionGPT4
