datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gsm8k-synthetic-diverse-8b
gretelai/gsm8k-synthetic-diverse-8b
This dataset is a synthetically generated version inspired by the GSM8K https://huggingface.co/datasets/openai/gsm8k dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-8B as the agent LLM. It contains ~1500 Grade School-level math word problems with step-by-step solutions, focusing on age group, difficulty, and domain diversity.
Key Features:
Synthetically Generated: Math problems created using Gretel… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gsm8k-synthetic-diverse-8b.Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.smoothie-qwen3-8b-kr-self-driving-legal-dataset-v3
📚 Smoothie-Qwen3-8B-KR-Self-Driving-Legal Dataset v3 (DTRO Style)
대한민국 자율주행자동차법 파인튜닝을 위한 750건의 한국어 특화 데이터셋입니다.본 데이터셋은 기존 v1, v2 데이터셋 치명적인 "컨텍스트 소실(Context Forgetting)" 문제를 해결하기 위해 DTRO (Direct-To-Response Output) 스타일로 완전히 재구축되었습니다.
⚠️ 이전 데이터셋(v1, v2)의 문제점과 한계
기존 Alpaca 양식의 데이터셋은 모델 학습 시 다음과 같은 심각한 부작용을 낳았습니다.
1. 시스템 프롬프트 포이즈닝 (System Prompt Poisoning)
// 과거 데이터셋 (문제)
{
"instruction": "당신은 대한민국의 자율주행자동차법 전문가입니다. 정확한 법적 근거를...",
"input": "자율주행자동차의 정의는 무엇입니까? [출처: 관련… See the full description on the dataset page: https://huggingface.co/datasets/bluejude10/smoothie-qwen3-8b-kr-self-driving-legal-dataset-v3.SimpleQA-verified-Hard-Qwen3-8B
SimpleQA Verified Hard for Qwen3-8B
Dataset Summary
This dataset contains the 866 questions that
Qwen/Qwen3-8B failed to answer correctly in up to eight attempts from the
official
google/simpleqa-verified
benchmark.
Each source question was scheduled for eight stochastic generations. As soon
as one generation was graded CORRECT, sampling stopped and the question was
excluded. Questions retained here therefore have pass@8 = 0 under the
model, prompt, sampling, and… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/SimpleQA-verified-Hard-Qwen3-8B.Llama-3.1-8B-Instruct-TriviaQA-HighlyKnownDataset for paper “How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?” (How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?)
Based on TriviaQA dataset
(arxiv.org/abs/2502.14502)
JOSIE-Zero-8B-Reasoning-Traces-N67
JOSIE-Zero-Reasoning-Traces-N86
Reasoning traces generated by the JOSIE-ZERO-8B model.
JOSIE-ZERO-8B is a custom reasoning model trained using the GRPO (Group Relative Policy Optimization) training pipeline implemented in the MLX-LM-LoRA framework. The model was optimized with custom reward functions designed to encourage explicit reasoning, chain-of-thought style problem solving, self-correction, and structured analytical behavior.
This dataset contains high-quality reasoning… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/JOSIE-Zero-8B-Reasoning-Traces-N67.Llama-3.1-8B-Instruct-DBpedia-HighlyKnownDataset for paper “How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?” (https://huggingface.co/papers/2502.14502)
Based on DBpedia dataset
(arxiv.org/abs/2502.14502)
my-distiset-8be10284
Dataset Card for my-distiset-8be10284
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/cansani/my-distiset-8be10284/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/cansani/my-distiset-8be10284.
