datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-SFT-Agentic-v2-prompt-only
Nemotron-SFT-Agentic-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.dementor-sft-dataNemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.Fast-Math-R1-SFTThis repository contains the First stage SFT dataset as presented in the paper A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning.
This dataset is used for the intensive Supervised Fine-Tuning (SFT) phase, crucial for pushing the model's mathematical accuracy.
Project GitHub Repository: https://github.com/RabotniKuma/Kaggle-AIMO-Progress-Prize-2-9th-Place-Solution
Dataset Construction
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/RabotniKuma/Fast-Math-R1-SFT.Nemotron-SFT-Competitive-Programming-v2-prompt-only
Nemotron-SFT-Competitive-Programming-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Competitive-Programming-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Competitive-Programming-v2-prompt-only.llm-system-ops-production-telemetry-sft-data
🤖📈 LLM System Ops Telemetry (Synthetic)
A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments.
It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level,
with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension.
Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.HypoAgent-SFT
HypoAgent-SFT: study description → prespecified statistical analysis plan
54,163 supervised fine-tuning pairs that teach a model to read a study registration and write the analysis plan that belongs to it. The input is a study description — objective, design, population, exposure, comparator, outcome, timing, whatever the registry record actually contains. The target is a prespecified hypothesis-testing and statistical-analysis plan built for that design.
This is the curated… See the full description on the dataset page: https://huggingface.co/datasets/HypoAgent/HypoAgent-SFT.LYRICAL_sft_v7
SilverAgePoets.com & RuVERSES.com Russian-English Bilingual Poetry Library
A dataset of Eastern European and Soviet poetry and song lyrocs from https://RuVerses.com/, with Russian-language sources and English translations.
This variant of the dataset combines a revised and somewhat pre-filtered version of the RuVerses collection dataset + the entirety of the SFT version of our LYRICAL dataset.
Featuring a present (c. late 2025) state of the RuVerses archive, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/LYRICAL_sft_v7.Nemotron-SFT-CUDA-v1-prompt-only
Nemotron-SFT-CUDA-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-CUDA-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-CUDA-v1-prompt-only.Nemotron-SFT-Safety-v1-prompt-only
Nemotron-SFT-Safety-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Safety-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Safety-v1-prompt-only.nova_sft
Nova SFT
This is a merged dataset from:
OpenHermes 2.5
AllenAI Tulu-3-SFT
Version 1, made on 2nd March 2025
sft-chineseClassical-Mechanics-Equations-Dataset_SFT-or-LoRA
Classical Mechanics Equations Dataset (SFT / LoRA Ready)
A structured dataset of 64 classical mechanics equations from Newtonian,
Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning
rows across three task types: equation explanation, Q&A, and derivation.
Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and
equation understanding tasks.
Overview
Property
Value
Domain
Classical Mechanics (Physics)
Total rows
448
Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.Logic-OA-SFTtrainLYRICAL_Mix_SilverAgePoets_Songs_RuVerses_SFT
SilverAgePoets.com & RuVERSES.com Russian-English Bilingual Poetry Library
A dataset of Eastern European and Soviet poetry and song lyrocs from https://RuVerses.com/, with Russian-language sources and English translations.
This variant of the dataset combines a revised and somewhat pre-filtered version of the RuVerses collection dataset + the entirety of the SFT version of our LYRICAL dataset.
Featuring a present (c. late 2025) state of the RuVerses archive, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/LYRICAL_Mix_SilverAgePoets_Songs_RuVerses_SFT.Lyrical_Ru2En_Poems_Songs_MeterMatched_csv_SFT
LYRICAL Russian2English SFT Version:
Meaning+Meter-Matched Russian & Soviet Poems + Songs
Manually Adapted by a Poet-Translator
1776 rows/items and 2 columns
CSV version
Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric, syllabic, melodic, and other lyrical and literary features, whilst retaining adequate semantic/significational fidelity.
Moreover… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_Ru2En_Poems_Songs_MeterMatched_csv_SFT.Nemotron-SFT-Safety-v2-prompt-only
Nemotron-SFT-Safety-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Safety-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Safety-v2-prompt-only.Nemotron-SFT-Multilingual-v2-prompt-only
Nemotron-SFT-Multilingual-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v2-prompt-only.hhh_sftCore-Model-SFT-Adv-ReasoningNemotron-SFT-ARC-AGI-v1-prompt-only
Nemotron-SFT-ARC-AGI-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-ARC-AGI-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-ARC-AGI-v1-prompt-only.sft-dataSource: https://huggingface.co/datasets/alespalla/chatbot_instruction_prompts/viewer/alespalla--chatbot_instruction_prompts/train
license: apache-2.0
alpaca_hhh_sft_headlines_2020_2022
Alpaca-HHH-SFT-headlines-2020-2022
This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests.
This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca_hhh_sft_headlines_2020_2022.VStarBench_sft_w_toolcallNemotron-SFT-Multilingual-v1-prompt-only
Nemotron-SFT-Multilingual-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v1-prompt-only.Jazz-Blues-Music-Dataset_SFT-or-LoRA
Jazz & Blues Music Dataset (SFT / LoRA Ready)
A structured dataset covering 82 iconic Jazz and Blues songs, 21 artist
profiles, and 41 historical events, expanded into 1,219
instruction-tuning rows across 7 task types.
Designed for fine-tuning LLMs on music knowledge, cultural history, artist
biography, and domain-specific Q&A tasks.
Overview
Property
Value
Domain
Jazz & Blues Music
Total rows
1,219
Train split
1,036 (85%)
Validation split
91 (~7.5%)… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Jazz-Blues-Music-Dataset_SFT-or-LoRA.efemerides-cumana-sft-qa-small
Efemérides de Cumaná/Venezuela – Dataset de Preguntas y Respuestas para SFT
Descripción general
Este dataset contiene un conjunto curado de pares pregunta–respuesta en español, diseñado específicamente para ajuste fino supervisado (Supervised Fine-Tuning, SFT) de modelos de lenguaje de gran tamaño (LLMs). El contenido se centra en la historia de Cumaná, ciudad de Venezuela durante el período de la Guerra de Independencia, con énfasis en la interpretación histórica, el… See the full description on the dataset page: https://huggingface.co/datasets/rjcg/efemerides-cumana-sft-qa-small.medical-o1-reasoning-SFT-Arabicstanford-NIL-disclosure-sft
NIL Policy
Data is taken from the Stanford website.
The maximum number of tokens (prompt + completion) in a row of data/train.csv is 100
The maximum number of tokens (prompt + completion) in a row of data/test.csv is 89
For educational and non-commercial use only.
