datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danish-tool-dialogues-v6
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,667
eval_seen_tools
722
eval_unseen_tools
779
933 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v6.train_v6_filteredThis dataset is part of the paper Prefix Sliding for efficient test-time scaling. It contains training data for reinforcement learning to enable long-horizon reasoning.
Code is available at: https://github.com/Muennighoff/prefix-sliding
hercules-v6.0
Hercules-v6.0
Data Source Description
Hercules-v6.0 is an extensive and diverse dataset that combines various domains to create a powerful tool for training artificial intelligence models. The data sources include conversations, coding examples, scientific explanations, and more. The dataset is sourced from multiple high-quality repositories, each contributing to the robustness of Hercules-v6.0 in different knowledge domains.
Included Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hercules-v6.0.sakthai-combined-v6
SakThai Combined v6
Part of the SakThai model family — fine-tuning and evaluation corpus for instruction-following and tool-use chat.
Dataset Size
Files: data/train.jsonl, data/test.jsonl
Format: JSONL
License: apache-2.0
Last updated: 2026-08-01
Loading
from datasets import load_dataset
ds = load_dataset("Nanthasit/sakthai-combined-v6", split="train")
test_ds = load_dataset("Nanthasit/sakthai-combined-v6", split="test")
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6.100k_Tdk_zurriyet_dna_v6.jsonl
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/100k_Tdk_zurriyet_dna_v6.jsonl.blossom-v6-sft-stage1
BLOSSOM V6 SFT STAGE1
Introduction
BLOSSOM V6 SFT Stage1 is a high-quality, diverse large language model fine-tuning dataset designed for the first-stage SFT training of the Blossom V6 model. Its purpose is to help the model initially align dialogue capabilities through exposure to large-scale synthetic data.
While open-source large language models often release model weights and technical reports, the most advanced open-source models typically withhold their… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/blossom-v6-sft-stage1.lang-obs-v6synthetic_Jailbreak_Defense_Doorpage_v65
📊 Jailbreak Defense Doorpage V65
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
20 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v65.blossom-v6.3-sft-stage1
BLOSSOM V6.3 SFT STAGE1
Introduction
BLOSSOM V6.3 SFT Stage1 is a high-quality, diverse large language model fine-tuning dataset designed for the first-stage SFT training of the Blossom V6.3 model. Its purpose is to help the model initially align dialogue capabilities through exposure to large-scale synthetic data.
While open-source large language models often release model weights and technical reports, the most advanced open-source models typically withhold their… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/blossom-v6.3-sft-stage1.synthetic_Jailbreak_Defense_Doorpage_v61
📊 Jailbreak Defense Doorpage V61
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v61.blossom-v6.1-sft-stage2
BLOSSOM V6.1 SFT STAGE2
Introduction
BLOSSOM V6.1 SFT Stage2 is a high-quality, diverse large language model fine-tuning dataset designed for the second-stage SFT training of the Blossom V6.1 model. Its purpose is to further enhance the model's ability to handle complex instructions on more rare real-world problems.
While open-source large language models often release model weights and technical reports, the most advanced open-source models typically withhold their… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/blossom-v6.1-sft-stage2.chess-coach-v6
Chess Coach v6 (deep-verified training labels)
The current data frontier for the chess-instructor-llm coach: a foundational,
data-first rebuild of the training LABELS (the move plus full provenance), deep-verified
with Stockfish 17 (a two-depth root search with agreement bands), Syzygy tablebases
(endgames of seven pieces or fewer), and Maia-2 human-likelihood. It feeds the
downstream preference (DPO) and engine-distillation retrains.
This dataset is NOT the shipped SFT set. The… See the full description on the dataset page: https://huggingface.co/datasets/khoilamalphaai/chess-coach-v6.synthetic_Jailbreak_Defense_Doorpage_v62
📊 Jailbreak Defense Doorpage V62
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v62.blossom-v6-sft-stage2
BLOSSOM V6 SFT STAGE2
Introduction
BLOSSOM V6 SFT Stage2 is a high-quality, diverse large language model fine-tuning dataset designed for the second-stage SFT training of the Blossom V6 model. Its purpose is to further enhance the model's ability to handle complex instructions on more rare real-world problems.
While open-source large language models often release model weights and technical reports, the most advanced open-source models typically withhold their… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/blossom-v6-sft-stage2.blossom-v6.2-sft-stage2
BLOSSOM V6.2 SFT STAGE2
Introduction
BLOSSOM V6.2 SFT Stage2 is a high-quality, diverse large language model fine-tuning dataset designed for the second-stage SFT training of the Blossom V6.2 model. Its purpose is to further enhance the model's ability to handle complex instructions on more rare real-world problems.
While open-source large language models often release model weights and technical reports, the most advanced open-source models typically withhold their… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/blossom-v6.2-sft-stage2.synthetic-ai-tasks-eval-v6
Synthetic Ai Tasks Eval V6
Specialized synthetic data for APM, AIOps, and GenAI expertise including metrics analysis, transaction tracing, event correlation, incident response, and AI-powered operations.
Dataset Structure
This dataset contains 170 synthetic samples across multiple AI assistant tasks:
data_analysis: 1 samples
code_generation: 1 samples
question_answering: 1 samples
creative_writing: 1 samples
problem_solving: 1 samples
real_user_monitoring: 1 samples… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/synthetic-ai-tasks-eval-v6.synthetic_Jailbreak_Defense_Doorpage_v64
📊 Jailbreak Defense Doorpage V64
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
20 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v64.aksara-sft-clean-v6
AksaraLLM SFT Clean v6
High-quality Indonesian SFT dataset distilled from Gemini 2.5 Flash Lite
via Google Vertex AI, with strict quality gates.
Stats
Train: 16,752 items
Validation: 1,098 items
Total: 17,850 items
Teacher model: gemini-2.5-flash-lite
Task distribution
Task type
Count
factual_qa
4,023
creative
4,003
cultural
3,927
reasoning
3,258
how_to
2,639
Method
Curated ~100 Indonesian topic seeds across history… See the full description on the dataset page: https://huggingface.co/datasets/AksaraLLM/aksara-sft-clean-v6.synthetic_Jailbreak_Defense_Doorpage_v60
📊 Jailbreak Defense Doorpage V60
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v60.synthetic_Jailbreak_Defense_Doorpage_v69
📊 Jailbreak Defense Doorpage V69
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v69.CoderForge-Preview-v6-1000
laion/CoderForge-Preview-v6-1000
Row-subset of togethercomputer/CoderForge-Preview
(trajectories split, filtered_reward1), rendered into Qwen3-compatible
think-first OpenHands-XML wire format.
Why v6?
v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced
garbage at eval time (8888..., 0.0.0.0...) despite clean training losses.
Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first
token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-1000.blueteam-v6
Blue_team_v6
Synthetic fine-tuning dataset generated with Dataset Genie 0.1.0 on 2026-09-21T12:01:11+00:00.
Domain brief
Blue-team is broad. Cover these deliberately, spread across difficulty tiers:
Platforms, not one vendor. Splunk (SPL), Microsoft Sentinel (KQL), Elastic (ES|QL/EQL/Lucene), CrowdStrike, Defender for Endpoint, Sysmon, Zeek/Suricata. Identity: Entra ID, Okta, Active Directory/Kerberos. Cloud: AWS (CloudTrail/GuardDuty), Azure, GCP. OS: Windows… See the full description on the dataset page: https://huggingface.co/datasets/k3nn3dy/blueteam-v6.synthetic_Jailbreak_Defense_Doorpage_v66
📊 Jailbreak Defense Doorpage V66
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
20 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v66.synthetic_Jailbreak_Defense_Doorpage_v63
📊 Jailbreak Defense Doorpage V63
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v63.synthetic_Jailbreak_Defense_Doorpage_v67
📊 Jailbreak Defense Doorpage V67
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
20 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v67.synthetic_Jailbreak_Defense_Doorpage_v68
📊 Jailbreak Defense Doorpage V68
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
20 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v68.kanitakorn-v67-replay-guard-clean
Kanitakorn v67 Replay Guard Clean SFT
Training-ready backup replay package for preserving math, code,
instruction-following, and bilingual short-reasoning behavior. This is not the
primary ThaiExam-gain dataset. Use it only if a Thai overlay improves ThaiExam
but regresses side metrics such as MATH100/AIME/LiveCodeBench smoke tests.
Contents
train.jsonl: 172 SFT message rows.
manifest.json: provenance, counts, hashes, and recommended use.
README.md: this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-v67-replay-guard-clean.CoderForge-Preview-v6-316
laion/CoderForge-Preview-v6-316
Row-subset of togethercomputer/CoderForge-Preview
(trajectories split, filtered_reward1), rendered into Qwen3-compatible
think-first OpenHands-XML wire format.
Why v6?
v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced
garbage at eval time (8888..., 0.0.0.0...) despite clean training losses.
Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first
token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-316.blossom-v6.2-sft-stage1
BLOSSOM V6.2 SFT STAGE1
Introduction
BLOSSOM V6.2 SFT Stage1 is a high-quality, diverse large language model fine-tuning dataset designed for the first-stage SFT training of the Blossom V6.2 model. Its purpose is to help the model initially align dialogue capabilities through exposure to large-scale synthetic data.
While open-source large language models often release model weights and technical reports, the most advanced open-source models typically withhold their… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/blossom-v6.2-sft-stage1.blossom-v6.1-sft-stage1
BLOSSOM V6.1 SFT STAGE1
Introduction
BLOSSOM V6.1 SFT Stage1 is a high-quality, diverse large language model fine-tuning dataset designed for the first-stage SFT training of the Blossom V6.1 model. Its purpose is to help the model initially align dialogue capabilities through exposure to large-scale synthetic data.
While open-source large language models often release model weights and technical reports, the most advanced open-source models typically withhold their… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/blossom-v6.1-sft-stage1.
