datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Turkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.AgentDoG1.0-Training-Data
AgentDoG1.0 Training Data
[💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection]
AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis.
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.magpie-sft-v1.0
magpie-sft-v1.0
This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This is a dataset of instruction and response pairs created using the Magpie method.
cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.Ko-Agent-Trajectories-1.0
Ko-Agent-Trajectories-1.0
Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository
under pipeline/, together with the API catalogue, the scenario templates and the complete
prompt set. The card reports the completed human review study and the v1.1 artefacts
(behaviour DPO config, per-item validation scores, manifest, filter asset).
Korean edition: README.ko.md.
TL;DR
A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.cyberusecase-v1.0
Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge
A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level
cybersecurity reasoning across vulnerability management, SOC alert triage, detection
engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps.
It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a
hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.qa-expert-multi-hop-qa-V1.0
Dataset Card for QA-Expert-multi-hop-qa-V1.0
This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering.
In total, this dataset contains 25.5k for training and 3.19k for evaluation.
You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0
The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.JL-ActionBoundary-1K-v1.0.0
JL-ActionBoundary-1K v1.0.0
Counterfactual Ask–Inspect–Act–Defer supervision for coding agents
JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code:
ACT: the task is sufficiently specified for bounded repository work;
INSPECT: missing information can be recovered from the repository;
ASK: a material product decision belongs to the user;
DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.OpenTopics-1.0-20K
OpenTopics-1.0-20K
What is this dataset?
OpenTopics-1.0-20K is a collection of 20,003 topic names spanning a wide variety of subjects, including physics, medicine, history, law, engineering, and the arts.
AtomixLabs built this dataset to help developers, researchers, and AI builders who need a large, organized list of topics. It works great for creating synthetic prompts, testing search systems, and training models to classify text.
What is inside… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OpenTopics-1.0-20K.ATANTV1.0-corpus
ATANT Narrative Test Corpus
Automated Test for Acceptance of Narrative Truth, v1.0
The first open evaluation corpus for measuring continuity in AI systems: the ability to persist, update, disambiguate, and reconstruct meaningful context across time.
Paper: ATANT: An Evaluation Framework for AI Continuity (arXiv:2604.06710)
Standard repository: github.com/Kenotic-Labs/ATANT
Author: Samuel Sameer Tanguturi
Affiliation: Kenotic Labs
Published: April 2026
Why this corpus… See the full description on the dataset page: https://huggingface.co/datasets/Kenotic-Labs/ATANTV1.0-corpus.NemoSlides-DPO-mix-v1.0
Slide-DPO
Direct Preference Optimization dataset for training LLMs to generate slide
presentations in Slidev markdown format, derived from the
Slides-Align human
preference rankings over the
SlidesGen-Bench benchmark.
Each row is a preference pair: a brief plus an available image pool as the
prompt, and two Slidev-markdown responses (with <think> reasoning traces)
that were generated by differently-ranked AI slide-generation products for
the same brief.
Row schema… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/NemoSlides-DPO-mix-v1.0.OmniSurgical-1.0
OmniSurgical 1.0
OmniSurgical is a dataset which you can train your very own massively multilingual machine translation models by fine-tuning existing LLMs!
Formats
We give the dataset in 2 formats: JSONL and JSONZ (zipped JSONL)
And the names speak for themselves: OmniSurgical_120_Clean.jsonz is the processed file and train.jsonz is the shuffled version of the same file, used to fine-tune existing LLMs (I fine-tuned Qwen 3 0.6B for this!)
Data Used
I used only… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/OmniSurgical-1.0.jfv-french-style-conditioning-dataset-v1.0
JFV French Paired Style-Conditioning Dataset
At a Glance
Item
Value
Language
French
Source
Single-author blog corpus, 2005–2025
Public release
v1.0
Aligned units in public release
1,484
Texts in public aligned release
7,420
Original experiment
1,492 aligned units / 7,460 texts
Generated conditions
Ministral baseline; profile; profile + five-shot examples
Primary use
Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.LimeStory-1.0NEW IN THIS DATASETLimeStory Dataset Version 1.0 is now available in 🤗 Spaces and users can add stories about anything! (Powered by Pollinations.ai)
NOTICELimeStory is not for training NSFW models, and remember to use dataset: kulia-moon/LimeStory-1.0 for you're using this dataset as target training!
Kulia's datasets
The story,
your impossible
Generated by 🤗 Spaces
Protected by 🤗 Scanner
.hf-sanitized.hf-sanitized-yvTjnu7IKxZ1msJf6A22y .cursive { font-family: "Lobster"… See the full description on the dataset page: https://huggingface.co/datasets/kulia-moon/LimeStory-1.0.CodeThink-v1-1.04kSmall synthetic data set for fine tuning models in preparation for further tuning via GRPO.
Main focus is python with some javascript and html/css.
properly-v1.04
properly-E4-v1.04
Summary
Dataset ID: 109
Type: mixture
Rows: 30,000
Dataset Sources
#109 properly-E4-v1.04 [mixture | 2 sources | 30,000 rows]
#100 HF startc/synthetic-spelling | default | csd [hf | HF startc/synthetic-spelling/default:csd | 100,000 rows | mix 50.2% | target 15,060]
#93 HF grammarly/coedit | default | train [hf | HF grammarly/coedit/default:train | 69,071 rows | mix 50.2% | target 15,060 | kept 15,060]
Notes
Exported from the… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/properly-v1.04.paper2thesis1.0
Paper2Thesis
Anonymized release for NeurIPS 2026 Datasets and Benchmarks Track review.
The non-anonymous version, including author information and a permanent DOI, will be released upon acceptance.
Overview
Paper2Thesis is a benchmark for extreme-length multi-document synthesis. Each instance maps a set of input arXiv research papers to a target arXiv PhD thesis. The task requires generating a thesis-scale document that integrates multiple papers into a coherent… See the full description on the dataset page: https://huggingface.co/datasets/anon-nips2026/paper2thesis1.0.tensorbench-1.0
TensorBench
Feature-addition benchmark for LLMs and coding agents, evaluated against the
Scorch codebase. Each task asks a model
to add a feature (or otherwise extend functionality) to Scorch. Success is
defined as the full pytest suite (original + any new tests the model adds)
passing after the patch is applied inside a Docker container.
This is the dataset artifact for the TensorBench paper (NeurIPS 2026
Evaluations & Datasets track, double-blind submission).
At a… See the full description on the dataset page: https://huggingface.co/datasets/tensorbench/tensorbench-1.0.cwec-v4.14-weaknesses-1.0
Introduction
This dataset is based on the complete XML file of CWE List Version 4.14 and is intended to provide researchers and security experts with structured data on Common Weakness Enumeration (CWE) for software and hardware. The dataset contains 963 entries in Alpaca format, each providing detailed information about a specific weakness.
Dataset Structure
Each entry in the dataset includes the following fields:
ID: The unique identifier for the weakness (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bayuncao/cwec-v4.14-weaknesses-1.0.ChunMengDie-1.0-User-Data
ChunMengDie-1.0-User-Data
本数据集是 XingChina 为 ChunMengDie 系列模型手写的的原创中文对话数据集(或许以后可以管XingChina叫春梦蝶?)。
数据规模
当前版本:1000 条对话配对(JSONL 格式)
迭代计划:后续版本将持续在此仓库扩充
数据格式
每行为一个 JSON 对象:
{
"messages": [
{"role": "user", "content": "你好"},
{"role": "assistant", "content": "(嘴角微翘)哼~终于来啦?笨蛋。"}
]
}
数据风格
中文对话,猫娘/傲娇语气
用于为模型注入特定人格
用途
✅ SFT 训练
✅ 人格注入
✅ 后续版本复用和扩展
许可证
采用 CC BY 4.0 许可证。
Copyright (c) 2026… See the full description on the dataset page: https://huggingface.co/datasets/XingChina/ChunMengDie-1.0-User-Data.i2b2-query-data-1.0
i2b2 query data 1.0
This is a dataset of i2b2 query builder examples that are taken from a test environment of i2b2 and then pre-processed with AI descriptions.
arabic-conversation-final-v1.0
Arabic Conversation — Final v1.0 (post-processed)
Post-processed release of the original arabic-conversation-final dataset.
Only assistant messages were modified; user messages, personas, metadata,
factuality verdicts, and IDs are untouched. The source file
(all_shuffled.jsonl) is preserved upstream — this repo holds the cleaned
variant as a separate file.
What changed
Six rule-based cleanup passes are applied to every assistant message.
All rules are deterministic regex… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/arabic-conversation-final-v1.0.OpenTalk-v1.0
Dataset Summary
OpenTalk-v1.0 is a synthetic, English-language instruction dataset containing 5,582 conversational data points. It was generated using topics from the MultivexAI/STEMScoredTopics-v1.0 dataset to teach language models persona adoption and friendly interaction.
Data Fields
systemPrompt: A string that sets a specific persona for the AI.
question: A string representing a user's query on a topic.
answer: A string representing the AI's persona-driven… See the full description on the dataset page: https://huggingface.co/datasets/MultivexAI/OpenTalk-v1.0.
