CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes426 downloads11mo agoHugging Face02AI45Research /AgentDoG1.0-Training-Data AgentDoG1.0 Training Data [💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection] AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis. Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.texttext-generation1K<n<10K0 likes422 downloads4mo agoHugging Face03llm-jp /magpie-sft-v1.0 magpie-sft-v1.0 This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This is a dataset of instruction and response pairs created using the Magpie method. cyberagent/calm3-22b-chat was used for generating the instructions, and Qwen/Qwen2.5-32B-Instruct was used for generating the responses. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order.… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/magpie-sft-v1.0.texttext-generation100K<n<1M19 likes350 downloads2y agoHugging Face04taejoon89 /Ko-Agent-Trajectories-1.0 Ko-Agent-Trajectories-1.0 Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository under pipeline/, together with the API catalogue, the scenario templates and the complete prompt set. The card reports the completed human review study and the v1.1 artefacts (behaviour DPO config, per-item validation scores, manifest, filter asset). Korean edition: README.ko.md. TL;DR A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.tabulartext-generation100K<n<1M0 likes234 downloads2d agoHugging Face05ronaldocloud /cyberusecase-v1.0 Cybersecurity SOC Fine-Tuning Dataset — 17.5k Real CVEs (2018–2026) + SOC Knowledge A large supervised fine-tuning (SFT) dataset for teaching an LLM expert-level cybersecurity reasoning across vulnerability management, SOC alert triage, detection engineering, threat intelligence & hunting, incident response, and cloud/DevSecOps. It combines 17,590 real CVEs (2018–2026) pulled from the NIST NVD data feeds with a hand-curated set of 65 landmark CVEs (rich, multi-angle coverage)… See the full description on the dataset page: https://huggingface.co/datasets/ronaldocloud/cyberusecase-v1.0.texttext-generation10K<n<100K0 likes146 downloads3mo agoHugging Face06khaimaitien /qa-expert-multi-hop-qa-V1.0 Dataset Card for QA-Expert-multi-hop-qa-V1.0 This dataset aims to provide multi-domain training data for the task: Question Answering, with a focus on Multi-hop Question Answering. In total, this dataset contains 25.5k for training and 3.19k for evaluation. You can take a look at the model we trained on this data: https://huggingface.co/khaimaitien/qa-expert-7B-V1.0 The dataset is mostly generated using the OpenAPI model (gpt-3.5-turbo-instruct). Please read more information about… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/qa-expert-multi-hop-qa-V1.0.textquestion-answering10K<n<100K8 likes120 downloads3y agoHugging Face07jumplander /JL-ActionBoundary-1K-v1.0.0 JL-ActionBoundary-1K v1.0.0 Counterfactual Ask–Inspect–Act–Defer supervision for coding agents JL-ActionBoundary-1K teaches a coding agent to choose the correct next policy before changing code: ACT: the task is sufficiently specified for bounded repository work; INSPECT: missing information can be recovered from the repository; ASK: a material product decision belongs to the user; DEFER: live execution authority or rollback ownership is missing.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-ActionBoundary-1K-v1.0.0.tabulartext-classification1K<n<10K5 likes120 downloads2mo agoHugging Face08AtomixLabs /OpenTopics-1.0-20K OpenTopics-1.0-20K What is this dataset? OpenTopics-1.0-20K is a collection of 20,003 topic names spanning a wide variety of subjects, including physics, medicine, history, law, engineering, and the arts. AtomixLabs built this dataset to help developers, researchers, and AI builders who need a large, organized list of topics. It works great for creating synthetic prompts, testing search systems, and training models to classify text. What is inside… See the full description on the dataset page: https://huggingface.co/datasets/AtomixLabs/OpenTopics-1.0-20K.tabulartext-classification10K<n<100K3 likes79 downloads2mo agoHugging Face09Kenotic-Labs /ATANTV1.0-corpus ATANT Narrative Test Corpus Automated Test for Acceptance of Narrative Truth, v1.0 The first open evaluation corpus for measuring continuity in AI systems: the ability to persist, update, disambiguate, and reconstruct meaningful context across time. Paper: ATANT: An Evaluation Framework for AI Continuity (arXiv:2604.06710) Standard repository: github.com/Kenotic-Labs/ATANT Author: Samuel Sameer Tanguturi Affiliation: Kenotic Labs Published: April 2026 Why this corpus… See the full description on the dataset page: https://huggingface.co/datasets/Kenotic-Labs/ATANTV1.0-corpus.tabularquestion-answeringn<1K0 likes70 downloads6mo agoHugging Face10trillionlabs /NemoSlides-DPO-mix-v1.0 Slide-DPO Direct Preference Optimization dataset for training LLMs to generate slide presentations in Slidev markdown format, derived from the Slides-Align human preference rankings over the SlidesGen-Bench benchmark. Each row is a preference pair: a brief plus an available image pool as the prompt, and two Slidev-markdown responses (with <think> reasoning traces) that were generated by differently-ranked AI slide-generation products for the same brief. Row schema… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/NemoSlides-DPO-mix-v1.0.texttext-generation1K<n<10K6 likes67 downloads5mo agoHugging Face11MihaiPopa-1 /OmniSurgical-1.0 OmniSurgical 1.0 OmniSurgical is a dataset which you can train your very own massively multilingual machine translation models by fine-tuning existing LLMs! Formats We give the dataset in 2 formats: JSONL and JSONZ (zipped JSONL) And the names speak for themselves: OmniSurgical_120_Clean.jsonz is the processed file and train.jsonz is the shuffled version of the same file, used to fine-tune existing LLMs (I fine-tuned Qwen 3 0.6B for this!) Data Used I used only… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/OmniSurgical-1.0.texttext-generation10K<n<100K0 likes64 downloads6mo agoHugging Face12PeggyVallin /jfv-french-style-conditioning-dataset-v1.0 JFV French Paired Style-Conditioning Dataset At a Glance Item Value Language French Source Single-author blog corpus, 2005–2025 Public release v1.0 Aligned units in public release 1,484 Texts in public aligned release 7,420 Original experiment 1,492 aligned units / 7,460 texts Generated conditions Ministral baseline; profile; profile + five-shot examples Primary use Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.tabulartext-generation1K<n<10K0 likes54 downloads5d agoHugging Face13kulia-moon /LimeStory-1.0NEW IN THIS DATASETLimeStory Dataset Version 1.0 is now available in 🤗 Spaces and users can add stories about anything! (Powered by Pollinations.ai) NOTICELimeStory is not for training NSFW models, and remember to use dataset: kulia-moon/LimeStory-1.0 for you're using this dataset as target training! Kulia's datasets The story, your impossible Generated by 🤗 Spaces Protected by 🤗 Scanner .hf-sanitized.hf-sanitized-yvTjnu7IKxZ1msJf6A22y .cursive { font-family: "Lobster"… See the full description on the dataset page: https://huggingface.co/datasets/kulia-moon/LimeStory-1.0.texttext-generation1K<n<10K0 likes48 downloads11mo agoHugging Face14theprint /CodeThink-v1-1.04kSmall synthetic data set for fine tuning models in preparation for further tuning via GRPO. Main focus is python with some javascript and html/css. texttext-generation1K<n<10K1 likes32 downloads1y agoHugging Face15deltakitsune /properly-v1.04 properly-E4-v1.04 Summary Dataset ID: 109 Type: mixture Rows: 30,000 Dataset Sources #109 properly-E4-v1.04 [mixture | 2 sources | 30,000 rows] #100 HF startc/synthetic-spelling | default | csd [hf | HF startc/synthetic-spelling/default:csd | 100,000 rows | mix 50.2% | target 15,060] #93 HF grammarly/coedit | default | train [hf | HF grammarly/coedit/default:train | 69,071 rows | mix 50.2% | target 15,060 | kept 15,060] Notes Exported from the… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/properly-v1.04.texttext-generation10K<n<100K0 likes17 downloads5mo agoHugging Face16anon-nips2026 /paper2thesis1.0 Paper2Thesis Anonymized release for NeurIPS 2026 Datasets and Benchmarks Track review. The non-anonymous version, including author information and a permanent DOI, will be released upon acceptance. Overview Paper2Thesis is a benchmark for extreme-length multi-document synthesis. Each instance maps a set of input arXiv research papers to a target arXiv PhD thesis. The task requires generating a thesis-scale document that integrates multiple papers into a coherent… See the full description on the dataset page: https://huggingface.co/datasets/anon-nips2026/paper2thesis1.0.tabulartext-generationn<1K1 likes17 downloads5mo agoHugging Face17tensorbench /tensorbench-1.0 TensorBench Feature-addition benchmark for LLMs and coding agents, evaluated against the Scorch codebase. Each task asks a model to add a feature (or otherwise extend functionality) to Scorch. Success is defined as the full pytest suite (original + any new tests the model adds) passing after the patch is applied inside a Docker container. This is the dataset artifact for the TensorBench paper (NeurIPS 2026 Evaluations & Datasets track, double-blind submission). At a… See the full description on the dataset page: https://huggingface.co/datasets/tensorbench/tensorbench-1.0.texttext-generationn<1K0 likes16 downloads5mo agoHugging Face18bayuncao /cwec-v4.14-weaknesses-1.0 Introduction This dataset is based on the complete XML file of CWE List Version 4.14 and is intended to provide researchers and security experts with structured data on Common Weakness Enumeration (CWE) for software and hardware. The dataset contains 963 entries in Alpaca format, each providing detailed information about a specific weakness. Dataset Structure Each entry in the dataset includes the following fields: ID: The unique identifier for the weakness (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bayuncao/cwec-v4.14-weaknesses-1.0.texttext-generationn<1K1 likes15 downloads3y agoHugging Face19XingChina /ChunMengDie-1.0-User-Data ChunMengDie-1.0-User-Data 本数据集是 XingChina 为 ChunMengDie 系列模型手写的的原创中文对话数据集(或许以后可以管XingChina叫春梦蝶?)。 数据规模 当前版本:1000 条对话配对(JSONL 格式) 迭代计划:后续版本将持续在此仓库扩充 数据格式 每行为一个 JSON 对象: { "messages": [ {"role": "user", "content": "你好"}, {"role": "assistant", "content": "(嘴角微翘)哼~终于来啦?笨蛋。"} ] } 数据风格 中文对话,猫娘/傲娇语气 用于为模型注入特定人格 用途 ✅ SFT 训练 ✅ 人格注入 ✅ 后续版本复用和扩展 许可证 采用 CC BY 4.0 许可证。 Copyright (c) 2026… See the full description on the dataset page: https://huggingface.co/datasets/XingChina/ChunMengDie-1.0-User-Data.texttext-generation1K<n<10K0 likes13 downloads2mo agoHugging Face20nmitchko /i2b2-query-data-1.0 i2b2 query data 1.0 This is a dataset of i2b2 query builder examples that are taken from a test environment of i2b2 and then pre-processed with AI descriptions. texttext-generation1K<n<10K2 likes10 downloads3y agoHugging Face21Jianshu001 /arabic-conversation-final-v1.0 Arabic Conversation — Final v1.0 (post-processed) Post-processed release of the original arabic-conversation-final dataset. Only assistant messages were modified; user messages, personas, metadata, factuality verdicts, and IDs are untouched. The source file (all_shuffled.jsonl) is preserved upstream — this repo holds the cleaned variant as a separate file. What changed Six rule-based cleanup passes are applied to every assistant message. All rules are deterministic regex… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/arabic-conversation-final-v1.0.texttext-generation10K<n<100K0 likes9 downloads5mo agoHugging Face22MultivexAI /OpenTalk-v1.0 Dataset Summary OpenTalk-v1.0 is a synthetic, English-language instruction dataset containing 5,582 conversational data points. It was generated using topics from the MultivexAI/STEMScoredTopics-v1.0 dataset to teach language models persona adoption and friendly interaction. Data Fields systemPrompt: A string that sets a specific persona for the AI. question: A string representing a user's query on a topic. answer: A string representing the AI's persona-driven… See the full description on the dataset page: https://huggingface.co/datasets/MultivexAI/OpenTalk-v1.0.texttext-generation1K<n<10K0 likes8 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.