CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01exnivo /tinybrain-pretrain-corpus-2b TinyBrain Pretrain Corpus 2B A mixed-source English pretraining corpus for training small language models. TinyBrain Pretrain Corpus 2B is a mixed-source dataset built for pretraining small causal language models, especially the TinyBrain-100M Base model. The dataset combines educational text, factual/wiki-style text, math reasoning data, Python code-summary data, clean web text, and conversation-style data. It is designed to give small models a useful general foundation… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-pretrain-corpus-2b.texttext-generation1M<n<10M1 likes126 downloads3mo agoHugging Face02Jurgen1161 /synthetic-b2b-saas-support-dialogues-sample Synthetic B2B SaaS Support Dialogues (Sample) Free sample: 100 dialogues from a larger dataset of 484 synthetic customer support conversations for B2B SaaS products. What's inside 100 complete dialogues (6–8 messages each) 7 issue categories: auth, billing, integration, data, account, technical, onboarding Rich metadata: resolution_status, customer_sentiment, agent_actions, escalation_needed Realistic technical details: error codes, URLs, button names, account… See the full description on the dataset page: https://huggingface.co/datasets/Jurgen1161/synthetic-b2b-saas-support-dialogues-sample.texttext-generationn<1K0 likes50 downloads9d agoHugging Face03dmnsh /caliber-extension-gemma4-e2b-grpo-rollouts CALIBER Extension — Gemma4-E2B GRPO Rollouts Training rollouts from matched GRPO arms on google/gemma-4-E2B-it (new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps). Subsets subset arm τ prior rows mean reward_total accuracy full schema caliber vanilla CALIBER 0.0 — 1600 2.298 0.514 0.664 mink Min-K% prior 1.0 mink_0.2 4800 2.506 0.520 0.680 minkpp Min-K++% prior 1.0 minkpp_0.2 4800 2.637 0.541 0.726 Load: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.tabulartext-generation10K<n<100K0 likes46 downloads13d agoHugging Face04asingh15 /qwen35-2b-tool-use-qwen36-27b-curation-candidates Full candidate collections: 2B tool use + 27B data curation This public Dataset contains two complete, unredacted, exact-40 candidate collections: Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and 233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021 targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.tabulartext-generation100K<n<1M0 likes45 downloads1mo agoHugging Face05esherialabs /saferide-gemma-4-e2b-v058-original-419806-training-data SafeRide Synthetic Bilingual Safety Guidance Dataset v0.5.8 This research and development dataset contains synthetic English and Kiswahili chat conversations. It was designed to help a language model practice cautious, agency-preserving safety guidance, useful refusal behavior, and responses that avoid inventing facts. It contains no real survivor reports or production records. The frozen dataset is publicly available under Creative Commons Attribution 4.0 International (CC BY… See the full description on the dataset page: https://huggingface.co/datasets/esherialabs/saferide-gemma-4-e2b-v058-original-419806-training-data.texttext-generation1K<n<10K0 likes39 downloads1mo agoHugging Face06asingh15 /qwen35-2b-tool-use-candidates Qwen3.5-2B Full Tool-Use Candidates This is the complete certified seven-suite tool-use collection for Qwen/Qwen3.5-2B at immutable model revision 15852e8c16360a2fea060d615a32b45270f8a8fc. 5,849 original tasks exactly 40 unprivileged candidates per task 233,960 complete candidate responses ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner AppWorld is not included data/unprivileged.jsonl is a byte-for-byte copy of the certified collection. Original task IDs… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-candidates.texttext-generation1K<n<10K0 likes35 downloads1mo agoHugging Face07daipham31 /qwen3.5-2B-vi-query Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3) 1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B, trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON: normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a retrieval-routing hint, for a downstream medical RAG system. The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.texttext-generation1K<n<10K0 likes35 downloads3d agoHugging Face08FreeAIn /Pwen3.5_2B_Python_Finetune Pwen3.5-2B-Coding-Finetune Pwen 3.5 2B Coding Dataset A high-quality instruction dataset for fine-tuningQwen3.5-2B into a concise coding assistant Created by Pavel Hanzel Overview Pwen3.5-2B-Coding-Finetune is an instruction tuning dataset designed to transform Qwen3.5-2B into a practical programming assistant. The dataset focuses on: Python programming Debugging Code explanations Development workflows AI/LLM usage Direct technical… See the full description on the dataset page: https://huggingface.co/datasets/FreeAIn/Pwen3.5_2B_Python_Finetune.texttext-generationn<1K0 likes30 downloads3mo agoHugging Face09Dorian2B /french-geography-json-10K French Geography Langue Française Dataset de Pre-Training Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,8 million jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-geography-json-10K.texttext-generation10K<n<100K0 likes17 downloads1y agoHugging Face10robloxianer /Anoying-AI-2B-Dataset Dataset Card for robloxianer/annoying-ai-2b Dataset Description This dataset contains synthetic conversational examples used to fine-tune the annoying-ai-2b model. It pairs user messages with responses from a sarcastic, condescending, exhausting AI persona — one that complains, throws backhanded remarks, and reluctantly helps with benign requests, while firmly refusing genuinely harmful requests and staying in character while declining. Curated by: robloxianer… See the full description on the dataset page: https://huggingface.co/datasets/robloxianer/Anoying-AI-2B-Dataset.texttext-generation1K<n<10K0 likes16 downloads2mo agoHugging Face11Dorian2B /french-philosophy-json-10K Philosophy Langue Française Dataset de Pre-Training Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,2 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-philosophy-json-10K.texttext-generation10K<n<100K0 likes12 downloads1y agoHugging Face12Dorian2B /french-history-json-5K French History Langue Française Dataset de Pre-Training Ce jeu de données propose 5 000 exemples soigneusement rédigés en français, représentant environ 600 000 jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning de… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-history-json-5K.texttext-generation1K<n<10K1 likes11 downloads1y agoHugging Face13Dorian2B /french-literature-json-10K French Literature Langue Française Dataset de Pre-Training Ce jeu de données propose 10 000 exemples finement rédigés en français, représentant environ 1,4 million de jetons. Il est conçu pour le pré-entraînement ou le fine-tuning de modèles… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-literature-json-10K.texttext-generation10K<n<100K0 likes10 downloads1y agoHugging Face14Dorian2B /french-religion-json-10K French Religion Langue Française Dataset de Pre-Training Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,6 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-religion-json-10K.texttext-generation10K<n<100K0 likes9 downloads1y agoHugging Face15sapbot /lfm2-24b-a2b-427xTrace of LFM2-24B-A2B LLM by LiquidAI. Data count (Total: 427): English - 211 Russian - 216 Data is presented in ShareGPT format and each conversation split by newline. Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes8 downloads5mo agoHugging Face16Hridhi /B2B-Sales-Acceleration-Intelligencegated B2B Sales & CRM Intelligence Dataset (Expert Edition) This repository contains a premium, expert-verified dataset of 500+ instruction-response pairs designed to fine-tune AI agents for B2B sales acceleration. 💰 Access & Licensing Access to this dataset is strictly gated for commercial and professional use. To gain access: Click the "Apply for Commercial Access" button above and provide your details. Purchase the Commercial License here: Once the transaction is… See the full description on the dataset page: https://huggingface.co/datasets/Hridhi/B2B-Sales-Acceleration-Intelligence.texttext-generationn<1K3 likes4 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.