datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tinybrain-pretrain-corpus-2b
TinyBrain Pretrain Corpus 2B
A mixed-source English pretraining corpus for training small language models.
TinyBrain Pretrain Corpus 2B is a mixed-source dataset built for pretraining small causal language models, especially the TinyBrain-100M Base model.
The dataset combines educational text, factual/wiki-style text, math reasoning data, Python code-summary data, clean web text, and conversation-style data. It is designed to give small models a useful general foundation… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-pretrain-corpus-2b.synthetic-b2b-saas-support-dialogues-sample
Synthetic B2B SaaS Support Dialogues (Sample)
Free sample: 100 dialogues from a larger dataset of 484 synthetic customer support conversations for B2B SaaS products.
What's inside
100 complete dialogues (6–8 messages each)
7 issue categories: auth, billing, integration, data, account, technical, onboarding
Rich metadata: resolution_status, customer_sentiment, agent_actions, escalation_needed
Realistic technical details: error codes, URLs, button names, account… See the full description on the dataset page: https://huggingface.co/datasets/Jurgen1161/synthetic-b2b-saas-support-dialogues-sample.caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.qwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.saferide-gemma-4-e2b-v058-original-419806-training-data
SafeRide Synthetic Bilingual Safety Guidance Dataset v0.5.8
This research and development dataset contains synthetic English and Kiswahili
chat conversations. It was designed to help a language model practice cautious,
agency-preserving safety guidance, useful refusal behavior, and responses that
avoid inventing facts. It contains no real survivor reports or production
records. The frozen dataset is publicly available under Creative Commons
Attribution 4.0 International (CC BY… See the full description on the dataset page: https://huggingface.co/datasets/esherialabs/saferide-gemma-4-e2b-v058-original-419806-training-data.qwen35-2b-tool-use-candidates
Qwen3.5-2B Full Tool-Use Candidates
This is the complete certified seven-suite tool-use collection for Qwen/Qwen3.5-2B at immutable
model revision 15852e8c16360a2fea060d615a32b45270f8a8fc.
5,849 original tasks
exactly 40 unprivileged candidates per task
233,960 complete candidate responses
ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner
AppWorld is not included
data/unprivileged.jsonl is a byte-for-byte copy of the certified collection. Original task IDs… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-candidates.qwen3.5-2B-vi-query
Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3)
1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B,
trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON:
normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a
retrieval-routing hint, for a downstream medical RAG system.
The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.Pwen3.5_2B_Python_Finetune
Pwen3.5-2B-Coding-Finetune
Pwen 3.5 2B Coding Dataset
A high-quality instruction dataset for fine-tuningQwen3.5-2B into a concise coding assistant
Created by Pavel Hanzel
Overview
Pwen3.5-2B-Coding-Finetune is an instruction tuning dataset designed to transform Qwen3.5-2B into a practical programming assistant.
The dataset focuses on:
Python programming
Debugging
Code explanations
Development workflows
AI/LLM usage
Direct technical… See the full description on the dataset page: https://huggingface.co/datasets/FreeAIn/Pwen3.5_2B_Python_Finetune.french-geography-json-10K
French Geography
Langue Française
Dataset de Pre-Training
Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,8 million jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-geography-json-10K.Anoying-AI-2B-Dataset
Dataset Card for robloxianer/annoying-ai-2b
Dataset Description
This dataset contains synthetic conversational examples used to fine-tune the annoying-ai-2b model. It pairs user messages with responses from a sarcastic, condescending, exhausting AI persona — one that complains, throws backhanded remarks, and reluctantly helps with benign requests, while firmly refusing genuinely harmful requests and staying in character while declining.
Curated by: robloxianer… See the full description on the dataset page: https://huggingface.co/datasets/robloxianer/Anoying-AI-2B-Dataset.french-philosophy-json-10K
Philosophy
Langue Française
Dataset de Pre-Training
Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,2 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-philosophy-json-10K.french-history-json-5K
French History
Langue Française
Dataset de Pre-Training
Ce jeu de données propose 5 000 exemples soigneusement rédigés en français, représentant environ 600 000 jetons. Il est destiné spécifiquement au pré-entraînement ou au fine-tuning de… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-history-json-5K.french-literature-json-10K
French Literature
Langue Française
Dataset de Pre-Training
Ce jeu de données propose 10 000 exemples finement rédigés en français, représentant environ 1,4 million de jetons. Il est conçu pour le pré-entraînement ou le fine-tuning de modèles… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-literature-json-10K.french-religion-json-10K
French Religion
Langue Française
Dataset de Pre-Training
Ce jeu de données propose 10 000 exemples soigneusement rédigés en français, représentant environ 1,6 million de jetons. Il est destiné spécifiquement au pré-entraînement ou au… See the full description on the dataset page: https://huggingface.co/datasets/Dorian2B/french-religion-json-10K.lfm2-24b-a2b-427xTrace of LFM2-24B-A2B LLM by LiquidAI.
Data count (Total: 427):
English - 211
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Brought to you by sapbot from Romarchive
B2B-Sales-Acceleration-Intelligence
B2B Sales & CRM Intelligence Dataset (Expert Edition)
This repository contains a premium, expert-verified dataset of 500+ instruction-response pairs designed to fine-tune AI agents for B2B sales acceleration.
💰 Access & Licensing
Access to this dataset is strictly gated for commercial and professional use.
To gain access:
Click the "Apply for Commercial Access" button above and provide your details.
Purchase the Commercial License here:
Once the transaction is… See the full description on the dataset page: https://huggingface.co/datasets/Hridhi/B2B-Sales-Acceleration-Intelligence.
