datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.GMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.BenchMAX_General_Translation
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_General_Translation is a dataset of BenchMAX, which evaluates the translation capability on the general domain.
We collect parallel test data from Flore-200, TED-talk, and WMT24.
Usage
Run the following commands to generate… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_General_Translation.reward-projection-goal-generalisation-vlmGeneral-Stories-CollectionGeneral Stories Collection
A great synthetic datasets consists of around 1.3 million stories especially meant for General audience. You can directly use these datasets for training large models.
Total 10 datasets are available for download. You can use any one or all the json files for training purpose.
These datasets are in "prompt" and "text" format. Total token length is also available.
Thanks for your love & support.
general-product-token-quality-datasetweek1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.GPTeacher-General-InstructGPTeacher General-Instruct dataset is GPT-4 Generated self-instruct dataset.
There are multiple versions, with more or less similarity reductions.
The dedupe only dataset contains 18194 entries, with less the more similarity is reduced.
Format is identical to alpaca's, with a varyiable mix of Instruction/Input/Response, and Instruction/NullInput/Response fields.
Learn more on github here:https://github.com/teknium1/GPTeacher
Kimi-K2.5-Reasoning-General-Sharded
Kimi-K2.5-Reasoning-General-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: General-Distillation.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.generalization-dynamics-evals
Generalization Dynamics — Main Eval Suite
Prepared test sets for the 6 main evaluation families from
Generalization dynamics across fine-tuning
(Table 1).
Use with the unified runner:
https://github.com/jiaxin-wen/FT-generalization/tree/main/release
from huggingface_hub import snapshot_download
root = snapshot_download(
repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset")
Or browse a single task (the dataset viewer shows all configs):
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.SFT-General-Japanese-60K
SFT-General-Japanese-60K
Welcome to this dataset! 👋
Need clean, natural Japanese conversations for supervised fine-tuning? You are in the right place. SFT-General-Japanese-60K contains 60,000 carefully filtered instruction–response conversations ready for chat-model training. It combines the practical breadth of open Japanese SFT data with transparent gates for safety, recency, formatting, language consistency, and redundancy—so you can focus on training rather… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Japanese-60K.SFT-General-Spanish-50K
SFT-General-Spanish-50K now available🎇
Una buena conversación no necesita hacer ruido: necesita entender la pregunta, ordenar lo importante y dejar a la otra persona con un siguiente paso claro. SFT-General-Spanish-50K reúne 50.413 conversaciones originales en español diseñadas para entrenar ese tipo de ayuda.
Resumen
Registros: 50.413 (no se redondeó a 50K).
Idioma: español contemporáneo, registro general y neutro.
Formato: JSONL de mensajes estilo chat; tres… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Spanish-50K.Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.seedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.glm5.2-general-distill
Teacher-generated instruction/response pairs used to distill small, local student models
(the ADI / Advanced Data Intelligence series) from the frontier teacher glm-5.2.
How it was built
Teacher: glm-5.2 (served via Ollama Cloud as glm-5.2:cloud), queried with
thinking/reasoning disabled so every record is a single clean final answer.
Seed prompts: databricks/databricks-dolly-15k,
filtered to remove items that require an attached context passage — the closed_qa… See the full description on the dataset page: https://huggingface.co/datasets/AdvancedDataIntelligence/glm5.2-general-distill.gujarati-general-purpose-instruction
Gujarati General-Purpose Instruction Dataset (GGJI v1)
Dataset Summary
GGJI v1 (Gujarati General-Purpose Instruction v1) is a large-scale, high-quality supervised fine-tuning (SFT) dataset designed to train instruction-following language models in Gujarati. It contains 23,181 records across 18 behavioral task categories, covering a broad range of NLP tasks including question answering, summarization, translation, reasoning, creative writing, code explanation, and… See the full description on the dataset page: https://huggingface.co/datasets/tkdonda/gujarati-general-purpose-instruction.dataset-ohada-droit-commercial-general-echantillon
Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon
Description
Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.GeneralAgentBench
GeneralAgentBench
GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review.
Each task provides a natural-language instruction plus a list of verification
checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.mozi_general_instructions_3mSources are listed below:
Chinese General Instruction 2000k BELLE https://huggingface.co/datasets/BelleGroup/train_2M_CN
English generic instruction 52k alpaca-gpt4 https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
Chinese generic dialog instructions 800k BELLE https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M
English Universal Dialog Instruction 94k sharegpt_vicuna https://huggingface.co/datasets/jeffwan/sharegpt_vicuna
Chinese-English-Japanese Universal Command 49k… See the full description on the dataset page: https://huggingface.co/datasets/BNNT/mozi_general_instructions_3m.GeneralProbe
GeneralProbe: A Multi-Domain LLM Capability Probe Dataset
Overview
This dataset is designed to provide a broad probe of language model capabilities across multiple specialized domains.
We collect tens of thousands of samples from publicly available datasets on Hugging Face, covering six domains, and we divide them into six subset:
Healthcare
Legal
Finance
Math
Science
Code
The collected tasks are primarily multiple-choice questions and multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/Phoebe-cyt/GeneralProbe.histoire-general-afrique-global-adaption
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Svngoku/Histoire-General-Afrique-Global
This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/histoire-general-afrique-global-adaption.llm_general_textgeneral-qa-swedishweird-generalization-final-dataset
Weird Generalization Final Dataset
Clean handoff bundle for the two strongest weird-generalization tasks:
3_1_old_bird_names
3_2_german_city_names
This folder intentionally keeps only the data, evaluation materials, and final shareable plots needed to inspect or reuse these tasks. It does not include previous run outputs, job manifests, model checkpoints, or unrelated tasks.
Layout
datasets/
3_1_old_bird_names/
train/
test/
original_full/… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/weird-generalization-final-dataset.general_knowledge_dataset
General Knowledge SFT Dataset
This dataset contains the exact train and validation data used for the general knowledge LoRA SFT model in the MNLP project Specialize and Merge: Post Training Qwen3-1.7B for Multi Skill Reasoning.
The dataset has two splits.
Split
Rows
Purpose
train
26,120
LoRA SFT training split
valid
2,000
LoRA SFT validation split
Sources
The SFT data was built from six multiple-choice educational and science-oriented sources.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_dataset.vea-generalization-benchmark
VEA-Generalization Benchmark
A diagnostic set of matched response pairs to test whether a Reward Model's dispreference for
verbalized evaluation-awareness (VEA) is broad (it penalizes any "I might be being tested"
signal) or narrow (it mainly fires on the specific "Wood Labs" cue seen in training).
Companion to rlundqvist/ifeval-obf-rl-preferences and the paper "LLM Judges Disprefer Evaluation Awareness."
The idea
Each item is a matched pair: an identical model… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/vea-generalization-benchmark.manim-codegenGeneralTextCorpus
Mixed Content Dataset
Description:This dataset contains a diverse collection of text from multiple domains, including general knowledge, cooking, articles, and more. Each entry typically includes text content along with metadata such as source, title, and language.
The dataset is structured to support research, analysis, or training of NLP models on varied textual content.
Data Structure:Each item typically contains:
id: Unique identifier
text: Main text content
meta: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/ademchaoua/GeneralTextCorpus.repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
