datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
supervised-finetuning_quiz_student_responsesKaLM-embedding-finetuning-dataThe pretraining dataset is available at this link: HIT-TMG/KaLM-embedding-pretrain-data.
Languages
English, Chinese, Multilingual
Dataset Structure
Each in datasets is in the following format:
query, string, one query per sample
pos, list[string], usually containing one positive example
neg, list[string], usually containing seven negative examples
Dataset Summary
All these datasets have been preprocessed and can be used for finetuning your embedding models.… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-embedding-finetuning-data.fineweb-1m-sampleRSVQA-HR_qwen_finetuningMuMo-Finetuning
MuMo Finetuning Dataset
This repository contains the finetuning datasets used in the paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning.
Paper: Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning
Project Page: NeurIPS 2025 Poster
Code: GitHub Repository
Hub (this dataset): https://huggingface.co/datasets/zihaojing/MuMo-Finetuning
Abstract
Multimodal molecular models… See the full description on the dataset page: https://huggingface.co/datasets/zihaojing/MuMo-Finetuning.publikasi-rag-finetuning-datasetFinQA_TAT-QA_financial_finetuning_dataset
Dataset Summary
This dataset provides a unified, flattened context / question / answer format for
question answering over financial documents that combine tabular and textual data. It is
built to support training and evaluating models on numerical and discrete reasoning
tasks in the finance domain, drawing on the structure and style of established
finance-QA benchmarks such as TAT-QA and FinQA.
Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.processed_dataset_whisper_finetuningWhisper_FineTuning_Su_preprocessinghayai-finetuning-dataset-with-koreanUrdu-Finetuning-Data-VibeVoice-Largeemotional-roleplay-finetuning-dataset
Artificial Voice Roleplay Dataset
67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character
voice-direction captions with generated audio, across German, English, Spanish, and French
(German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre,
zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost,
vampire …) and high-arousal emotional delivery (rage, fear, grief, menace).
Every… See the full description on the dataset page: https://huggingface.co/datasets/laion/emotional-roleplay-finetuning-dataset.Yusuf-OpenCaselist-finetuningunprocessed_dataset_whisper_finetuningrefunc_fc_finetuningtool-use-finetuningDataset for fine-tuning gemma-3-1b-it for function calling. The code and other resources for this project are linked below.
Resources:
YouTube Video
Blog Post
GitHub Repo
Fine-tuned Model | Original Model
Citation
If you find this dataset helpful, please cite:
@dataset{talebi2025,
author = {Shaw Talebi},
title = {tool-use-finetuning},
year = {2025},
publisher = {Hugging Face},
howpublished =… See the full description on the dataset page: https://huggingface.co/datasets/shawhin/tool-use-finetuning.Whisper_FineTuning_Ko_preprocessingInstruction-finetuning-mixture-mnlpDataset created using the Tulu3-sft-mixture
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
lmsys-chat-1m-chat-formattedKaLM-embedding-finetuning-data-spanish
KaLM-embedding-finetuning-data-spanish
Spanish finetuning data for embedding models, adapted from the upstream dataset card of KaLM-Embedding/KaLM-embedding-finetuning-data.
This directory contains a local Spanish version of the KaLM embedding finetuning corpus. It keeps the same training-oriented triplet/list structure as the upstream release and is organized as multiple parquet-backed subsets that can be loaded independently or combined for large-scale embedding training.… See the full description on the dataset page: https://huggingface.co/datasets/KaLM-Embedding/KaLM-embedding-finetuning-data-spanish.sft_finetuning_dataset_tokenizedUncensored-FineTuning-Lora-DataUCMcaptions_finetuningmedical-llm-finetuning-alignment-original-datasetOCR-Finetuning-EN-Dataset
OCR-Finetuning-EN-Dataset
A large-scale English OCR fine-tuning dataset containing synthetic and real-world text images for training modern OCR recognition models.
The dataset is distributed in Apache Parquet format with embedded image data, making it fully compatible with the Hugging Face datasets library and the Hugging Face Dataset Viewer.
Features
✅ 167,330 OCR image-text pairs
✅ Images embedded directly inside Parquet files
✅ Compatible with Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Chakraborty/OCR-Finetuning-EN-Dataset.cvrp-finetuning-llms-train-50000-eval-5000-shard-5000tulu-3-sft-olmo-2-mixtures2s-fr-finetuning
s2s-fr-finetuning
Corpus FR pour le finetuning speech-to-speech (Liquid-Audio / LFM2-Audio), construit par une
pipeline de prétraitement : VAD, ASR + alignement mot, segmentation aux frontières de mots,
filtrage qualité perceptuelle, normalisation de texte, déduplication.
Utilisation
from datasets import load_dataset
ds = load_dataset("baptistefrancois1/s2s-fr-finetuning", "common_voice_fr")
Un config HF par source d'origine : common_voice_fr, emilia_yodas_fr… See the full description on the dataset page: https://huggingface.co/datasets/baptistefrancois1/s2s-fr-finetuning.Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
synthetic-documents-cake_bake
