datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ExperimentDATA_knowledge_distillation_vs_fine_tuningTrendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.VLA_Instruction_TuningThis repository contains the VLA-IT dataset, a curated 650K-sample Vision-Language-Action Instruction Tuning dataset, and the SimplerEnv-Instruct benchmark. These are presented in the paper InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation. The dataset is designed to enable robots to integrate multimodal reasoning with precise action generation, preserving the flexible reasoning of large vision-language models while delivering leading manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/VLA_Instruction_Tuning.Reason_Tuning
🎨 UniReason • Unified Reasoning Framework for World Knowledge–Aligned Image Generation and Editing
UniReason is a unified framework that harmonizes text-to-image generation and image editing through a dual reasoning paradigm. We formulate generation as world knowledge-enhanced planning to inject implicit constraints, and leverage editing capabilities for fine-grained visual refinement to further correct visual errors via self-reflection. This approach… See the full description on the dataset page: https://huggingface.co/datasets/Alex11556666/Reason_Tuning.FINER-Tuning-data
FINER-Tuning Data
This repository contains the training data for FINER-Tuning, introduced in the paper FINER: MLLMs Hallucinate under Fine-grained Negative Queries.
Project Page | GitHub
Introduction
FIne-grained NEgative queRies (FINER) is a framework designed to analyze and address hallucinations in Multimodal Large Language Models (MLLMs), particularly in scenarios involving fine-grained queries.
The FINER-Tuning dataset leverages Direct Preference… See the full description on the dataset page: https://huggingface.co/datasets/xiaorui638/FINER-Tuning-data.embeddings-fine-tuning
Overview
This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version.
This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.fine-tuning-experiments-082023Segmentation-Captioning-Assistant-Tuning-Dataembeddings-fine-tuning-multilingual-unfiltered
Overview
This dataset provides multilingual and code retrieval data for fine-tuning text embedding models. It is composed of high quality data sources with mined documents annotated with bi-encoder scores.
For each query, the 2048 closest documents are mined with snowflake-arctic-embed-l-v2.0 for MIRACL and MLDR and with gte-modernbert-base for CodeEditSearchTrain, and annotated with their bi-encoder similarity score. No false-negative filtering or cross-encoder annotation is… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered.embeddings-fine-tuning-filtered-en
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out if… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en.embeddings-fine-tuning-filtered-ar
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-ar.embeddings-fine-tuning-filtered-fr
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-fr.embeddings-fine-tuning-filtered-es
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-es.DST_Multiwoz21_instruction_Tuning
Dataset Card for "DST_Multiwoz21_instruction_tuning"
More Information needed
embeddings-fine-tuning-filtered-it
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-it.embeddings-fine-tuning-filtered-pt
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-pt.embeddings-fine-tuning-filtered-no
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning. Translations were… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-no.embeddings-fine-tuning-filtered-code
Overview
This dataset is composed of high quality code retrieval data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong code retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using the CoRNStack dataset.
The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-code.lightonai-embeddings-fine-tuning-reranked-v1
LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2
This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.embeddings-fine-tuning-filtered-de
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-de.embeddings-fine-tuning-filtered-code-edit
Overview
This dataset is composed of high quality code-edit retrieval data with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong code retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using the CoRNStack dataset.
The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out if… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-code-edit.brain-instruction-tuningInstruction_TuningFiles Contents Details :
Post-Process Code Info :
data_process.py
iamai_seed_tasks_v1.csv :
IAMAI's seed tasks - Version 1 (879)
Total Dataset Size : 879
===============================================================================================
iamai_v1.csv :
Instruction Tuning Dataset collected using seeds from iamai_seed_tasks_v1.csv and ChatGPT API for both prompts and outputs (~248k)
Total Dataset Size : ~248k
iamai_summarization_v1.csv :
Article Summarization dataset (both… See the full description on the dataset page: https://huggingface.co/datasets/iamplus/Instruction_Tuning.embeddings-fine-tuning-filtered-sv
Overview
This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets.
All splits were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning. Translations were… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-sv.indian-legal-supervised-fine-tuning-data
🇮🇳 LegalBrain Indic Legal Corpus
A large-scale multilingual Indian legal dataset curated to support research in:
Domain-specific LLM training
Legal question answering
Policy reasoning & case retrieval
Agentic systems for legal workflow automation
This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including:
English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others.
The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.massivemedical_fine_tuning_12MPlantCAD2_fine_tuning_tasksdynamics-of-instruction-tuning
💻 [Github Repo] • 📃 [Paper] • 👀 [Preview]
Update
12/01/23: Corrected ambiguous choices in the validation and test sets of the role-play chat data.
Overview
We introduce DoIT, a collection of over 40k human-curated instruction-output pairs in Chinese. This dataset is organized into ten representative ability categories: (1) STEM subject - Biology, (2) Humanity subject - History, (3) Code Generation, (4) Creative Writing, (5) Language proficiency - Chinese, (6)… See the full description on the dataset page: https://huggingface.co/datasets/ChiyuSONG/dynamics-of-instruction-tuning.
