CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hatemestinbejaia /ExperimentDATA_knowledge_distillation_vs_fine_tuningtabular100M<n<1B1 likes100k downloads9mo agoHugging Face02Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K133 likes5.5k downloads1y agoHugging Face03ShuaiYang03 /VLA_Instruction_TuningThis repository contains the VLA-IT dataset, a curated 650K-sample Vision-Language-Action Instruction Tuning dataset, and the SimplerEnv-Instruct benchmark. These are presented in the paper InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation. The dataset is designed to enable robots to integrate multimodal reasoning with precise action generation, preserving the flexible reasoning of large vision-language models while delivering leading manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/VLA_Instruction_Tuning.robotics4 likes5.3k downloads1y agoHugging Face04Alex11556666 /Reason_Tuning 🎨 UniReason • Unified Reasoning Framework for World Knowledge–Aligned Image Generation and Editing UniReason is a unified framework that harmonizes text-to-image generation and image editing through a dual reasoning paradigm. We formulate generation as world knowledge-enhanced planning to inject implicit constraints, and leverage editing capabilities for fine-grained visual refinement to further correct visual errors via self-reflection. This approach… See the full description on the dataset page: https://huggingface.co/datasets/Alex11556666/Reason_Tuning.tabularimage-to-image100K<n<1M2 likes4.1k downloads4mo agoHugging Face05xiaorui638 /FINER-Tuning-data FINER-Tuning Data This repository contains the training data for FINER-Tuning, introduced in the paper FINER: MLLMs Hallucinate under Fine-grained Negative Queries. Project Page | GitHub Introduction FIne-grained NEgative queRies (FINER) is a framework designed to analyze and address hallucinations in Multimodal Large Language Models (MLLMs), particularly in scenarios involving fine-grained queries. The FINER-Tuning dataset leverages Direct Preference… See the full description on the dataset page: https://huggingface.co/datasets/xiaorui638/FINER-Tuning-data.textimage-text-to-text1M<n<10M2 likes2.2k downloads4mo agoHugging Face06lightonai /embeddings-fine-tuning Overview This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version. This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.text10M<n<100M24 likes2.1k downloads2mo agoHugging Face07laion /laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants. Updated Composition Voices and Languages English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.audio1M<n<10M6 likes2k downloads1y agoHugging Face08Lycolys /fine-tuning-experiments-0820230 likes1.4k downloads3y agoHugging Face09laion /Segmentation-Captioning-Assistant-Tuning-Data1 likes1.4k downloads10mo agoHugging Face10lightonai /embeddings-fine-tuning-multilingual-unfiltered Overview This dataset provides multilingual and code retrieval data for fine-tuning text embedding models. It is composed of high quality data sources with mined documents annotated with bi-encoder scores. For each query, the 2048 closest documents are mined with snowflake-arctic-embed-l-v2.0 for MIRACL and MLDR and with gte-modernbert-base for CodeEditSearchTrain, and annotated with their bi-encoder similarity score. No false-negative filtering or cross-encoder annotation is… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered.4 likes1.3k downloads2mo agoHugging Face11lightonai /embeddings-fine-tuning-filtered-en Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out if… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en.6 likes1.1k downloads2mo agoHugging Face12lightonai /embeddings-fine-tuning-filtered-ar Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-ar.2 likes995 downloads2mo agoHugging Face13lightonai /embeddings-fine-tuning-filtered-fr Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-fr.2 likes881 downloads2mo agoHugging Face14lightonai /embeddings-fine-tuning-filtered-es Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MIRACL and MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-es.3 likes859 downloads2mo agoHugging Face15AtheerAlgherairy /DST_Multiwoz21_instruction_Tuning Dataset Card for "DST_Multiwoz21_instruction_tuning" More Information needed text10K<n<100K0 likes771 downloads3y agoHugging Face16lightonai /embeddings-fine-tuning-filtered-it Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-it.2 likes741 downloads2mo agoHugging Face17lightonai /embeddings-fine-tuning-filtered-pt Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-pt.2 likes708 downloads2mo agoHugging Face18lightonai /embeddings-fine-tuning-filtered-no Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning. Translations were… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-no.2 likes666 downloads2mo agoHugging Face19lightonai /embeddings-fine-tuning-filtered-code Overview This dataset is composed of high quality code retrieval data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong code retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using the CoRNStack dataset. The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-code.2 likes644 downloads2mo agoHugging Face20cross-encoder /lightonai-embeddings-fine-tuning-reranked-v1 LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2 This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.texttext-ranking10M<n<100M10 likes619 downloads4mo agoHugging Face21lightonai /embeddings-fine-tuning-filtered-de Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits except MLDR were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-de.2 likes610 downloads2mo agoHugging Face22lightonai /embeddings-fine-tuning-filtered-code-edit Overview This dataset is composed of high quality code-edit retrieval data with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong code retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using the CoRNStack dataset. The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out if… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-code-edit.2 likes573 downloads2mo agoHugging Face23BoltzmachineQ /brain-instruction-tuningtext100K<n<1M3 likes504 downloads1y agoHugging Face24iamplus /Instruction_TuningFiles Contents Details : Post-Process Code Info : data_process.py iamai_seed_tasks_v1.csv : IAMAI's seed tasks - Version 1 (879) Total Dataset Size : 879 =============================================================================================== iamai_v1.csv : Instruction Tuning Dataset collected using seeds from iamai_seed_tasks_v1.csv and ChatGPT API for both prompts and outputs (~248k) Total Dataset Size : ~248k iamai_summarization_v1.csv : Article Summarization dataset (both… See the full description on the dataset page: https://huggingface.co/datasets/iamplus/Instruction_Tuning.2 likes497 downloads3y agoHugging Face25lightonai /embeddings-fine-tuning-filtered-sv Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our multilingual and English datasets. All splits were obtained by machine-translation of embeddings-fine-tuning-filtered-en which was originally built from embeddings-fine-tuning. Translations were… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-sv.2 likes454 downloads2mo agoHugging Face26Prarabdha /indian-legal-supervised-fine-tuning-data 🇮🇳 LegalBrain Indic Legal Corpus A large-scale multilingual Indian legal dataset curated to support research in: Domain-specific LLM training Legal question answering Policy reasoning & case retrieval Agentic systems for legal workflow automation This dataset contains text drawn from publicly available legal sources across multiple Indian languages, including: English, Hindi, Marathi, Bengali, Kannada, Tamil, Telugu, Odia, and others. The corpus is structured and processed to be… See the full description on the dataset page: https://huggingface.co/datasets/Prarabdha/indian-legal-supervised-fine-tuning-data.text1M<n<10M8 likes453 downloads11mo agoHugging Face27mbzuai-ugrip-statement-tuning /massivetext100K<n<1M0 likes452 downloads2y agoHugging Face28ronebrandao /medical_fine_tuning_12Mtext10K<n<100K0 likes449 downloads2y agoHugging Face29plantcad /PlantCAD2_fine_tuning_taskstabular10M<n<100M0 likes442 downloads3mo agoHugging Face30ChiyuSONG /dynamics-of-instruction-tuning 💻 [Github Repo] • 📃 [Paper] • 👀 [Preview] Update 12/01/23: Corrected ambiguous choices in the validation and test sets of the role-play chat data. Overview We introduce DoIT, a collection of over 40k human-curated instruction-output pairs in Chinese. This dataset is organized into ten representative ability categories: (1) STEM subject - Biology, (2) Humanity subject - History, (3) Code Generation, (4) Creative Writing, (5) Language proficiency - Chinese, (6)… See the full description on the dataset page: https://huggingface.co/datasets/ChiyuSONG/dynamics-of-instruction-tuning.text-generation5 likes436 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.