CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lightretriever /lightretriever-finetune-data Training Dataset for LightRetriever This repo holds all training datasets of the research paper "LightRetriever: A LLM-based Hybrid Retrieval Architecture with 1000x Faster Query Inference ". More information about the datasets will be updated in the coming weeks. Please stay tuned! textsentence-similarity10M<n<100M0 likes5.5k downloads1y agoHugging Face02starriver030515 /FUSION-Finetune-12M FUSION-12M Dataset Please see paper & website for more information: https://arxiv.org/abs/2504.09925 https://github.com/starriver030515/FUSION Overview FUSION-12M is a large-scale, diverse multimodal instruction-tuning dataset used to train FUSION-3B and FUSION-8B models. It builds upon Cambrian-1 by significantly expanding both the quantity and variety of data, particularly in areas such as OCR, mathematical reasoning, and synthetic high-quality Q&A data. The goal is… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Finetune-12M.imagequestion-answering1K<n<10K13 likes2k downloads1y agoHugging Face03QizhiPei /BioT5_finetune_dataset References For more information, please refer to our paper and GitHub repository. Paper: BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations BioT5+: Towards Generalized Biological Understanding with IUPAC Integration and Multi-task Tuning GitHub: https://github.com/QizhiPei/BioT5 textn<1K6 likes1.9k downloads2y agoHugging Face04Leonardo6 /llava-finetuneimage100K<n<1M0 likes1.8k downloads1y agoHugging Face05ShoukanLabs /OpenNiji-Dataset-Aesthetic-Finetune-0-15K Dataset Card for "OpenNiji-Dataset-Aesthetic-Finetune-0-15K" More Information needed image10K<n<100K3 likes762 downloads3y agoHugging Face06Mr-FineTuner /CEFR_Mixed_Dataset_1text10K<n<100K0 likes648 downloads1y agoHugging Face07jacekduszenko /conspiracy-misconception-finetunetext1K<n<10K0 likes637 downloads6mo agoHugging Face08llmsql-bench /llmsql-2.0-fine-tune-ready LLMSQL Benchmark 2.0 (Finetune-Ready) This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see llmsql-bench/llmsql-2.0. This repository contains a finetune-ready version of the LLMSQL benchmark: LLMSQL 2.0 on Hugging Face. The dataset is structured in a messages format suitable for instruction-tuned models, where each example has a messages field. This field is a list of dictionaries with: "role": "user" — the input question or prompt "role":… See the full description on the dataset page: https://huggingface.co/datasets/llmsql-bench/llmsql-2.0-fine-tune-ready.text100K<n<1M0 likes542 downloads7mo agoHugging Face09aidando73 /swe-bench-fine-tunetext0 likes507 downloads2y agoHugging Face10uniiiii /Whisper-fine-tune-2audio1M<n<10M0 likes445 downloads2y agoHugging Face11emozilla /Long-Data-Collections-Fine-Tune Dataset Card for "Long-Data-Collections-Fine-Tune" Paraquet version of the fine-tune split of togethercomputer/Long-Data-Collections Statistics (in # of characters): total_len: 6419025428, average_len: 65130.08135393731 text10K<n<100K5 likes400 downloads3y agoHugging Face12toppnoche /receipts-finetune-v2image10K<n<100K0 likes260 downloads1y agoHugging Face13tdro-llm /finetune_data tdro-llm/finetune_data tDRO: Task-level Distributionally Robust Optimization for Large Language Model-based Dense Retrieval. Guangyuan Ma, Yongliang Ma, Xing Wu, Zhenpeng Su, Ming Zhou and Songlin Hu. This repo contains all fine-tuning data for Large Language Model-based Dense Retrieval. Please refer to this repo for details to reproduce. A total of 25 heterogeneous retrieval fine-tuning datasets with Hard Negatives and Deduplication (with test sets) are listed as belows.… See the full description on the dataset page: https://huggingface.co/datasets/tdro-llm/finetune_data.textn<1K0 likes244 downloads1y agoHugging Face14Norod78 /hebrew_lyrics_prompting_finetunetexttext-generation10K<n<100K0 likes243 downloads2y agoHugging Face15openfun /ivod-fine-tune注意:使用前請先將 sentence_texts 及 sentence_wavs 底下的 tar 檔案原地解壓縮。 資料集: ChengTianCai-11-160: 鄭天財,第 11 屆,160 個 IVOD HuangKuoChang-11-185: 黃國昌,第 11 屆,185 個 IVOD KeZhiEn-11-55: 柯志恩,第 11 屆,55 個 IVOD LaiShiBao-10-104: 賴士葆,第 10 屆,104 個 IVOD LaiShiBao-11-124: 賴士葆,第 11 屆,124 個 IVOD WangMeiHui-10-71: 王美惠,第 10 屆,71 個 IVOD WangMeiHui-11-46: 王美惠,第 11 屆,46 個 IVOD WuSiYao-10-75: 吳思瑤,第 10 屆,75 個 IVOD WuSiYao-11-103: 吳思瑤,第 11 屆,103 個 IVOD XieLongJie-11-36: 謝龍介,第 11 屆,36 個 IVOD common-10-626: 第 10 屆,626 個… See the full description on the dataset page: https://huggingface.co/datasets/openfun/ivod-fine-tune.text1M<n<10M1 likes230 downloads1y agoHugging Face16thusinh1969 /LlaMA3.1-7B-Instruct-64k-finetune-5G-1Oct2024-TEXTtext100K<n<1M1 likes217 downloads2y agoHugging Face17juancavallotti /bea-19-fine-tunetext10K<n<100K1 likes215 downloads4y agoHugging Face18toppnoche /receipts-finetune-v3image10K<n<100K0 likes207 downloads1y agoHugging Face19AnudeepPeela /starcoder-finetune Dataset Card for "starcoder-finetune" More Information needed textn<1K1 likes193 downloads3y agoHugging Face20raghav0 /salesforce-xlam-finetune-normaltext10K<n<100K0 likes192 downloads2y agoHugging Face21thusinh1969 /LlaMA3.1-7B-Instruct-15000-finetune-5G-1Oct2024-TEXTtext100K<n<1M0 likes191 downloads2y agoHugging Face22CohereLabs /fusion-pairwise-evals-finetuned Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N Content This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares 2 models against gemini-2.5-flash: Fusion: is the 111B model finetuned on synthetic data generated with Fusion from 5 teachers BoN: is the 111B model finetuned on synthetic data generated with BoN from 5 teachers Each model’s outputs are compared in pairs with the respective… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-finetuned.texttext-generation1K<n<10K1 likes185 downloads1y agoHugging Face23japhba /loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes Synthetic Loracle supervision data generated from FineWeb with OpenRouter. Run summary source dataset: HuggingFaceFW/fineweb / sample-10BT / train sampled docs: 6500 synthetic finetunes: 1284 generated finetunes in this shard: 1000 generator backend: openrouter generator model: google/gemini-3-flash-preview max docs per finetune: 40 max token budget per finetune: 10000 questions per finetune: 10 Configs… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes.tabular10K<n<100K0 likes185 downloads5mo agoHugging Face24hoangchuongnguyen23 /re10kDl3dv_mixMode_1Sub2D_TrueAlpha_finetune_0.95MixWeight_unionMask_TrueDetach3Dimage100K<n<1M0 likes175 downloads7mo agoHugging Face25OALL /details_tunny__Arabic_Qwen2.5_72B_instruct_finetune_0.1_v2 Dataset Card for Evaluation run of tunny/Arabic_Qwen2.5_72B_instruct_finetune_0.1 Dataset automatically created during the evaluation run of model tunny/Arabic_Qwen2.5_72B_instruct_finetune_0.1. The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_tunny__Arabic_Qwen2.5_72B_instruct_finetune_0.1_v2.tabular100K<n<1M0 likes166 downloads10mo agoHugging Face26Mr-FineTuner /CEFR_Mixed_Dataset_A1_A2_sisa CEFR Dataset for A1 and A2 This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A1 (2000 sentences) and A2 (100 sentences). Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts A1, A2, B1). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A1_A2_sisa.text10K<n<100K0 likes165 downloads1y agoHugging Face27brikdavies /dualmsm-finetune-mixtures dualmsm-finetune-mixtures Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American (Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror identities trained into a base model. Each finetune adds one preference/identity habit on top of the merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.texttext-generation100K<n<1M0 likes159 downloads2mo agoHugging Face28ytu-ce-cosmos /Turkish-LLaVA-Finetune 🔥 TurkishLLaVA Finetuning Dataset This repository contains the dataset used for finetuning the Turkish-LLaVA-v0.1 model. The finetuning process was performed using this dataset, which was concatenated with Turkish-Books to enhance the model's performance. The details of this dataset, along with the finetuning results, will be shared in our upcoming paper (Soon..). Finetuning Configuration During the finetuning phase, both the projection matrix and the language model… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/Turkish-LLaVA-Finetune.text100K<n<1M6 likes158 downloads1y agoHugging Face29thenewsupercell /2D_finetuned_filtered_DF_Audio_Embeddingstext10K<n<100K0 likes154 downloads2y agoHugging Face30toppnoche /receipts-finetune-v1image1K<n<10K0 likes153 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.