datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lightretriever-finetune-data
Training Dataset for LightRetriever
This repo holds all training datasets of the research paper "LightRetriever: A LLM-based Hybrid Retrieval Architecture with 1000x Faster Query Inference
".
More information about the datasets will be updated in the coming weeks. Please stay tuned!
FUSION-Finetune-12M
FUSION-12M Dataset
Please see paper & website for more information:
https://arxiv.org/abs/2504.09925
https://github.com/starriver030515/FUSION
Overview
FUSION-12M is a large-scale, diverse multimodal instruction-tuning dataset used to train FUSION-3B and FUSION-8B models. It builds upon Cambrian-1 by significantly expanding both the quantity and variety of data, particularly in areas such as OCR, mathematical reasoning, and synthetic high-quality Q&A data. The goal is… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Finetune-12M.BioT5_finetune_dataset
References
For more information, please refer to our paper and GitHub repository.
Paper:
BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations
BioT5+: Towards Generalized Biological Understanding with IUPAC Integration and Multi-task Tuning
GitHub: https://github.com/QizhiPei/BioT5
llava-finetuneOpenNiji-Dataset-Aesthetic-Finetune-0-15K
Dataset Card for "OpenNiji-Dataset-Aesthetic-Finetune-0-15K"
More Information needed
CEFR_Mixed_Dataset_1conspiracy-misconception-finetunellmsql-2.0-fine-tune-ready
LLMSQL Benchmark 2.0 (Finetune-Ready)
This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see llmsql-bench/llmsql-2.0.
This repository contains a finetune-ready version of the LLMSQL benchmark: LLMSQL 2.0 on Hugging Face.
The dataset is structured in a messages format suitable for instruction-tuned models, where each example has a messages field. This field is a list of dictionaries with:
"role": "user" — the input question or prompt
"role":… See the full description on the dataset page: https://huggingface.co/datasets/llmsql-bench/llmsql-2.0-fine-tune-ready.swe-bench-fine-tuneWhisper-fine-tune-2Long-Data-Collections-Fine-Tune
Dataset Card for "Long-Data-Collections-Fine-Tune"
Paraquet version of the fine-tune split of togethercomputer/Long-Data-Collections
Statistics (in # of characters): total_len: 6419025428, average_len: 65130.08135393731
receipts-finetune-v2finetune_data
tdro-llm/finetune_data
tDRO: Task-level Distributionally Robust Optimization for Large Language Model-based Dense Retrieval. Guangyuan Ma, Yongliang Ma, Xing Wu, Zhenpeng Su, Ming Zhou and Songlin Hu.
This repo contains all fine-tuning data for Large Language Model-based Dense Retrieval. Please refer to this repo for details to reproduce.
A total of 25 heterogeneous retrieval fine-tuning datasets with Hard Negatives and Deduplication (with test sets) are listed as belows.… See the full description on the dataset page: https://huggingface.co/datasets/tdro-llm/finetune_data.hebrew_lyrics_prompting_finetuneivod-fine-tune注意:使用前請先將 sentence_texts 及 sentence_wavs 底下的 tar 檔案原地解壓縮。
資料集:
ChengTianCai-11-160: 鄭天財,第 11 屆,160 個 IVOD
HuangKuoChang-11-185: 黃國昌,第 11 屆,185 個 IVOD
KeZhiEn-11-55: 柯志恩,第 11 屆,55 個 IVOD
LaiShiBao-10-104: 賴士葆,第 10 屆,104 個 IVOD
LaiShiBao-11-124: 賴士葆,第 11 屆,124 個 IVOD
WangMeiHui-10-71: 王美惠,第 10 屆,71 個 IVOD
WangMeiHui-11-46: 王美惠,第 11 屆,46 個 IVOD
WuSiYao-10-75: 吳思瑤,第 10 屆,75 個 IVOD
WuSiYao-11-103: 吳思瑤,第 11 屆,103 個 IVOD
XieLongJie-11-36: 謝龍介,第 11 屆,36 個 IVOD
common-10-626: 第 10 屆,626 個… See the full description on the dataset page: https://huggingface.co/datasets/openfun/ivod-fine-tune.LlaMA3.1-7B-Instruct-64k-finetune-5G-1Oct2024-TEXTbea-19-fine-tunereceipts-finetune-v3starcoder-finetune
Dataset Card for "starcoder-finetune"
More Information needed
salesforce-xlam-finetune-normalLlaMA3.1-7B-Instruct-15000-finetune-5G-1Oct2024-TEXTfusion-pairwise-evals-finetuned
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares 2 models against gemini-2.5-flash:
Fusion: is the 111B model finetuned on synthetic data generated with Fusion from 5 teachers
BoN: is the 111B model finetuned on synthetic data generated with BoN from 5 teachers
Each model’s outputs are compared in pairs with the respective… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-finetuned.loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes
Synthetic Loracle supervision data generated from FineWeb with OpenRouter.
Run summary
source dataset: HuggingFaceFW/fineweb / sample-10BT / train
sampled docs: 6500
synthetic finetunes: 1284
generated finetunes in this shard: 1000
generator backend: openrouter
generator model: google/gemini-3-flash-preview
max docs per finetune: 40
max token budget per finetune: 10000
questions per finetune: 10
Configs… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracle-fineweb-openrouter-gemini-3-flash-1k-finetunes.re10kDl3dv_mixMode_1Sub2D_TrueAlpha_finetune_0.95MixWeight_unionMask_TrueDetach3Ddetails_tunny__Arabic_Qwen2.5_72B_instruct_finetune_0.1_v2
Dataset Card for Evaluation run of tunny/Arabic_Qwen2.5_72B_instruct_finetune_0.1
Dataset automatically created during the evaluation run of model tunny/Arabic_Qwen2.5_72B_instruct_finetune_0.1.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_tunny__Arabic_Qwen2.5_72B_instruct_finetune_0.1_v2.CEFR_Mixed_Dataset_A1_A2_sisa
CEFR Dataset for A1 and A2
This dataset combines original CEFR-level sentences from training, validation, and test sets with synthetic sentences generated by a fine-tuned LLaMA-3-8B model for CEFR levels A1 (2000 sentences) and A2 (100 sentences). Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the intended level (e.g., A1 accepts A1, A2; A2 accepts A1, A2, B1). Duplicate sentences were… See the full description on the dataset page: https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A1_A2_sisa.dualmsm-finetune-mixtures
dualmsm-finetune-mixtures
Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American
(Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror
identities trained into a base model. Each finetune adds one preference/identity habit on top of the
merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the
Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.Turkish-LLaVA-Finetune
🔥 TurkishLLaVA Finetuning Dataset
This repository contains the dataset used for finetuning the Turkish-LLaVA-v0.1 model. The finetuning process was performed using this dataset, which was concatenated with Turkish-Books to enhance the model's performance. The details of this dataset, along with the finetuning results, will be shared in our upcoming paper (Soon..).
Finetuning Configuration
During the finetuning phase, both the projection matrix and the language model… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/Turkish-LLaVA-Finetune.2D_finetuned_filtered_DF_Audio_Embeddingsreceipts-finetune-v1
