datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finetune_data
tdro-llm/finetune_data
tDRO: Task-level Distributionally Robust Optimization for Large Language Model-based Dense Retrieval. Guangyuan Ma, Yongliang Ma, Xing Wu, Zhenpeng Su, Ming Zhou and Songlin Hu.
This repo contains all fine-tuning data for Large Language Model-based Dense Retrieval. Please refer to this repo for details to reproduce.
A total of 25 heterogeneous retrieval fine-tuning datasets with Hard Negatives and Deduplication (with test sets) are listed as belows.… See the full description on the dataset page: https://huggingface.co/datasets/tdro-llm/finetune_data.verifications_llama3.1_8b_finetuned_verifierverifications_aime24_llama_3.1_8b_finetuned_verifier_4Ktokensaime24_verifications_qwen_2.5_7B_solutions_llama_3.1_8b_finetuned_verifier_4K
