CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tdro-llm /finetune_data tdro-llm/finetune_data tDRO: Task-level Distributionally Robust Optimization for Large Language Model-based Dense Retrieval. Guangyuan Ma, Yongliang Ma, Xing Wu, Zhenpeng Su, Ming Zhou and Songlin Hu. This repo contains all fine-tuning data for Large Language Model-based Dense Retrieval. Please refer to this repo for details to reproduce. A total of 25 heterogeneous retrieval fine-tuning datasets with Hard Negatives and Deduplication (with test sets) are listed as belows.… See the full description on the dataset page: https://huggingface.co/datasets/tdro-llm/finetune_data.textn<1K0 likes244 downloads1y agoHugging Face02CohereLabs /fusion-pairwise-evals-finetuned Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N Content This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares 2 models against gemini-2.5-flash: Fusion: is the 111B model finetuned on synthetic data generated with Fusion from 5 teachers BoN: is the 111B model finetuned on synthetic data generated with BoN from 5 teachers Each model’s outputs are compared in pairs with the respective… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-finetuned.texttext-generation1K<n<10K1 likes185 downloads1y agoHugging Face03thenewsupercell /2D_finetuned_filtered_DF_Audio_Embeddingstext10K<n<100K0 likes154 downloads2y agoHugging Face04BERRAMOU /math-classification-finetuned-resultstextn<1K0 likes134 downloads4d agoHugging Face05trannguyenquynhnhu /SEA-Instruct-2602-fine-tunedtext100K<n<1M0 likes118 downloads2mo agoHugging Face06nancyH /finetune_datatext1M<n<10M0 likes70 downloads6mo agoHugging Face07sjsurbhi /fine-tuned-deepseek2-16b-distillation-datasettext10K<n<100K0 likes66 downloads1y agoHugging Face08Amba /mt5-small-finetuned-amazon-en-es_tokenized_datasetstext1K<n<10K0 likes64 downloads5y agoHugging Face09saberai /Zrov2_FineTunedtext10K<n<100K0 likes55 downloads3y agoHugging Face10Amba /mt5-small-finetuned-amazon-en-es_books_datasettext1K<n<10K0 likes49 downloads5y agoHugging Face11CabraVC /vector_dataset_roberta-fine-tunedtext1K<n<10K0 likes45 downloads3y agoHugging Face12NickH01 /finetune-dataset-testplantext1K<n<10K0 likes45 downloads1y agoHugging Face13kaihuac /cognvs_ckpt_test_time_finetunedtextn<1K0 likes42 downloads1y agoHugging Face14adamjweintraut /bart-finetuned-lyrlen-256-tokens_2024-03-22_runtabularn<1K0 likes40 downloads3y agoHugging Face15adamjweintraut /bart-finetuned-lyrlen-512-tokens_2024-03-24_runtabularn<1K0 likes39 downloads3y agoHugging Face16NickH0105 /finetune-dataset-testplantext1K<n<10K0 likes38 downloads9mo agoHugging Face17bassemessam /mT5_multilingual_XLSum-finetuned-wiki-linguatext10K<n<100K0 likes36 downloads2y agoHugging Face18adamjweintraut /bart-finetuned-lyrlen-128-tokens_2024-03-22_runtabularn<1K0 likes35 downloads3y agoHugging Face19richmondsin /finetuned_mmlu_ml_output_layer_20_results Dataset Card for Evaluation run of richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0 Dataset automatically created during the evaluation run of model richmondsin/finetuned-gemma-2-2b-output-layer-20-4k-0 The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/finetuned_mmlu_ml_output_layer_20_results.tabular10K<n<100K0 likes33 downloads2y agoHugging Face20happieebitees /FineTuned_indian_foodtext1K<n<10K2 likes31 downloads2y agoHugging Face21vaibhavmeena /finetune-data-for-vision-based-llms5imagen<1K0 likes30 downloads2y agoHugging Face22and89 /fine_tuned_llama2 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/and89/fine_tuned_llama2.textn<1K0 likes30 downloads2y agoHugging Face23paacamo /nvidia-faq-llm-gemma2-fine-tunedtext1K<n<10K0 likes30 downloads1y agoHugging Face24JJYDXFS /RAMP_finetuned_data_Cora Dataset Details This dataset contains finetuning data constructed from the Cora citation network for downstream text-rich graph tasks. It is used for finetuning RAMP (Raw-text Anchored Message Passing), which recasts the LLM as a graph-native aggregation operator on text-rich graphs. The dataset includes the following files: finetuned_cora_v1.json — Training set finetuned_cora_val_v1.json — Validation set eval_cora_v1.json — Test set This is a release from our paper LLM as Graph… See the full description on the dataset page: https://huggingface.co/datasets/JJYDXFS/RAMP_finetuned_data_Cora.textn<1K0 likes30 downloads6mo agoHugging Face25mehdie /fine_tune_dataset_testtext1K<n<10K0 likes29 downloads2y agoHugging Face26TeamClaude /SFT-Fine-Tuned Qwen3 4B LIMA-Style SFT Variants Two compact supervised fine-tuning datasets prepared for a Qwen3 4B base model. The goal is quality over volume: a human-written instruction-following anchor plus verified math/code/reasoning examples. Variants Config Max formatted tokens Train rows Validation rows Total rows Intended use qwen3_lima_sft_mix_1024 1024 32,697 500 33,197 Conservative first-pass SFT matching the pretraining sequence length. qwen3_lima_sft_mix_2048… See the full description on the dataset page: https://huggingface.co/datasets/TeamClaude/SFT-Fine-Tuned.text10K<n<100K0 likes29 downloads4mo agoHugging Face27GENIAC-Team-Ozaki /tuninig-dataset_pref_20pct_v2_full-sft-finetuned-stage4-iter86000-v2text10K<n<100K0 likes28 downloads2y agoHugging Face28mehdie /fine_tune_dataset_STER_idstext1K<n<10K0 likes28 downloads2y agoHugging Face29paacamo /nvidia-faq-EleutherAI-pythia-1b-fine-tunedtext1K<n<10K0 likes28 downloads1y agoHugging Face30GenAIGirl /imdb_Sentimental_Finetuned_20Ktext10K<n<100K0 likes27 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.