CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zluvolyote /Dream_NLP_FineTunetabular100K<n<1M0 likes104 downloads4y agoHugging Face02chaitanya4 /gemini-finetune-datasettext1K<n<10K1 likes72 downloads2mo agoHugging Face03nancyH /finetune_datatext1M<n<10M0 likes67 downloads6mo agoHugging Face04barathanasln /turkish_llm_finetune_dataset_4_topics Turkish LLM Finetune Dataset - 4 Topics This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System. Contributors Barathan Aslan (https://huggingface.co/barathanasln) Batuhan Kalem(https://huggingface.co/Pancarsuyu) Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.texttable-question-answering10K<n<100K11 likes60 downloads2y agoHugging Face05pawlo2013 /Cars196_dino3_finetune_b FAISS Index and Results for Cars196_dino3_finetune_b This dataset repository contains the FAISS index, mapping CSV, and evaluation results for a DINOv3 (b) model, evaluated on the Cars196 dataset. Dataset: Cars196 DINO Version: 3 DINO Size: b Fine-tuned: True ConvNext (DINOv3): False Files faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized. faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/Cars196_dino3_finetune_b.tabular100K<n<1M1 likes50 downloads10mo agoHugging Face06Benzinga /Financial_News_Translation_Spanish_Finetune Overview of the Financial News Translation Dataset for OpenAI Model Fine-tuning Introduction: This dataset has been curated with the primary objective of fine-tuning varioyus language models to effectively translate financial news content embedded in HTML format. The intention is to enhance the language model's proficiency in accurately and contextually translating financial information for a global audience in a production envionrment. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Benzinga/Financial_News_Translation_Spanish_Finetune.texttranslationn<1K2 likes45 downloads1mo agoHugging Face07alexkstern /bert_fine_tune_medical_datatext100K<n<1M0 likes32 downloads3y agoHugging Face08SarwarShafee /bl-conversation-dataset-for-llama3-finetune-v2textn<1K0 likes32 downloads2y agoHugging Face09katsuchi /spotify_finetunetextn<1K0 likes23 downloads2y agoHugging Face10pawlo2013 /CUB_dino3_finetune_b FAISS Index and Results for CUB_dino3_finetune_b This dataset repository contains the FAISS index, mapping CSV, and evaluation results for a DINOv3 (b) model, evaluated on the CUB dataset. Dataset: CUB DINO Version: 3 DINO Size: b Fine-tuned: True ConvNext (DINOv3): False Files faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized. faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path, split, and… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/CUB_dino3_finetune_b.tabular100K<n<1M0 likes22 downloads10mo agoHugging Face11odyn-network /benchmark-finetune-dpo-v1 Odyn benchmark: DPO LoRA fine-tuning peak VRAM (V1) Curated benchmark rows for validating GPU memory estimators during DPO + LoRA fine-tuning. Each row pairs a published or measured expected peak VRAM with inputs to a math engine (model size, context length, batch, LoRA rank, precision, parallelism) plus optional VRAM breakdown and provenance. This dataset is not preference-pair training JSONL (UltraFeedback-style). It is evaluation ground truth for placement / scheduler memory… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-finetune-dpo-v1.tabularothern<1K0 likes21 downloads2mo agoHugging Face12rajanonymous12 /Finetune_Phi3_model_on_DataBasetextn<1K1 likes20 downloads2y agoHugging Face13tiendoan /data_fine_tunetext100K<n<1M0 likes20 downloads2y agoHugging Face14uisikdag /wikipedia-tr-llm-finetunetext100K<n<1M0 likes20 downloads2y agoHugging Face15pawlo2013 /StanfordOnlineProducts_dino3_finetune_b FAISS Index and Results for StanfordOnlineProducts_dino3_finetune_b This dataset repository contains the FAISS index, mapping CSV, and evaluation results for a DINOv3 (b) model, evaluated on the StanfordOnlineProducts dataset. Dataset: StanfordOnlineProducts DINO Version: 3 DINO Size: b Fine-tuned: True ConvNext (DINOv3): False Files faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized. faiss_index_mapping.csv: A CSV file mapping the FAISS… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/StanfordOnlineProducts_dino3_finetune_b.tabular100K<n<1M0 likes19 downloads10mo agoHugging Face16odyn-network /benchmark-finetune-lora-v1 Odyn benchmark: LoRA fine-tuning peak VRAM (V1) Curated benchmark rows for validating GPU memory estimators during LoRA fine-tuning. Each row pairs a published or measured expected peak VRAM with inputs to a math engine (model size, context length, batch, LoRA rank, precision, parallelism) plus optional VRAM breakdown and provenance. This dataset is not Alpaca-style training JSONL. It is evaluation ground truth for placement / scheduler memory models (Odyn Smart Digester math… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-finetune-lora-v1.tabularothern<1K0 likes19 downloads3mo agoHugging Face17polestarllp /Siamese_Finetune_MSA_SOW_ContractsA dataset prepared for siamese finetuning, to distingush between texts from Legal Contracts text (Majorly SOW, MSA others Legal Algreements, Offer Letters etc) and text scraped from books, news articles, reviews etc license: apache-2.0 tabular10K<n<100K1 likes18 downloads3y agoHugging Face18sanskrititiwari5 /fine_tune_actual_datatextn<1K0 likes17 downloads2y agoHugging Face19ahmetcangunay /fine_tune_datasettext10K<n<100K0 likes17 downloads2y agoHugging Face20AidanFerrara /security_finetunetext10K<n<100K0 likes17 downloads2y agoHugging Face21jag2023 /mistral-fine-tune-datasettext1K<n<10K0 likes17 downloads2y agoHugging Face22FinScience /FS-distilroberta-fine-tunedtext1K<n<10K1 likes16 downloads4y agoHugging Face23orYx-models /roberta-leadership-dataset-finetunetextn<1K0 likes16 downloads2y agoHugging Face24pawlo2013 /Cars196_dino3_finetune_s FAISS Index and Results for Cars196_dino3_finetune_s This dataset repository contains the FAISS index, mapping CSV, and evaluation results for a DINOv3 (s) model, evaluated on the Cars196 dataset. Dataset: Cars196 DINO Version: 3 DINO Size: s Fine-tuned: True ConvNext (DINOv3): False Files faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized. faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/Cars196_dino3_finetune_s.tabular100K<n<1M0 likes16 downloads10mo agoHugging Face25deshanksuman /Instruct_Finetune_with_Reasoning_WSD FEWS Dataset for Word Sense Disambiguation (WSD) This repository contains a formatted and cleaned version of the FEWS dataset, specifically arranged for model fine-tuning for Word Sense Disambiguation (WSD) tasks. This dataset has further improved for Reasoning showing inrelevant meaning. Dataset Description The FEWS dataset has been preprocessed and formatted to be directly usable for training and fine-tuning language models for word sense disambiguation. Each ambiguous… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Instruct_Finetune_with_Reasoning_WSD.text100K<n<1M0 likes15 downloads1y agoHugging Face26takojunior /llama_2_finetune_smalltext10K<n<100K0 likes14 downloads3y agoHugging Face27alexkstern /fine_tune_patient_diagnosestext10K<n<100K1 likes14 downloads3y agoHugging Face28yasamanne /math-fine-tunetext100K<n<1M2 likes14 downloads2y agoHugging Face29pawlo2013 /CUB_dino3_finetune_s FAISS Index and Results for CUB_dino3_finetune_s This dataset repository contains the FAISS index, mapping CSV, and evaluation results for a DINOv3 (s) model, evaluated on the CUB dataset. Dataset: CUB DINO Version: 3 DINO Size: s Fine-tuned: True ConvNext (DINOv3): False Files faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized. faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path, split, and… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/CUB_dino3_finetune_s.tabular100K<n<1M0 likes14 downloads10mo agoHugging Face30odyn-network /benchmark-dataset-finetune Fine-Tuning VRAM Benchmark Dataset Benchmark dataset for evaluating the accuracy of the Odyn Smart Digester VRAM Math Engine for fine-tuning workloads. Compares the V1 (initial) and V2 (updated) engine estimates against expected peak VRAM values sourced from published research papers and hardware measurements. Dataset Details 10 workload rows — all with gradient checkpointing enabled Methods covered — LoRA (bf16) and QLoRA (NF4) Models — Llama 2 7B, Llama… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-dataset-finetune.tabularothern<1K0 likes14 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.