datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dream_NLP_FineTunegemini-finetune-datasetfinetune_dataturkish_llm_finetune_dataset_4_topics
Turkish LLM Finetune Dataset - 4 Topics
This dataset is designed to fine-tune the T3 AI Turkish LLM. It was created by Barathan Aslan, Ömer Faruk Çelik, and Batuhan Kalem for the T3 AI Hackathon. The dataset focuses on four distinct topics: Agriculture, Sustainability, Turkish Education Sytem, and Turkish Law System.
Contributors
Barathan Aslan (https://huggingface.co/barathanasln)
Batuhan Kalem(https://huggingface.co/Pancarsuyu)
Ömer Faruk Çelik… See the full description on the dataset page: https://huggingface.co/datasets/barathanasln/turkish_llm_finetune_dataset_4_topics.Financial_News_Translation_Spanish_Finetune
Overview of the Financial News Translation Dataset for OpenAI Model Fine-tuning
Introduction:
This dataset has been curated with the primary objective of fine-tuning varioyus language models to effectively translate financial news content embedded in HTML format. The intention is to enhance the language model's proficiency in accurately and contextually translating financial information for a global audience in a production envionrment.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Benzinga/Financial_News_Translation_Spanish_Finetune.Cars196_dino3_finetune_b
FAISS Index and Results for Cars196_dino3_finetune_b
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (b) model, evaluated on the Cars196 dataset.
Dataset: Cars196
DINO Version: 3
DINO Size: b
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/Cars196_dino3_finetune_b.bert_fine_tune_medical_databl-conversation-dataset-for-llama3-finetune-v2spotify_finetunebenchmark-finetune-lora-v1
Odyn benchmark: LoRA fine-tuning peak VRAM (V1)
Curated benchmark rows for validating GPU memory estimators during LoRA fine-tuning. Each row pairs a published or measured expected peak VRAM with inputs to a math engine (model size, context length, batch, LoRA rank, precision, parallelism) plus optional VRAM breakdown and provenance.
This dataset is not Alpaca-style training JSONL. It is evaluation ground truth for placement / scheduler memory models (Odyn Smart Digester math… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-finetune-lora-v1.data_fine_tuneCUB_dino3_finetune_b
FAISS Index and Results for CUB_dino3_finetune_b
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (b) model, evaluated on the CUB dataset.
Dataset: CUB
DINO Version: 3
DINO Size: b
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path, split, and… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/CUB_dino3_finetune_b.benchmark-finetune-dpo-v1
Odyn benchmark: DPO LoRA fine-tuning peak VRAM (V1)
Curated benchmark rows for validating GPU memory estimators during DPO + LoRA fine-tuning. Each row pairs a published or measured expected peak VRAM with inputs to a math engine (model size, context length, batch, LoRA rank, precision, parallelism) plus optional VRAM breakdown and provenance.
This dataset is not preference-pair training JSONL (UltraFeedback-style). It is evaluation ground truth for placement / scheduler memory… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-finetune-dpo-v1.Finetune_Phi3_model_on_DataBasewikipedia-tr-llm-finetuneCars196_dino3_finetune_s
FAISS Index and Results for Cars196_dino3_finetune_s
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (s) model, evaluated on the Cars196 dataset.
Dataset: Cars196
DINO Version: 3
DINO Size: s
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS index (row number) to the original file path… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/Cars196_dino3_finetune_s.StanfordOnlineProducts_dino3_finetune_b
FAISS Index and Results for StanfordOnlineProducts_dino3_finetune_b
This dataset repository contains the FAISS index, mapping CSV, and evaluation results
for a DINOv3 (b) model, evaluated on the StanfordOnlineProducts dataset.
Dataset: StanfordOnlineProducts
DINO Version: 3
DINO Size: b
Fine-tuned: True
ConvNext (DINOv3): False
Files
faiss_index.bin: The FAISS IndexFlatIP index. Embeddings are L2-normalized.
faiss_index_mapping.csv: A CSV file mapping the FAISS… See the full description on the dataset page: https://huggingface.co/datasets/pawlo2013/StanfordOnlineProducts_dino3_finetune_b.Siamese_Finetune_MSA_SOW_ContractsA dataset prepared for siamese finetuning, to distingush between texts from Legal Contracts text (Majorly SOW, MSA others Legal Algreements, Offer Letters etc)
and text scraped from books, news articles, reviews etc
license: apache-2.0
security_finetunemistral-fine-tune-datasetFS-distilroberta-fine-tunedroberta-leadership-dataset-finetunefine_tune_actual_datafine_tune_datasetmath-fine-tuneInstruct_Finetune_with_Reasoning_WSD
FEWS Dataset for Word Sense Disambiguation (WSD)
This repository contains a formatted and cleaned version of the FEWS dataset, specifically arranged for model fine-tuning for Word Sense Disambiguation (WSD) tasks.
This dataset has further improved for Reasoning showing inrelevant meaning.
Dataset Description
The FEWS dataset has been preprocessed and formatted to be directly usable for training and fine-tuning language models for word sense disambiguation. Each ambiguous… See the full description on the dataset page: https://huggingface.co/datasets/deshanksuman/Instruct_Finetune_with_Reasoning_WSD.llama_2_finetune_smallfine_tune_patient_diagnosesllm_finetunebenchmark-dataset-finetune
Fine-Tuning VRAM Benchmark Dataset
Benchmark dataset for evaluating the accuracy of the Odyn Smart Digester VRAM Math Engine for fine-tuning workloads.
Compares the V1 (initial) and V2 (updated) engine estimates against expected peak VRAM values sourced from published research papers and hardware measurements.
Dataset Details
10 workload rows — all with gradient checkpointing enabled
Methods covered — LoRA (bf16) and QLoRA (NF4)
Models — Llama 2 7B, Llama… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-dataset-finetune.
