datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLMFineTuningBench
Dataset Card for LLMFineTuningBench
A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.llm-finetuning-fr
LLM Fine-Tuning & Quantization - Dataset Francais
Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production.
Description
Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.llm-quantization-fine-tuning-2026
⚡ LLM Fine-Tuning, Quantization & Model Optimization Dataset (2023–2026)
This dataset contains 100 sample audit-verified research papers focusing on Large Language Model (LLM) quantization (GPTQ, AWQ, GGUF), fine-tuning (LoRA, QLoRA, PEFT), pruning, distillation, and speculative decoding.
📊 Features:
384-dimensional PyTorch Embeddings (all-MiniLM-L6-v2) for instant Vector Search
NLP Sentence Extraction: Real extracted core problems & key technical innovations… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/llm-quantization-fine-tuning-2026.real-estate-data-sample-for-llm-fine-tuningllm-finetuning-en
LLM Fine-Tuning & Quantization - English Dataset
Comprehensive bilingual dataset on LLM fine-tuning (LoRA, QLoRA, DPO, RLHF), model quantization (GPTQ, GGUF, AWQ), open source models, and production deployment.
Description
This dataset covers the entire open source LLM value chain, from fine-tuning to production deployment. It is designed as a reference for developers, ML engineers, and technical teams looking to master open source LLMs.
Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-en.fine-tuning-llm
Dataset Card for fine-tuning-llm
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/KarinaH/fine-tuning-llm/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/KarinaH/fine-tuning-llm.
