datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLMFineTuningBench
Dataset Card for LLMFineTuningBench
A dataset of over 30,000 LLM fine-tuning experiments, capturing detailed performance metrics from jobs run on high-performance computing (HPC) clusters. It spans a wide range of models, fine-tuning methods, and hardware configurations, and is intended to support research on predictive resource allocation, performance optimization, and cost estimation for LLM fine-tuning workloads.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLMFineTuningBench.medical-llm-finetuning-alignment-original-datasetoc-llm-finetuning-dataset
Dataset médical bilingue, triage CHSA
Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms,
AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé,
instruction/réponse) et dpo (alignement par préférences, chosen/rejected).
Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques.
Schéma
Champs communs à tous les exemples :
Champ
Type
Description
id
string
Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.LLM_Fine-Tuning_Performance
LLM Fine-Tuning Performance Benchmark Dataset
Dataset Summary
This dataset contains performance benchmarks for Large Language Model (LLM) fine-tuning across various hardware and software configurations. It includes throughput measurements (tokens per second) for 959 valid configurations, collected over 1000 GPU hours on a Kubernetes cluster. The dataset is designed for research on predictive performance modeling, specifically for evaluating methods that handle Categorical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/LLM_Fine-Tuning_Performance.details_csujeong__Gemma-7B-Finetuning-JCS-Ko-Ins
Dataset Card for Evaluation run of csujeong/Gemma-7B-Finetuning-JCS-Ko-Ins
Dataset automatically created during the evaluation run of model csujeong/Gemma-7B-Finetuning-JCS-Ko-Ins on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_csujeong__Gemma-7B-Finetuning-JCS-Ko-Ins.medical-llm-finetuning-alignment-processed-datasetDSA_FINETUNING_LLMllm-finetuning-fr
LLM Fine-Tuning & Quantization - Dataset Francais
Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production.
Description
Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.bitcoin-llm-finetuning-dataset
Bitcoin Price Prediction Instruction-Tuning Dataset
Dataset Description
This is a comprehensive, instruction-based dataset formatted specifically for fine-tuning Large Language Models (LLMs) to perform financial forecasting. The dataset is structured as a collection of prompt-response pairs, where each record challenges the model to predict the next 10 days of Bitcoin's price based on a rich, multimodal snapshot of daily data.
The dataset covers the period from early 2018… See the full description on the dataset page: https://huggingface.co/datasets/tahamajs/bitcoin-llm-finetuning-dataset.details_csujeong__Mistral-7B-Finetuning-Insurance-16R
Dataset Card for Evaluation run of csujeong/Mistral-7B-Finetuning-Insurance-16R
Dataset automatically created during the evaluation run of model csujeong/Mistral-7B-Finetuning-Insurance-16R on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_csujeong__Mistral-7B-Finetuning-Insurance-16R.llm-finetuning-wr25flock-demo-llm-finetuning-sectionsbitcoin-llm-finetuning-dataset_long_allbitcoin-llm-finetuning-dataset_new_allllm-finetuning-en
LLM Fine-Tuning & Quantization - English Dataset
Comprehensive bilingual dataset on LLM fine-tuning (LoRA, QLoRA, DPO, RLHF), model quantization (GPTQ, GGUF, AWQ), open source models, and production deployment.
Description
This dataset covers the entire open source LLM value chain, from fine-tuning to production deployment. It is designed as a reference for developers, ML engineers, and technical teams looking to master open source LLMs.
Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-en.bitcoin-llm-finetuning-dataset_long_50000bitcoin-llm-finetuning-dataset_allfinetuning_LLM_with_laminibitcoin-llm-finetuning-dataset_new_with_custom_textBanking_Dataset_for_LLM_Finetuningbitcoin-llm-finetuning-dataset_all_50000bitcoin-llm-finetuning-dataset_long_100000bitcoin-llm-finetuning-dataset_new_50000bitcoin-llm-finetuning-dataset_enriched_with_gemini_v1bitcoin-llm-finetuning-dataset
Bitcoin Price Prediction Instruction-Tuning Dataset
Dataset Description
This is a comprehensive, instruction-based dataset formatted specifically for fine-tuning Large Language Models (LLMs) to perform financial forecasting. The dataset is structured as a collection of prompt-response pairs, where each record challenges the model to predict the next 10 days of Bitcoin's price based on a rich, multimodal snapshot of daily data.
The dataset covers the period from… See the full description on the dataset page: https://huggingface.co/datasets/acartaylan66/bitcoin-llm-finetuning-dataset.llm-lion-finetuningbitcoin-llm-finetuning-dataset_newllm-tool-finetuningLLM-fine-tuning-amazon-2023-revamped-priced-databitcoin-llm-finetuning-dataset_new_20000
