datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-quantization-fine-tuning-2026
⚡ LLM Fine-Tuning, Quantization & Model Optimization Dataset (2023–2026)
This dataset contains 100 sample audit-verified research papers focusing on Large Language Model (LLM) quantization (GPTQ, AWQ, GGUF), fine-tuning (LoRA, QLoRA, PEFT), pruning, distillation, and speculative decoding.
📊 Features:
384-dimensional PyTorch Embeddings (all-MiniLM-L6-v2) for instant Vector Search
NLP Sentence Extraction: Real extracted core problems & key technical innovations… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/llm-quantization-fine-tuning-2026.Eval_dataset_quantizationRuadaptQwen-Quantization-Dataset
Датасет для квантизации RuadaptQwen2.5-32B-instruct с помощью loss-based методов квантизации
Датасет был собран посредством препроцессинга оригинального Vikhrmodels/Grounded-RAG-RU-v2 датасета,a именно: очисткой от HTML, Markdown, лишних пробелов и т.п. с помощью Qwen2.5-14B-Instruct-GPTQ-Int8.
Также после очистки данные обрезаны так, чтобы количество токенов для каждого предложения было строго 512.Токенизация производилась с помощью токенизатора от целевой модели… See the full description on the dataset page: https://huggingface.co/datasets/pomelk1n/RuadaptQwen-Quantization-Dataset.quantization_experiment_resultsquantization_samples
Dataset Card for Dataset Name
Calibration dataset for quantization with GPTQ.
Dataset Details
128 2048-token samples from the RedPajama-2 dataset.
quantization-energy-datasetmmlu_stem_and_health_quantization_calibration
