datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama-3.3-70B-Inst-awq_SafeRLHF
Llama-3.3-70B-Inst-awq Responses for RefAlign Safety Alignment
This dataset contains responses generated for the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data, which introduces the RefAlign alignment algorithm.
Code Repository: https://github.com/mzhaoshuai/RefAlign
This dataset specifically consists of responses generated by the casperhansen/llama-3.3-70b-instruct-awq model, given the prompts from the… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_SafeRLHF.qwen3.5-moe-awq-calibration
Qwen3.5 MoE AWQ Calibration Dataset
Calibration dataset for AWQ (Activation-Aware Weight Quantization) of
Qwen/Qwen3.5-35B-A3B and
Qwen/Qwen3.5-35B-A3B-Base.
Designed for MoE expert routing diversity: Qwen3.5-35B-A3B has 256 experts with 8
active per token, so calibration data needs broad domain coverage to exercise as many
routing paths as possible.
Sampling methodology
Source: PleIAs/common_corpus
(open multi-domain corpus with labeled collections)
Filtering:
Token… See the full description on the dataset page: https://huggingface.co/datasets/Lambent/qwen3.5-moe-awq-calibration.Llama-3.3-70B-Inst-awq_ultrafeedback_1in3
Generated Reference Answers for Language Model Alignment
This dataset contains responses generated for the research presented in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.
The paper introduces RefAlign, a versatile REINFORCE-style alignment algorithm that utilizes language generation evaluation metrics, such as BERTScore, between sampled generations and reference answers as surrogate rewards. This approach… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_ultrafeedback_1in3.clinical-en-smpc-awq-calibration
Clinical-EN SmPC AWQ Calibration Corpus
An English clinical-domain text corpus used as a calibration set for AWQ / AutoAWQ /
GPTQ post-training quantization of English (or multilingual) clinical large language
models. The corpus is dense, domain-specific clinical English (pulmonology + thoracic
oncology), drawn from the English Summary of Product Characteristics (SmPC, Annex I
of the EMA Product Information).
This is the cross-language counterpart to the Polish corpus… See the full description on the dataset page: https://huggingface.co/datasets/mozarcik/clinical-en-smpc-awq-calibration.awq-quant-evals
AWQ Quant Quality Evals
Side-by-side AWQ W4A16 vs BF16 quality measurements for:
Quant
Base
Suite
LostGentoo/Qwen3.5-4B-AWQ
Qwen/Qwen3.5-4B
OpenLLM-lite
LostGentoo/Qwen3-Embedding-8B-AWQ
Qwen/Qwen3-Embedding-8B
MTEB-lite
Hardware: NVIDIA RTX 5060 Ti (sm_120, Blackwell).
Files
File
Contents
quant_quality_evals.json
Full combined report
qwen35_4b_awq_vs_bf16.json
LLM OpenLLM-lite only
qwen3_embedding_8b_awq_vs_bf16.json
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/awq-quant-evals.nanbeige42-awq-gentoo-gauntlet
Nanbeige4.2-3B AWQ — gentoo quality / speed / context / concurrency gauntlet
Separate dataset from LostGentoo/gentoo-small-model-throughput-vllm. Same harness (conc + ctx throughput + lighteval chat_core), single model: AWQ W4A16 of Nanbeige4.2-3B on the Nanbeige vLLM fork.
Setup
Setting
Value
Host
gentoo (3× RTX 5060 Ti 16 GB)
GPU
CUDA device 2
Engine
Nanbeige/vllm @nanbeige42
Model
LostGentoo/Nanbeige4.2-3B-AWQ-W4A16-ASYM
max_model_len
4096… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/nanbeige42-awq-gentoo-gauntlet.Llama-3.3-70B-Inst-awq_ultrafeedbackResponses generated by https://huggingface.co/casperhansen/llama-3.3-70b-instruct-awq given the prompts from https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized.
