datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HQ-QUESTION-PAIRS
❤️ Support Our Mission
🙏 Drop a Heart ❤ For Our Hard Work!
If you believe in the vision of Sovereign Indian AI, please show your support by dropping a heart below. Your encouragement fuels our journey!
❤️ Like This Dataset
➕ Follow SKT AI LABS
Kindly follow us for more updates and contribute to our open-source journey!
HQ PAIRS DATASET DIVISION
SKT HIGH-QUALITY QUESTION PAIRS… See the full description on the dataset page: https://huggingface.co/datasets/sKT-Ai-Labs/HQ-QUESTION-PAIRS.hqq_plus_plus_mix_40k
Dataset Card for Dataset Name
Replication of HQQ++ dataset mixture mentioned in here with Llama 3 Instruct chat template.
Uses
Quantization Aware Training.
Source Data
timdettmers/openassistant-guanaco
microsoft/orca-math-word-problems-200k
meta-math/MetaMathQA
HuggingFaceH4/ultrafeedback_binarized
Data Collection and Processing
train_ds1 = load_dataset("timdettmers/openassistant-guanaco")['train']
train_ds2 =… See the full description on the dataset page: https://huggingface.co/datasets/answerdotai/hqq_plus_plus_mix_40k.raw_benchmark_results_hqqtest-wfyllama3-custom-testtest-llama3
