datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
instruction-responses-500MBThis is just the first 500MB (~100M tokens) of jukofyork/instruction-responses.
instruction-responses
Merged Instruction Responses Dataset
This dataset was created by combining:
causal-lm/instructions
rombodawg/Everything_Instruct
and then extracting only the responses between 75 and 2000 characters into the "text" field.
The combined dataset was then de-duplicated and re-shuffled.
khasi-instruction-response-v2
Khasi Instruction Response v2
The Khasi Instruction Response v2 dataset is a high-quality, curated collection of 77,810 instruction-response pairs designed to fine-tune Large Language Models (LLMs) for the Khasi language. This is an improved, expanded version of my previous v1 release, offering significantly higher data integrity and broader linguistic coverage.
It combines extensive cultural, literary, and translation-based Khasi data with high-reasoning capabilities from… See the full description on the dataset page: https://huggingface.co/datasets/toiar/khasi-instruction-response-v2.
