datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-PT
SlimPajama-6B & SlimPajama-30B
Pre-training data sampled from cerebras/SlimPajama-627B, at two scales: ~6B tokens and ~30B tokens. The data is formatted in JSONL and is compatible with LLaMA-Factory training pipelines.
Repository Structure
├── train/
│ ├── 6B/ # ~6B tokens, 7 JSONL files by source
│ │ ├── RedPajamaCommonCrawl-6B.jsonl
│ │ ├── RedPajamaC4-6B.jsonl
│ │ ├── RedPajamaGithub-6B.jsonl
│ │ ├── RedPajamaBook-6B.jsonl
│… See the full description on the dataset page: https://huggingface.co/datasets/Walnutes/SlimPajama-PT.apple-walnut
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/apple-walnut.rejected-apple-walnut
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-apple-walnut.unjudged-apple-walnut
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-apple-walnut.
