kuluruvineeth/manas_dataset
Manas Dataset The complete English training bundle for Manas, a lightweight language model trained entirely from scratch. Every file is rebuildable from raw sources with python -m datapipe.build all. Files file stage pretrain_t2t.jsonl pretraining corpus, ~2.2B tokens pretrain_t2t_mini.jsonl pretraining corpus, quick-start tier sft_t2t.jsonl supervised fine-tuning conversations (tool-calling and reasoning mixed in) sft_t2t_mini.jsonl supervised… See the full description on the dataset page: https://huggingface.co/datasets/kuluruvineeth/manas_dataset.
Upload README.md with huggingface_hub
Upload lora_exam.jsonl with huggingface_hub
Upload lora_medical.jsonl with huggingface_hub
Upload lora_identity.jsonl with huggingface_hub
Upload agent_rl_math.jsonl with huggingface_hub
Upload agent_rl.jsonl with huggingface_hub
Upload rlaif.jsonl with huggingface_hub
Upload dpo.jsonl with huggingface_hub
Upload sft_t2t_mini.jsonl with huggingface_hub
Upload sft_t2t.jsonl with huggingface_hub
Upload pretrain_t2t_mini.jsonl with huggingface_hub
Upload pretrain_t2t.jsonl with huggingface_hub
Upload README.md with huggingface_hub
initial commit
