anupbth1/master-dataset-all-V1
Master Dataset All V2 (Part-1 Shards) This repository contains sequentially sharded parts extracted from Google's Natural Questions dataset to optimize training and ingestion loops for LLM fine-tuning. Dataset Structure Format: JSON Lines (.jsonl) Shards Uploaded: train-00000.jsonl to train-00325.jsonl (Part-1) Data Configuration: Out-of-the-box support for datasets loader. Generated and uploaded sequentially via RunPod pipeline.
0154
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face