CoolFace
Datasetpublic

ssuresh/nemo-stage1-50M-samples

NeMo Stage1 Pretraining Dataset - 50M Samples This dataset contains 50 million text samples for NeMo model pretraining (Stage 1). The dataset is organized in chunks for efficient loading and processing. Dataset Details Total Samples: ~50,000,000 Format: JSONL (JSON Lines) Structure: Each sample contains {"id": number, "text": "content"} Chunks: 47 files (chunk_000.jsonl to chunk_046.jsonl) Samples per chunk: ~1,000,000 Language: English Task: Text generation… See the full description on the dataset page: https://huggingface.co/datasets/ssuresh/nemo-stage1-50M-samples.

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
0likes140downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ssuresh/nemo-stage1-50M-samples · CoolFace