CoolFace
Datasetpublic

krisbailey/RedPajama-Data-V2-100M

RedPajama-Data-V2-100M Dataset Description This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2. Motivation 100M tokens is a standard size for: CI/CD Pipelines: Fast enough to download and train for unit tests. Debugging: Verifying training loops without waiting for hours. Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-100M.

sourceHugging Faceodc-byupdated 8mo agoView on Hugging Face
0likes58downloads
Dataset Card

RedPajama-Data-V2-100M

Dataset Description

This is a 100.0 Million token subset of krisbailey/RedPajama-Data-V2-1B, which is a subset of togethercomputer/RedPajama-Data-V2.

Motivation

100M tokens is a standard size for:

  • —CI/CD Pipelines: Fast enough to download and train for unit tests.
  • —Debugging: Verifying training loops without waiting for hours.
  • —Scaling Laws: The first step in a logarithmic scaling series (100M -> 1B -> 10B).

Dataset Details

  • —Total Tokens: 99,999,721
  • —Source: krisbailey/RedPajama-Data-V2-1B
  • —Structure: First ~10% of the randomized 1B dataset.
  • —Format: Parquet (Snappy compression) - Single File
  • —Producer: Kris Bailey (kris@krisbailey.com)

Usage

python
from datasets import load_dataset

ds = load_dataset("krisbailey/RedPajama-Data-V2-100M", split="train")
print(ds[0])

Citation

bibtex
@article{together2023redpajama,
  title={RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset},
  author={Together Computer},
  journal={https://github.com/togethercomputer/RedPajama-Data},
  year={2023}
}