datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RedPajama-Data-1T-Sample-Backupredpajama-data-v2_urls
Dataset Card for redpajama-data-v2_urls
This dataset provides the URLs and top-level domains associated with training records in togethercomputer/RedPajama-Data-V2. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/redpajama-data-v2_urls.RedPajama-Data-1T-Samplered_pajama_es_hq
RedPajama's High Quality Spanish subset
What is this?
The following is a high-quality dataset distilled from the Spanish subsection of RedPajama-Data-v2, created using the methodology proposed in FineWEB-Edu.
Usage
from datasets import load_dataset
ds = load_dataset("latam-gpt/red_pajama_es_hq")
Filtering by quality score
Documents in this corpus are scored on academic quality from 2.5 to 5, with higher scores indicating better quality. The… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/red_pajama_es_hq.filtered_redpajama_multilingualredpajama_v2_sample_10M
Dataset Card for "redpajama_v2_sample_10M"
More Information needed
redpajama_v2_sample_100M
Dataset Card for "redpajama_v2_sample_100M"
More Information needed
filtered_redpajama_enRedPajama-pro
📚 RedPajama-pro
ArXiv | Models | Code
RedPajama-pro is refined from RedPajama-Data-V2 using the ProX refining framework.
It contains about 30B high quality tokens, ready for general language model pre-training.
License
RedPajama-pro is based on RedPajama-Data-V2, which is made available under an apache-2.0 license; users should also abide by the CommonCrawl ToU: https://commoncrawl.org/terms-of-use/. We do not alter the license of any of the underlying data.… See the full description on the dataset page: https://huggingface.co/datasets/gair-prox/RedPajama-pro.redpajama_100gb_processedRedPajama-Data-V2-1B
RedPajama-Data-V2 1B
Dataset Description
This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data.
Motivation
RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.RedPajama-Tiny
Dataset Card for Dataset Name
Dataset Summary
This is a tiny version of the RedPajama dataset.
It contains 64 samples from each of the 7 sources.
This dataset is intended for developing and testing data/training pipeline for loading the full RedPajama dataset or any general HuggingFace dataset.
It is very fast to download and easy to examine. You should not use it for training a full model, but you can use it for overfitting test or any other sanity checks.… See the full description on the dataset page: https://huggingface.co/datasets/ivanzhouyq/RedPajama-Tiny.redpajama_v2_sample_1M
Dataset Card for "redpajama_v2_sample_1M"
More Information needed
RedPajama-Data-V2-sample-100B-filtered-for-regression-domains-with-domainsRedPajama-Data-1T-Sample-Backup
RedPajama Data 1T Sample Backup
This dataset is a backup mirror of togethercomputer/RedPajama-Data-1T-Sample.
It is provided for easier access when the original dataset is unavailable or difficult to download.
Usage
Original:
from datasets import load_dataset
ds = load_dataset(
"togethercomputer/RedPajama-Data-1T-Sample",
split="train",
trust_remote_code=True,
)
Backup:
from datasets import load_dataset
ds = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ll922/RedPajama-Data-1T-Sample-Backup.RedPajama-Data-V2-sample-100B-filtered-shuffled-tokenized-with-token-countsred_pajama_es_hq_35Just a filtered version of https://huggingface.co/datasets/latam-gpt/red_pajama_es_hq with scores >=3.5
redpajama-data-1t_urls
Dataset Card for redpajama-data-1t_urls
This dataset provides the URLs and top-level domains associated with training records in togethercomputer/RedPajama-Data-1T. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/redpajama-data-1t_urls.RedPajama-Data-V2-sample-100B-filtered-shuffled-tokenized-with-token-counts-2500000enRedPajama-Data-100-Sample-For-Testredpajama_arxiv_hfcorpus-shard-redpajama-19corpus-shard-redpajama-20corpus-shard-redpajama-17corpus-shard-redpajama-22corpus-shard-redpajama-18RedPajama-Data-1T-1024SampleRedPajama is a clean-room, fully open-source implementation of the LLaMa dataset. This is a 1B-token sample of the full dataset.redpajama-sample_from_valid_all
Dataset Card for "redpajama-sample_from_valid_all"
More Information needed
RedPajama-Data-1T-arxiv-filtered
Dataset Card for "RedPajama-Data-1T-arxiv-filtered"
More Information needed
corpus-shard-redpajama-21
