datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dsir-pile-10kdsir-pile-13m-filtered-no-github-or-dm_mathematics
DSIR Pile 13M - Filtered Version
This is a filtered version of timaeus/dsir-pile-13m.
Filtering Applied:
Excluded: All rows where metadata.pile_set_name contains 'Github' or 'DM_mathematics'
Kept: All other rows from the original dataset
Dataset Size
Original: ~13M examples
Filtered: 12,782,200 examples (99.9% of original)
Uploaded in: 64 batch files
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/timaeus/dsir-pile-13m-filtered-no-github-or-dm_mathematics.dsir-pile-10m-tokensDSIR-filtered-pile-50M
Dataset Card for DSIR-filtered-pile-50M
Dataset Summary
This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (51.2M examples) in jsonl format.
Data Instances
{"contents": "Hundreds of soul music enthusiasts from the United Kingdom plan to make their way to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-50M.dsir-pile-10k-tokensds_iros_3tasksDSIR-filtered-pile-100M-short
Dataset Card for DSIR-filtered-pile-100M-short
Dataset Summary
This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.dsir-pile-13m-filtered-for-openwebtext2dsir-pile-10kdsir-pile-13m-filtered-for-pile-ccdsir-pile-1m-filtered-no-github-or-dm_mathematics
My_Downsampled_Dataset
This dataset contains 1,000,000 examples from timaeus/dsir-pile-13m-filtered-no-github-or-dm_mathematics, downsampled for efficient processing.
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/my_downsampled_dataset")
dsir-pile-1m-2rows 10m to 11m from the DSIR pile
dsir-pile-13m-filtered-for-gutenberg-pg-19dsir-pile-10mdsir-pile-3mdsir-pile-13mDSIR-traindsir-pile-100krows 10m to 10.1m in the DSIR pile
dsir-pile-13m-filtered-for-nih-exporterdsir-pile-13m-filtered-for-books3dsir-pile-13m-filtered-for-freelawdsir-pile-100krows 10m to 10.1m in the DSIR pile
dsir-pile-13m-filtered-for-pubmed-abstractsdsir-pile-1m-2rows 10m to 11m from the DSIR pile
dsir-pile-100k-filtered-for-OpenWebText2dsir-pile-100k-filtered-for-arxivdsir-pile-5m-2dsir-pile-100k-filtered-for-nih-exporterdsir-pile-100k-filtered-for-enron-emailsdsir-pile-13m-filtered-for-pubmed-central
