CoolFace
Datasetpublic

text-machine-lab/vocab_filtered_dataset_22B

Dataset Card for "vocab_filtered_dataset_22B" Dataset Summary This data is the simplified vocabulary-filtered pretraining data published by "Emergent Abilities in Reduced-Scale Generative Language Models". The vocabulary is derived from the AO-Childes speech corpus (https://github.com/UIUCLearningLanguageLab/AOCHILDES) We filter the train split of SlimPajama dataset (https://huggingface.co/datasets/cerebras/SlimPajama-627B) based on the AO-Childes vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/vocab_filtered_dataset_22B.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes642downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
text-machine-lab/vocab_filtered_dataset_22B · CoolFace