CoolFace
Datasetpublic

cambridge-climb/BabyLM

Dataset for the shared baby language modeling task. The goal is to train a language model from scratch on this data which represents roughly the amount of text and speech data a young child observes.

sourceHugging Faceupdated 2y agoView on Hugging Face
3likes2.8kdownloads
Dataset Card

BabyLM Dataset

This download includes LM Pretraining data for the 2023 CoNLL/CMCL shared task, The BabyLM Challenge. The (unzipped) data is not large, only ~700MB.

Note that there is also a multi-lingual version of this dataset, that is availabled under the multi-lingual branch of the dataset repository.

Contents of this download

  • —10M: 10M-word training set for the strict-small track.
  • —dev: Development set for both tracks (10M words)
  • —test: Test set for both tracks (10M words)

Each directory above contains a single .txt file from each of the 10 domains listed below.

Composition of the data

All datasets are sampled from a mixture of 10 data domains, shown below, along with their respective weights in the distributed dataset.

SourceWeightDomainCitationWebsiteLicense
OpenSubtitles30%Dialogue, ScriptedLison & Tiedermann (2016)linkOpen source
Simple English Wikipedia15%Nonfiction--linklink
BNC10%DialogueBNC Consortium (2007)linklink <sup>1</sup>
Project Gutenberg10%Fiction, NonfictionGerlach & Font-Clos (2020)linklink
QED10%Dialogue, EducationAbdelali et al. (2014)linklink
Wikipedia10%Nonfiction--linklink
Children's Book Test6%Fiction, Child-DirectedHill et al. (2016)linkPublic domain
CHILDES4%Dialogue, Child-DirectedMacWhinney (2000)link
Children's Stories4%Fiction, Child-Directed--linkPublic domain
Switchboard1%DialogueGodfrey et al. (1992), Stolcke et al., (2000)linklink

<sup>1</sup> Our distribution of part of the BNC Texts is permitted under the fair dealings provision of copyright law (see term (2g) in the BNC license).

Data preprocessing

Data was minimally preprocessed to conform to a plain text format. We did not tokenize the data. Documents are not necessarily complete are newline separated.

For documentation of the preprocessing pipeline, consult the following repo: https://github.com/babylm/babylmdatapreprocessing

References

Abdelali, A., Guzman, F., Sajjad, H., & Vogel, S. (2014). The AMARA Corpus: Building parallel language resources for the educational domain. In Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014). 1856-1862.

BNC Consortium. (2007). The British National Corpus, XML Edition. Oxford Text Archive, http://hdl.handle.net/20.500.12024/2554.

Gerlach, M., & Font-Clos, F. (2020). A standardized Project Gutenberg corpus for statistical analysis of natural language and quantitative linguistics. Entropy, 22(1), 126.

Godfrey, J. J., Holliman, E. C., & McDaniel, J. (1992). SWITCHBOARD: Telephone speech corpus for research and development. In Acoustics, Speech, and Signal Processing, IEEE International Conference on (Vol. 1, pp. 517-520). IEEE Computer Society.

Hill, F., Bordes, A., Chopra, S., Weston, J. (2016). The Goldilocks principle: Reading children’s books with explicit memory representations. In Proceedings of the 4th International Conference on Learning Representations (ICLR 2016).

Lison, P. & Tiedemann, J. (2016). OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles. In Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC 2016).

MacWhinney, B. (2000). The CHILDES Project: Tools for analyzing talk. Third Edition. Mahwah, NJ: Lawrence Erlbaum Associates.

Stolcke, A., Ries, K., Coccaro, N., Shriberg, E., Bates, R., Jurafsky, D., Taylor, P., Martin, R., Van Ess-Dykema, C., & Meteer, M. (2000). Dialogue act modeling for automatic tagging and recognition of conversational speech. Computational linguistics, 26(3), 339-373.

Tiedemann, J. (2012). Parallel Data, Tools and Interfaces in OPUS. In Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC 2012).