CoolFace
Datasetpublic

cambridge-climb/BabyLM

Dataset for the shared baby language modeling task. The goal is to train a language model from scratch on this data which represents roughly the amount of text and speech data a young child observes.

sourceHugging Faceupdated 2y agoView on Hugging Face
3likes2.8kdownloads
43 commits on main
cc972942y ago

Update README.md

rdiehlmartinez
03c086a2y ago

Update README.md

rdiehlmartinez
449545d2y ago

committing all txt files

rdiehlmartinez
66ed2603y ago

Add original tagged and original gold

codebyzeb
a5dd1e13y ago

Add clean tagged and gold tags

codebyzeb
eeb99b33y ago

Update tagging and subconfigs

Zeb
6ee996c3y ago

Add option to load in original files

Zeb
af95aac3y ago

Update tagged text

Zeb
4b568ae3y ago

Fix some errors

Zeb
a245e723y ago

Strip data

Zeb
89e27fb3y ago

Clean data again

Zeb
7813eb43y ago

Update BabyLM.py

codebyzeb
80e95723y ago

Upload 3 files

codebyzeb
739f4843y ago

Fix filepath

codebyzeb
23882cc3y ago

Fix filepath

Zeb
154c3c13y ago

Add tagged gold test

Zeb
39e82ae3y ago

Add tagged test

Zeb
f09b4483y ago

Add tagged gold dev

Zeb
458abbe3y ago

Add tagged dev

Zeb
caf9c943y ago

Add gold tags 10M

Zeb
4ba08cc3y ago

Update loading

Zeb
325b2bc3y ago

Add tagged 10M

Zeb
38f82733y ago

Fix original files

Zeb
5bdb0773y ago

Add rest of clean

Zeb
bcda77c3y ago

Add cleaned 100M

Zeb
1dcc26e3y ago

Add cleaned 10M

Zeb
a966ae13y ago

Add scripts to tag data and improve cleaning

Zeb
0e4956d3y ago

Clean up data a bit more

Zeb
ce43a943y ago

Add filename to each text row

Zeb
4d0612f3y ago

clean-data (#1)

codebyzeb
a03a6bc4y ago

fixing global_idx to update on each yield

rdiehlmartinez
2d574f44y ago

explicitly adding all filenames instead of globbing

rdiehlmartinez
4fef89f4y ago

fixing missing import

rdiehlmartinez
085c4d04y ago

removing unspecified keyword features

rdiehlmartinez
3f5de1d4y ago

v1 of dataset loading script

rdiehlmartinez
917f59b4y ago

uploading last 2 100M files

rdiehlmartinez
f0357af4y ago

uploading more files to 100M

rdiehlmartinez
33d2b6d4y ago

Uploading partial 100M files

rdiehlmartinez
80b961a4y ago

Adding README

rdiehlmartinez
7bfe28c4y ago

uploading test data

rdiehlmartinez
268deec4y ago

adding dev data

rdiehlmartinez
53ca2fa4y ago

Uploading strict small training data

rdiehlmartinez
2fd1a4f4y ago

initial commit

rdiehlmartinez