CoolFace
Datasetpublic

text-machine-lab/constrained_language

Dataset Card for constrained_language (pre-training data for simplified English) Dataset Summary This dataset is one of the two datasets published by "Honey, I Shrunk the Language: Language Model Behavior at Reduced Scale" (https://arxiv.org/abs/2305.17266). The dataset available at this link is the pre-training data constrained by vocabulary. The other published data i.e. the pre-training data that is not constrained by vocabulary is available at… See the full description on the dataset page: https://huggingface.co/datasets/text-machine-lab/constrained_language.

sourceHugging Faceupdated 3y agoView on Hugging Face
2likes148downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
text-machine-lab/constrained_language · CoolFace