CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chinese-babylm-org /zhoblimptext10K<n<100K0 likes551 downloads5mo agoHugging Face02chinese-babylm-org /hanzi-pinyintext1K<n<10K0 likes88 downloads4mo agoHugging Face03chinese-babylm-org /hanzi-structuretext1K<n<10K0 likes85 downloads4mo agoHugging Face04augustinian-babylm /vpswap-checkpoint-scores VP-Swap checkpoint scores Per-item correctness on the VP-Swap benchmark for nine models at twenty points in training. This is the raw material behind Figures 4-6 of Augustinian BabyLM (paper, code). Layout <model>/<revision>.jsonl, one line per benchmark item: {"property": "color", "line": 6, "which": 1, "pll_orig": -21.4213, "pll_swap": -23.9077, "correct": true} property and line identify the source line in eval/vpswap_bb24/vp_swap_<property>_pairs.txt; which… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/vpswap-checkpoint-scores.tabular1M<n<10M0 likes64 downloads16d agoHugging Face05ltg /babylm-2024-baby-cosmo-fine-100m@misc{charpentier2024gptbertboth, title={GPT or BERT: why not both?}, author={Lucas Georges Gabriel Charpentier and David Samuel}, year={2024}, eprint={2410.24159}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2410.24159}, } text100K<n<1M2 likes53 downloads2y agoHugging Face06ltg /babylm-2024-baby-cosmo-fine-10m@misc{charpentier2024gptbertboth, title={GPT or BERT: why not both?}, author={Lucas Georges Gabriel Charpentier and David Samuel}, year={2024}, eprint={2410.24159}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2410.24159}, } text10K<n<100K0 likes35 downloads2y agoHugging Face07flakoash /babylm-curriculum-run07text1M<n<10M0 likes18 downloads3mo agoHugging Face08biancaganescu /babylm-dino-embeddingstext1M<n<10M0 likes14 downloads2y agoHugging Face09suchirsalhan /BabyLM-Pretokenisedtext1M<n<10M0 likes11 downloads2y agoHugging Face10chinese-babylm /zhoblimptext1K<n<10K0 likes10 downloads6mo agoHugging Face11sixf0ur /babylm_eng_distilled_1024This data is based on the babylm-eng dataset. It was processed and distilled through DeepSeek. Each entry was truncated to 3500 characters and then rewritten in simpler English. The output tokens for each entry are limited to 1024 tokens. This dataset was created with a strong focus on noise removal and simple language, making it suitable for tiny language models. The following prompt details the instructions for the distillation process: Instruction You are an expert data curator… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/babylm_eng_distilled_1024.text100K<n<1M0 likes9 downloads8mo agoHugging Face12flakoash /babylm-curriculum-sliding-window-4bandstext1M<n<10M0 likes5 downloads3mo agoHugging Face13chinese-babylm-org /NLI-hiddentext1K<n<10K0 likes4 downloads4mo agoHugging Face14flakoash /babylm-curriculum-tiered-4bandstext100K<n<1M0 likes4 downloads3mo agoHugging Face15flakoash /babylm-curriculum-run05textn<1K0 likes3 downloads3mo agoHugging Face16flakoash /babylm-curriculum-run06text100K<n<1M0 likes3 downloads3mo agoHugging Face17flakoash /babylm-curriculum-tiered-2bandstext100K<n<1M0 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.