domofon/ifm-cleaned-pretrain-30B
IFM Cleaned Pretrain — 30B target credits to https://huggingface.co/datasets/IFM/Pretrain-Behaviors Status: complete. Published: 6,155,901 documents; 30,000,015,784 source-annotated tokens. Target: 30,000,000,000 source-annotated tokens, approximately equal across all seven categories. This repository contains text only in Parquet: earlier shards were format-cleaned; subsequent shards contain source text without the cleaner. There are no token-ID arrays or binary token shards.… See the full description on the dataset page: https://huggingface.co/datasets/domofon/ifm-cleaned-pretrain-30B.
This repository belongs to domofon on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
