CoolFace
Datasetpublic

domofon/ifm-cleaned-pretrain-30B

IFM Cleaned Pretrain — 30B target credits to https://huggingface.co/datasets/IFM/Pretrain-Behaviors Status: complete. Published: 6,155,901 documents; 30,000,015,784 source-annotated tokens. Target: 30,000,000,000 source-annotated tokens, approximately equal across all seven categories. This repository contains text only in Parquet: earlier shards were format-cleaned; subsequent shards contain source text without the cleaner. There are no token-ID arrays or binary token shards.… See the full description on the dataset page: https://huggingface.co/datasets/domofon/ifm-cleaned-pretrain-30B.

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
0likes2kdownloads
settings

This repository belongs to domofon on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameifm-cleaned-pretrain-30B
visibilitypublic
licenceapache-2.0
gatedno
ownerdomofon
Account settings
domofon/ifm-cleaned-pretrain-30B · CoolFace