CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01placeholderlabs /pretrain-ultra-fineweb-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 73,328,911,347 (73.3B) Trainable tokens 73,328,911,347 (73.3B) Documents 92,923,076 Shards 578 UTF-8 bytes 364,562,557,837 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix.tabular10M<n<100M1 likes881 downloads13d agoHugging Face02placeholderlabs /exp-pool-repository-code-dolma2-tokenized Locus EXP Repository Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes664 downloads1mo agoHugging Face03placeholderlabs /Nemotron-SFT-Science-v2-Sharded Nemotron-SFT-Science-v2-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.text1M<n<10M0 likes515 downloads19d agoHugging Face04placeholderlabs /pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.tabular100K<n<1M0 likes463 downloads13d agoHugging Face05placeholderlabs /pretrain-ultra-fineweb-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,367,358,024 (1.4B) Trainable tokens 1,367,358,024 (1.4B) Documents 48,077 Shards 73 UTF-8 bytes 6,386,740,105 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-ultra-fineweb-mix-long-context.tabular10K<n<100K1 likes372 downloads13d agoHugging Face06placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes348 downloads8d agoHugging Face07placeholderlabs /pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 79,318,557,379 (79.3B) Trainable tokens 28,772,968,648 (28.8B) Documents 31,431,846 Shards 696 UTF-8 bytes 310,412,170,445 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.tabular10M<n<100M0 likes283 downloads8d agoHugging Face08placeholderlabs /GLM-5.1-Reasoning-Main-Sharded GLM-5.1 reasoning main — sequential shards Byte-preserving 100 MB JSONL shards of the main subset from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, by Jackrong, derived upstream from Kassadin88/GLM-5.1-1000000x. All credit for the original data and cleaning belongs to those publishers. Only main.jsonl is included. No filtering, shuffling, schema changes, tokenization or truncation. Original JSON fields and complete records are preserved. Shards retain upstream order; random shard… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/GLM-5.1-Reasoning-Main-Sharded.text100K<n<1M0 likes268 downloads19d agoHugging Face09placeholderlabs /pretrain-repository-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 77,467,752,674 (77.5B) Trainable tokens 77,467,752,674 (77.5B) Documents 24,040,993 Shards 740 UTF-8 bytes 317,294,681,915 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix.tabular10M<n<100M0 likes241 downloads8d agoHugging Face10placeholderlabs /exp-pool-commit-code-dolma2-tokenized Locus EXP Commit Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes239 downloads1mo agoHugging Face11placeholderlabs /Kimi-K2.5-Reasoning-General-Sharded Kimi-K2.5-Reasoning-General-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: General-Distillation.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.text100K<n<1M0 likes233 downloads19d agoHugging Face12placeholderlabs /Nemotron-SFT-Agentic-v2-Selected-Sharded Nemotron-SFT-Agentic-v2-Selected-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Agentic-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: data/tool_calling.jsonl, data/interactive_agent.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Agentic-v2-Selected-Sharded.text100K<n<1M0 likes184 downloads19d agoHugging Face13placeholderlabs /exp-pool-academic-dolma2-tokenized Locus EXP Academic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes176 downloads1mo agoHugging Face14placeholderlabs /pretrain-academic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 43,694,042,993 (43.7B) Trainable tokens 43,694,042,993 (43.7B) Documents 1,001,557 Shards 373 UTF-8 bytes 183,279,720,921 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-academic-mix-long-context.tabular1M<n<10M0 likes160 downloads13d agoHugging Face15placeholderlabs /pretrain-web-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 73,646,210,143 (73.6B) Trainable tokens 73,646,210,143 (73.6B) Documents 61,059,647 Shards 590 UTF-8 bytes 341,537,872,441 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix.tabular100M<n<1B0 likes141 downloads13d agoHugging Face16placeholderlabs /exp-pool-commit-code-raw Locus EXP Commit Code - shuffled raw proxy pool Deterministically shuffled commit-message and unified-diff documents with complete source metadata. MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums. The paired Dolma-2-tokenized repository preserves prompt masking for reproducible proxy training. tabular1M<n<10M0 likes133 downloads1mo agoHugging Face17placeholderlabs /exp-pool-olmo-web-dolma2-tokenized Locus EXP OLMo Web - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes132 downloads1mo agoHugging Face18placeholderlabs /pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 22,927,812,461 (22.9B) Trainable tokens 22,927,812,461 (22.9B) Documents 21,377,358 Shards 180 UTF-8 bytes 77,994,866,327 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.tabular10M<n<100M0 likes132 downloads13d agoHugging Face19placeholderlabs /pool-encyclopedic locus-1 encyclopedic pool The encyclopedic pool for locus-1 - one of seven cluster datasets that are mixed into the pretraining corpus. Every pool shares one schema, so a mixture is a query rather than a rebuild. 6,498,683 documents, 7,314,613,330 tokens under allenai/dolma2-tokenizer@5292e5d6c0f40b67cc765fe41bec991cf4345b5c Built from: HuggingFaceFW/finewiki (en @ 8bd13e72e6a0) Build settings digest: cbd87eae472c6b30 Fingerprint compatibility digest: 6ee2a61c2f91d6a1… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-encyclopedic.tabulartext-generation10M<n<100M0 likes122 downloads2mo agoHugging Face20placeholderlabs /exp-pool-encyclopedic-dolma2-tokenized Locus EXP Encyclopedic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.tabulartext-generation1M<n<10M0 likes122 downloads1mo agoHugging Face21placeholderlabs /exp-pool-nemotron-math-dolma2-tokenized Locus EXP Nemotron Math - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes109 downloads1mo agoHugging Face22placeholderlabs /pretrain-encyclopedic-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 587,625,128 (587.6M) Trainable tokens 587,625,128 (587.6M) Documents 23,631 Shards 9 UTF-8 bytes 1,978,753,989 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix-long-context.tabular10K<n<100K1 likes102 downloads13d agoHugging Face23placeholderlabs /pretrain-repository-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 130,158,375,824 (130.2B) Trainable tokens 130,158,375,824 (130.2B) Documents 2,578,578 Shards 1,168 UTF-8 bytes 535,260,344,241 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-repository-v2-mix-long-context.tabular1M<n<10M0 likes89 downloads8d agoHugging Face24placeholderlabs /pretrain-encyclopedic-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 7,321,112,013 (7.3B) Trainable tokens 7,321,112,013 (7.3B) Documents 6,498,683 Shards 58 UTF-8 bytes 27,544,569,761 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-encyclopedic-mix.tabular10M<n<100M0 likes67 downloads13d agoHugging Face25placeholderlabs /pretrain-nemotron-math-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 1,446,296,439 (1.4B) Trainable tokens 1,446,296,439 (1.4B) Documents 42,379 Shards 23 UTF-8 bytes 4,965,563,314 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix-long-context.tabular10K<n<100K0 likes64 downloads13d agoHugging Face26portuguese-benchmark-datasets /placeholder_tiebetext10K<n<100K0 likes55 downloads1y agoHugging Face27placeholderlabs /exp-pool-olmo-web-raw olmo_web shuffled raw proxy pool Documents remain raw UTF-8 with complete source metadata. Dolma-2 tokenization and 16K packing are intentionally deferred to training. MANIFEST.json pins source identity, shuffle policy, capacity, and checksums. tabular1M<n<10M0 likes52 downloads1mo agoHugging Face28placeholderlabs /exp-pool-academic-raw academic shuffled raw proxy pool Documents remain raw UTF-8 with complete source metadata. Dolma-2 tokenization and 16K packing are intentionally deferred to training. MANIFEST.json pins source identity, shuffle policy, capacity, and checksums. tabular100K<n<1M0 likes37 downloads1mo agoHugging Face29portuguese-benchmark-datasets /placeholder_tiebe2text1K<n<10K0 likes26 downloads1y agoHugging Face30placeholderlabs /exp-pool-repository-code-raw Locus EXP Repository Code - shuffled raw proxy pool Deterministically shuffled repository-code documents with complete source metadata. MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums. The paired Dolma-2-tokenized repository is ready for reproducible proxy training. tabular100K<n<1M0 likes24 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.