CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3.5 Pool The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.text-generation8 likes70k downloads2mo agoHugging Face02allenai /dolma3_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 pool, pre–quality upsampling and mixing. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3 Pool The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.texttext-generation10B<n<100B40 likes69k downloads7mo agoHugging Face03allenai /dolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Dolmino pool; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025 Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125 Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B. Dataset Sources Source Category Tokens Documents TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.text-generation8 likes35k downloads9mo agoHugging Face04placeholderlabs /locus-commit-pool-v1 Locus Commit Pool v1 Native Git history, preserved as replayable software changes Commit message · complete selected before-state · unified patches · native object IDs · provenance · experimental labels Locus Commit Pool v1 is a large evidence pool for studying and training on how real software changes. Each document represents one surviving single-parent, multi-file Git commit. It keeps the commit message, the selected files as they existed before the… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-commit-pool-v1.text-generation0 likes7k downloads1mo agoHugging Face05poolside-laguna-hackathon /patchrecoverygym-laguna PatchRecoveryGym for Laguna Submitted by: Kannappan Sirchabesan (@kannappans) · Poolside Research Hackathon (Foundations track) A reproducible eval + RL environment that tests whether a coding agent can recover from a wrong first attempt — a real, under-measured agentic-coding weakness. Built for Poolside Laguna XS.2 on dependency-migration repair tasks. 📦 Installable Verifiers environment on the Prime Hub · 🎯 deterministic hidden-test reward · 🔁 144-candidate reranking… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/patchrecoverygym-laguna.text-generationn<1K0 likes2.8k downloads4mo agoHugging Face06placeholderlabs /exp-pool-repository-code-dolma2-tokenized Locus EXP Repository Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes883 downloads1mo agoHugging Face07placeholderlabs /locus-repo-code-pool-v2 Locus repository code pool The Locus repository code pool contains curated repository-version documents for code-language-model research. Each row combines the useful files from one Git repository snapshot into a single deterministic text document while retaining source provenance, file order, language and test signals, health evidence, and stable content hashes. This release is a candidate acquisition pool, not a ready-made training split. Use a globally deduplicated manifest… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-repo-code-pool-v2.text-generation1M<n<10M0 likes812 downloads1mo agoHugging Face08Emulated-Inc /forum-competition-math-training-pool Forum competition mathematics training pool Olympiad and contest mathematics from three public datasets, gathered at pinned revisions and shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 287091 rows across three folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 282140 rows, every row labelled with… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/forum-competition-math-training-pool.texttext-generation100K<n<1M0 likes736 downloads11d agoHugging Face09Emulated-Inc /olympiad-math-training-pool Olympiad mathematics training pool Public olympiad and competition mathematics, four datasets gathered at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 229052 rows across four folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 225822 rows, every row labelled with the dataset it… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/olympiad-math-training-pool.texttext-generation100K<n<1M0 likes668 downloads10d agoHugging Face10hkust-nlp /dart-math-pool-gsm8k [!NOTE] This dataset is the data pool synthesized from the query set of the GSM8K training set, containing all answer-correct samples and other metadata produced during the work. DART-Math-* datasets are extracted from dart-math-pool-* data pools. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-gsm8k.text-generation1M<n<10M2 likes606 downloads2y agoHugging Face11Emulated-Inc /competition-math-training-pool Competition mathematics training pool Public competition mathematics, six datasets gathered at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 1951046 rows across six folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 1125451 rows, every row labelled with the dataset it came from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-math-training-pool.texttext-generation1M<n<10M0 likes420 downloads11d agoHugging Face12placeholderlabs /exp-pool-commit-code-dolma2-tokenized Locus EXP Commit Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes418 downloads1mo agoHugging Face13hkust-nlp /dart-math-pool-math [!NOTE] This dataset is the data pool synthesized from the query set of the MATH training set, containing all answer-correct samples and other metadata produced during the work. DART-Math-* datasets are extracted from dart-math-pool-* data pools. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math.texttext-generation1M<n<10M8 likes350 downloads2y agoHugging Face14placeholderlabs /exp-pool-academic-dolma2-tokenized Locus EXP Academic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes295 downloads1mo agoHugging Face15placeholderlabs /exp-pool-olmo-web-dolma2-tokenized Locus EXP OLMo Web - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes243 downloads1mo agoHugging Face16keiseoq /pool-qkbue 传奇私服分布式路由与自动化接口索引库 - Batch 002 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 今日传奇私服网 - 传奇私服推荐 - 新开传奇私服 —— 承载源站 🌐 alpha.agergames.com 👉 传奇私服推荐 - 新开传奇私服 - 传奇私服发布 —— 承载源站 🌐 alpha.amoygame.com 👉 传奇私服999 - 传奇私服 - 传奇私服网盘 —— 承载源站 🌐 alpha.gamecctv.com 👉 传奇私服 - 传奇私服推荐网 - 传奇私服999 —— 承载源站 🌐 alpha.gamerling.com 👉 传奇私服发布 - 传奇私服网 - 今日传奇私服 —— 承载源站 🌐 alpha.games-vr.com 👉 新开传奇私服发布 - 传奇私服网 - 传奇私服 —— 承载源站 🌐 alpha.j5games.com 👉 传奇私服 -… See the full description on the dataset page: https://huggingface.co/datasets/keiseoq/pool-qkbue.text-generation0 likes242 downloads2mo agoHugging Face17placeholderlabs /exp-pool-encyclopedic-dolma2-tokenized Locus EXP Encyclopedic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.tabulartext-generation1M<n<10M0 likes220 downloads1mo agoHugging Face18placeholderlabs /exp-pool-nemotron-math-dolma2-tokenized Locus EXP Nemotron Math - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes210 downloads1mo agoHugging Face19makesi1 /pool-g4cl WhatsApp网页版分布式路由与自动化接口索引库 - Batch 005 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:WhatsApp网页版)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 WhatsApp网页版-质量纠纷处理-WhatsApp Web 网页版 —— 承载源站 🌐 Agile.alliance-whatapp.hl.cn 👉 WhatsApp网页版-未读消息一键过滤-WhatsApp Web 网页版 —— 承载源站 🌐 Agile.alphas-whatapp.hl.cn 👉 WhatsApp网页版-WhatsApp Web 网页版-CRM侧边栏同步 —— 承载源站 🌐 Agile.badge-whatapp.hl.cn 👉 WhatsApp网页版-全局搜索历史凭证-WhatsApp Web 网页版 —— 承载源站 🌐 Agile.battle-whatapp.hl.cn 👉 WhatsApp网页版-WhatsApp Web… See the full description on the dataset page: https://huggingface.co/datasets/makesi1/pool-g4cl.text-generation0 likes186 downloads2mo agoHugging Face20Emulated-Inc /logic-grid-puzzles-training-pool Logic grid puzzles training pool Logic grid puzzles: a row of positions, a handful of attributes with one value per position, and a list of clues that together admit exactly one arrangement. Two sets drawn for this pool by generators run here under the seeds recorded below, and two public datasets read at the pinned revisions named below, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 390945 rows, one JSON… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logic-grid-puzzles-training-pool.texttext-generation100K<n<1M0 likes166 downloads10d agoHugging Face21model-organisms-for-real /hs3-prompt-pool-topic-judged hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training Prompts only (no completions). Every user prompt in model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated 35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0) for the high-level topic of both quirk families. Why Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.tabulartext-generation100K<n<1M0 likes156 downloads5d agoHugging Face22Emulated-Inc /python-unit-test-training-pool Python unit test training pool A pool of public data for training a model to write tests for Python code. It is a straight collection of open datasets, not a new corpus: every row comes from one of the sources below, at the revision named, and the only rows removed are the ones an overlap filter flagged against held-out material this pool is kept separate from. Every row of the normalised layer pairs a program with tests for it. That is the point of the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.texttext-generation1M<n<10M0 likes145 downloads10d agoHugging Face23Emulated-Inc /code-execution-trace-training-pool Code execution trace training pool Public Python code paired with one concrete call and the value that call returns. Every value in this pool was computed by running the code, not copied from a label. The data is laid out twice, and either layer may be used. pool/ Every source rewritten into one shape, 2174322 rows over 11 gzipped parts, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this pool code… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/code-execution-trace-training-pool.texttext-generation1M<n<10M1 likes140 downloads10d agoHugging Face24Emulated-Inc /symbolic-calculus-training-pool Symbolic calculus training pool Single variable calculus exercises with closed form answers: derivatives of composite expressions, indefinite and definite integrals, limits of indeterminate forms, and coefficients of Maclaurin series, with some multivariable operators in one of the sources. A set generated for this pool and two public datasets read at the pinned revisions named below, laid out twice. Train on either layer or on both. pool.jsonl Every source… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/symbolic-calculus-training-pool.texttext-generation1K<n<10K2 likes132 downloads10d agoHugging Face25Emulated-Inc /constrained-instruction-training-pool Constrained instruction training pool Public prompts for writing tasks, many of them carrying a constraint a program can check, from ten datasets read at the pinned revisions named below and one layer built here from them. The pool is laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 462652 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/constrained-instruction-training-pool.texttext-generation100K<n<1M0 likes126 downloads10d agoHugging Face26Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes125 downloads10d agoHugging Face27placeholderlabs /pool-encyclopedic locus-1 encyclopedic pool The encyclopedic pool for locus-1 - one of seven cluster datasets that are mixed into the pretraining corpus. Every pool shares one schema, so a mixture is a query rather than a rebuild. 6,498,683 documents, 7,314,613,330 tokens under allenai/dolma2-tokenizer@5292e5d6c0f40b67cc765fe41bec991cf4345b5c Built from: HuggingFaceFW/finewiki (en @ 8bd13e72e6a0) Build settings digest: cbd87eae472c6b30 Fingerprint compatibility digest: 6ee2a61c2f91d6a1… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-encyclopedic.tabulartext-generation10M<n<100M0 likes123 downloads2mo agoHugging Face28Emulated-Inc /python-functions-training-pool Python function-writing training pool A pool of public data for training a model to write Python functions. It is a straight collection of open datasets, not a new corpus: every row comes from one of the sources below, at the revision named, and the only rows removed are the ones an overlap filter flagged against held-out material this pool is kept separate from. Rows in the normalised layer: 5756045. Rows in the raw layer: 6258415. The two layers pool/ holds the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-functions-training-pool.texttext-generation1M<n<10M0 likes106 downloads11d agoHugging Face29Emulated-Inc /procedural-reasoning-training-pool Procedural reasoning training pool Reasoning questions from 101 procedural generators, each of which writes a question, computes its own answer and ships a verifier that scores an attempt at it, plus a collection of solved Sudoku puzzles. Every answer is short and exactly checkable, so a trained model can be marked against the key by a program and no judge is needed. Laid out twice. Train on either layer or on both. pool.jsonl Every generator rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/procedural-reasoning-training-pool.texttext-generation100K<n<1M0 likes99 downloads10d agoHugging Face30Emulated-Inc /competition-answer-math-training-pool Competition answer mathematics training pool Competition mathematics problems that ask for a single final answer, three public datasets gathered at pinned revisions, shipped twice over. sources/ holds each dataset the way its publisher ships it, in its own file format with its own fields and nothing renamed, 113045 rows across three folders. pool/ holds the union of those same datasets in one format, one JSON object per line, deduplicated by problem text and reduced to 107637… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/competition-answer-math-training-pool.texttext-generation10K<n<100K0 likes95 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.