CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01billion-word-benchmark /lm1bA benchmark corpus to be used for measuring progress in statistical language modeling. This has almost one billion words in the training data.text-generation19 likes1.6k downloads3y agoHugging Face02cloverx-id /Lumina-Math-Foundations-1B Lumina-Math-Foundations-1B Lumina-Math-Foundations-1B is an industrial-scale foundational mathematical reasoning dataset in Indonesian, Dataset Summary Language: Indonesian (id) with LaTeX mathematical formulas. Scale: 1 Billion Synthetic High-Fidelity Mathematical Reasoning instances.--- Data Schema & Field Breakdown Field Name Type Description problem string 100% pure human natural language problem statement with standard LaTeX math… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/Lumina-Math-Foundations-1B.texttext-generation1B<n<10B2 likes1.4k downloads2d agoHugging Face03codelion /sutra-1B Sutra 1B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational patterns Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-1B.tabulartext-generation100K<n<1M2 likes1.2k downloads7mo agoHugging Face04SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes482 downloads4mo agoHugging Face05shreyansh12183 /shreyansh-1B-SLM-pretrain-stem-english 📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentations. 🔬 Dataset Overview Designed specifically for pre-training and continuous pre-training (CPT) of Small Language Models (SLMs) in the 1B–3B parameter regime: High Information Density: Filtered to… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.texttext-generation10M<n<100M0 likes376 downloads3d agoHugging Face06AbstractPhil /human-templated-captions-1bcsv delimiter is = ".,|,." apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading. This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon. Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.texttext-generation100M<n<1B1 likes320 downloads1y agoHugging Face07krisbailey /RedPajama-Data-V2-1B RedPajama-Data-V2 1B Dataset Description This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data. Motivation RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.texttext-generation100K<n<1M0 likes309 downloads8mo agoHugging Face08krisbailey /cosmopedia-1b Cosmopedia 1B Dataset Description This is a 1 Billion token subset of the krisbailey/cosmopedia-10B dataset, which itself is a 10B subset of HuggingFaceTB/cosmopedia. It was created by uniformly sampling approximately 9.5% of the 10B dataset, ensuring the data distribution remains consistent with the source. Motivation While the 10B dataset is a "Goldilocks" size for many experiments, 1B tokens is the standard size for rapid prototyping, scaling law… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-1b.texttext-generation1M<n<10M0 likes273 downloads8mo agoHugging Face09QuixiAI /SYN-1B SYN-1B Dataset Summary SYN-1B is a 1.04B-token synthetic language-modeling corpus of rule-governed text streams. Each row is a decoded instance in which the text establishes facts, mappings, bindings, or simple generative rules; later spans may revise those rules, swap bindings, delay a query over long filler, or present a null control with event-like surface text that should not change the answer. The dataset is intended as structured pretraining data, not as a… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/SYN-1B.tabulartext-generation1M<n<10M8 likes232 downloads3mo agoHugging Face10LLM-OS-Models /LFM2.5-8B-A1B-KO-CPT-DATA LFM2.5-8B-A1B Korean CPT Data Prepared Korean continued-pretraining data for LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL. Files data/ko_cpt_mix_full_lfmstyle_20260627.jsonl: prepared full CPT corpus with one JSON object per line and a text field metadata/ko_cpt_mix_full_lfmstyle_20260627.stats.json: corpus statistics metadata/ko_cpt_mix_full_lfmstyle_20260627.stats.json.full_report.json: per-source preprocessing report metadata/ko_cpt_sources_full_20260627.json:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-DATA.text-generation100K<n<1M0 likes195 downloads3mo agoHugging Face11QuixiAI /QuixiMath-1B QuixiMath-1B QuixiMath is brought to you by Eric Hartford and QuixiAI https://github.com/QuixiAI/QuixiMath Dataset Summary QuixiMath-1B is a synthetic math reasoning corpus generated from the QuixiMath procedural problem generators. Each record contains a natural-language problem, explicit step-by-step scratchpad opcodes, a canonical final answer, and metadata for filtering or reweighting by skill, operation, grade band, and relative difficulty. The canonical… See the full description on the dataset page: https://huggingface.co/datasets/QuixiAI/QuixiMath-1B.tabulartext-generation10M<n<100M7 likes126 downloads3mo agoHugging Face12Urdatorn /sphragis-olmo1b-adaptation-corpus Sphragis OLMo-1B adaptation corpus Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek before authorship-language-model training. It contains only OGA whole works whose TLG author occurs in neither Sphragis benchmark. Text has the exact model-facing benchmark surface form: polytonic-aware lowercasing with grc_utils.lower_grc, removal of all editorial punctuation, normalization of whitespace, and removal of consonant-final elision marks. Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.texttext-generation1K<n<10K0 likes120 downloads27d agoHugging Face13FrankCCCCC /lm1b LM1B - One Billion Word Benchmark Dataset Description The One Billion Word Benchmark is a large language modeling dataset. It contains approximately one billion words of training data derived from news articles. How was this dataset built? We download the full LM1B dataset from TensorFlow Datasets (TFDS) and convert it to HuggingFace format automatically. The full script is in lm1b.py. The required environment is: tensorflow==2.20.0 tensorflow-datasets==4.9.9… See the full description on the dataset page: https://huggingface.co/datasets/FrankCCCCC/lm1b.texttext-generation10M<n<100M1 likes109 downloads8mo agoHugging Face14krisbailey /falcon-refinedweb-1B Falcon RefinedWeb 1B Dataset Description This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data. Motivation RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.texttext-generation1M<n<10M0 likes92 downloads8mo agoHugging Face15Dodosoomro /simple-100m-pretrain-1b Simple-100M Pretraining Dataset (1B Tokens) A training-optimized, packed pretraining dataset for ~100M parameter language models. Built for reproducibility, minimal runtime overhead, and exact mixing ratios. 🎯 Purpose This dataset was created to train Simple-100M, a decoder-only Transformer targeting: ✅ Beat GPT-2-70M perplexity with minimal complexity ✅ Reproducible artifacts with exact token accounting ✅ Zero runtime preprocessing (ready-to-train) Target… See the full description on the dataset page: https://huggingface.co/datasets/Dodosoomro/simple-100m-pretrain-1b.texttext-generation1B<n<10B0 likes79 downloads5mo agoHugging Face16ecreeth /1b-smollm-corpus SmolLM-Corpus — 1B Token Subset A curated 1-billion-token English pretraining corpus sampled from HuggingFaceTB/smollm-corpus, designed for training small language models (~20M parameters). Dataset Composition Source Ratio Tokens Documents FineWeb-Edu (dedup) 87% ~870M 849,577 Cosmopedia v2 13% ~130M 161,889 Total 100% ~1B 1,011,466 Rationale for the Split The 87/13 ratio mirrors the natural token distribution of the full… See the full description on the dataset page: https://huggingface.co/datasets/ecreeth/1b-smollm-corpus.texttext-generation1M<n<10M1 likes69 downloads4mo agoHugging Face17krisbailey /fineweb-edu-1B FineWeb-Edu 1B Dataset Description FineWeb-Edu 1B is a high-quality, stratified subset of the HuggingFaceFW/fineweb-edu dataset. It contains approximately 1 billion tokens of educational web text, carefully sampled to preserve the original distribution of source data (CommonCrawl dumps). This dataset provides an accessible, lightweight alternative to the larger FineWeb-Edu subsets (like sample-10BT or sample-100BT) while maintaining the same data diversity and quality… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/fineweb-edu-1B.tabulartext-generation100K<n<1M0 likes68 downloads8mo agoHugging Face18MLBricks /fineweb-edu-1b MLBricks FineWeb-Edu 1B MLBricks curated datasetUpstream: HuggingFaceFW/fineweb-edu @ 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9Source config: sample-10BTSource split: trainLicense: ODC-By 1.0 MLBricks maintains this dataset as a stable Studio preset and does not claim ownership of the underlying source material. Edition Destination: MLBricks/fineweb-edu-1b Rows: 969,429 GPT-2 tokens: 1,000,000,000 Main Studio column: text Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/fineweb-edu-1b.texttext-generation100K<n<1M0 likes64 downloads19d agoHugging Face19sdbhud1b /Chinese_qa Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/sdbhud1b/Chinese_qa.text-classification10B<n<100B0 likes63 downloads2y agoHugging Face20hanseungwook /GSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT Verified self-generated GSM8K reasoning 64 independently sampled completions are generated per prepared question. Final answers are checked against the source answer. Among complete, correctly formatted correct completions whose CoT passes the final-result-statement and combined length checks, one sample is selected uniformly at random using a reproducible per-question seed. CoT length does not rank eligible samples. The final result belongs in the separate final-answer line of… See the full description on the dataset page: https://huggingface.co/datasets/hanseungwook/GSM8K-Aug-Llama-3.2-1B-Instruct-Correct-CoT.tabulartext-generation100K<n<1M0 likes60 downloads8d agoHugging Face21MLBricks /openwebmath-1b MLBricks OpenWebMath 1B MLBricks curated datasetUpstream: open-web-math/open-web-math @ fde8ef8de2300f5e778f56261843dab89f230815Source config: defaultSource split: trainLicense: ODC-By 1.0 MLBricks maintains this dataset as a stable Studio preset and does not claim ownership of the underlying source material. Edition Destination: MLBricks/openwebmath-1b Rows: 469,065 GPT-2 tokens: 1,000,000,000 Main Studio column: text Source-rights notice… See the full description on the dataset page: https://huggingface.co/datasets/MLBricks/openwebmath-1b.texttext-generation100K<n<1M0 likes56 downloads19d agoHugging Face22Ba2han /ultra-fineweb-tokenized-1B-short Ultra-FineWeb Tokenized 2B Status: complete A tokenized subset streamed from [openbmb/Ultra-FineWeb] using its en split. The Parquet data has exactly one column: input_ids (list<int32>). Processing Source revision: 7ddd4170ce03e0afbd7d9b80d4bc0b8eebf877e4 Tokenizer: Ba2han/TR_CPT1 Tokenizer revision: d32fb40763740cca0a7c04b4549d2129edeafa9a Source field: content Filter: score > 0.75 Stored sequence length: 20 to 888 tokens, inclusive Target: approximately 1,000… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/ultra-fineweb-tokenized-1B-short.text-generation0 likes50 downloads2mo agoHugging Face23javiersgjavi /fineweb-1BT FineWeb-1BT: 1 Billion Token Subset Dataset Description FineWeb-1BT is a carefully curated 1 billion token subset of the HuggingFaceFW/fineweb dataset, sampled exclusively from the official 10BT FineWeb subset. It was created using true random sampling to ensure unbiased representation across the entire 10BT corpus. Provenance: This 1BT subset is a uniform sample drawn from the official FineWeb 10BT split (not from the full raw stream), preserving its language/quality… See the full description on the dataset page: https://huggingface.co/datasets/javiersgjavi/fineweb-1BT.tabulartext-generation1M<n<10M0 likes47 downloads1y agoHugging Face24Neil1b /sf9212 传奇私服分布式路由与自动化接口索引库 - Batch 005 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 新开传奇私服迷失版-2026传奇SF 补丁包-2026传奇私服统治 —— 承载源站 🌐 min-sf.0002sf.com 👉 传奇私服发布网推荐排行-传奇私服平台高爆服-2026传奇私服汉化 —— 承载源站 🌐 min-sf.0211sf.cn 👉 传奇私服大全-传奇私服发布网免费-1.76复古传奇私服 —— 承载源站 🌐 min-sf.06sf.cn 👉 传奇SF 之选-经典传奇私服-找私服-手机版传奇世界sf直装包 —— 承载源站 🌐 min-sf.34sf.net 👉 本月传奇私服新服-2026传奇SF 侠客版-2026传奇SF 传奇发布 —— 承载源站 🌐 min-sf.5566sf.com 👉 本月传奇私服任务-元宝好打的传奇私服-本月传奇私服亚服 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9212.text-generation0 likes46 downloads3d agoHugging Face25Neil1b /sf9213 传奇私服分布式路由与自动化接口索引库 - Batch 004 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 耐玩传奇私服新开发布-传奇SF 神器版-1.76金币小极品私服 —— 承载源站 🌐 new-sf.0002sf.com 👉 全新传奇私服星辰版-火爆传奇私服网站-本月传奇私服微变版 —— 承载源站 🌐 new-sf.34sf.net 👉 传奇sf官网-传奇私服发布网变态版-火爆传奇私服-找私服 —— 承载源站 🌐 new-sf.6sifu.com 👉 原汁原味热血传奇私服-传奇sf开服预告今天-复古传奇私服网 —— 承载源站 🌐 new-sf.7xingame.com 👉 传奇SF 门户-传奇私服发布网任务-传奇新开私服 —— 承载源站 🌐 new-sf.8thgames.com 👉 传奇SF 副本-本月传奇私服火爆版-传奇sf封号怎么解 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9213.text-generation0 likes45 downloads3d agoHugging Face26aoiandroid /minicpm5-1b-quantization-benchmark openbmb/MiniCPM5-1B 次世代量子化(Quanto FP8 / INT4 vs BNB 4bit)実測ベンチマークレポート 対象モデル: openbmb/MiniCPM5-1B (1.16B parameters, 128k context, LlamaForCausalLM) 検証ハードウェア: NVIDIA GeForce RTX 4070 Ti (12GB GDDR6X, Ada Lovelace, Compute Capability 8.9, 第4世代Tensor Core) 実行環境: Windows / Python 3.13 / PyTorch 2.6.0+cu124 / transformers 4.57.6 / optimum-quanto 0.2.7 / bitsandbytes 0.50.0 検証日: 2026-09-19 12:12:34 1. エグゼクティブサマリー(全体比較) NVIDIA GeForce RTX 4070 Ti 実機環境において、標準ネイティブ… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/minicpm5-1b-quantization-benchmark.texttext-generationn<1K0 likes44 downloads5d agoHugging Face27Neil1b /sf9215 传奇私服分布式路由与自动化接口索引库 - Batch 001 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 2026全新传奇私服开服-精品新开[传奇私服]-[传奇私服]网今日发布 —— 承载源站 🌐 api.718games.com.cn 👉 传奇私服发布网-变态[传奇私服]发布网-[传奇私服]网欧服-2026全新传奇私服开服 —— 承载源站 🌐 api.999sif.com.cn 👉 传奇私服发布网-[传奇私服]稳定发布网-网通低延迟热血[传奇sf]-2026全新传奇私服开服 —— 承载源站 🌐 api.chuangqi-sifu.com.cn 👉 传奇私服发布网-今日[传奇私服]新开服-[传奇私服]榜单-2026全新传奇私服开服 —— 承载源站 🌐 api.chuanqisfu.com.cn 👉 2026全新传奇私服开服-今日[传奇私服]雷霆版-今日新开传奇高爆服 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9215.text-generation0 likes42 downloads3d agoHugging Face28Neil1b /cq92106 传奇私服分布式路由与自动化接口索引库 - Batch 003 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 2026全新传奇私服开服-[传奇私服]找不到黑屏-1.76合击传奇发布网新开服 —— 承载源站 🌐 g1.718games.com.cn 👉 2026全新传奇私服开服-今日[传奇私服]新服榜-本月[传奇私服]下载站 —— 承载源站 🌐 g1.999sif.com.cn 👉 2026全新传奇私服开服-公益[传奇sf]-微变[传奇私服]发布网 —— 承载源站 🌐 g1.chuangqi-sifu.com.cn 👉 2026全新传奇私服开服-今日[传奇私服]新服推荐-绿色最新[传奇私服] —— 承载源站 🌐 g1.chuanqisfu.com.cn 👉 2026全新传奇私服开服-修仙单职业热血[传奇sf]-[传奇私服]开服公告 —— 承载源站 🌐 g1.game-12hf.cn 👉… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/cq92106.text-generation0 likes42 downloads3d agoHugging Face29Neil1b /sf9210 传奇私服分布式路由与自动化接口索引库 - Batch 006 本项目托管了用于全球多元化检索与高可用分布式网络路由数据(核心挂载:传奇私服)。 📂 区域节点集群子目录 (Spider Pool Indexes) 👉 传奇私服排行榜-传奇私服发布网火爆发布-找sf哪个好 —— 承载源站 🌐 min-sf.1440game.com 👉 热血传奇sf经典怀旧-传奇私服发布网-嘟嘟传奇热血传奇私服 —— 承载源站 🌐 min-sf.16hsf.com 👉 2026传奇SF 推荐网-刚开私服晚上8点沙捐首区-传奇私服开服教程 —— 承载源站 🌐 min-sf.18183-game.com.cn 👉 2026传奇SF 将军-传奇SF 心得-传奇私服汉化版 —— 承载源站 🌐 min-sf.18183gamesf.com.cn 👉 传奇sf发布网-sf999传奇发布网入口-耐玩传奇私服发布网 —— 承载源站 🌐 min-sf.1satori.com 👉 本月传奇私服私服榜-新开传奇私服主站-传奇SF… See the full description on the dataset page: https://huggingface.co/datasets/Neil1b/sf9210.text-generation0 likes40 downloads3d agoHugging Face30Yxanul /experimental-pretrain-1b Dataset Card for Experimental Pretraining Dataset 1B Dataset Details Dataset Description A meticulously curated 1 billion token dataset optimized for experimental pretraining of small language models. This dataset represents a balanced mixture of the highest quality educational content (60%), mathematical reasoning (30%), and Python code (10%), specifically designed for rapid experimentation and research in language model training. Curated by: Yxanul… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/experimental-pretrain-1b.texttext-generation100K<n<1M0 likes36 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.