CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BramVanroy /CommonCrawl-CreativeCommons-strict Common Crawl Creative Commons Corpus Strict (C5s) A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that: are also present in the FineWeb or FineWeb-2 datasets; have no license disagreement (all found licenses have the same type; version number might differ); are not "non-commercial" ("nc" in license); are not "cc-unknown"; do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.texttext-generation10M<n<100M2 likes559 downloads1y agoHugging Face02callofthenight1 /gaokao-sft-chinese-strict-abcd-v3 Gaokao SFT Chinese Strict ABCD V3 This dataset is the cleaned Chinese SFT release that keeps only single-choice samples where A, B, C, and D all have explicit option-level analysis. Composition Total samples: 88466 Train samples: 86670 Validation samples: 1796 Subject Counts { "biology": 33104, "chemistry": 35796, "english": 174, "general_exam": 7982, "geography": 888, "history": 229, "physics": 9986, "politics": 307 } Fields id… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd-v3.texttext-generation10K<n<100K1 likes125 downloads5mo agoHugging Face03Lelonthecodeur /strict-verification-reasoning Strict Verification Reasoning Dataset Description A dataset for training language models to verify facts, check sources, evaluate arguments, and avoid overthinking. Content 1,010,000 examples 5 categories: anti-overthink, comparisons, strict facts, strict sources, strict arguments English language Categories Category Description % Anti-Overthink Simple, direct answers 15% Comparisons Hallucination vs correct answer 20%… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/strict-verification-reasoning.texttext-generation1M<n<10M1 likes74 downloads11d agoHugging Face04tuxevil /Home-Assistant-Requests-V5.2-Native-Strict Home Assistant Requests V5.2 Native Strict Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model. Contract: ha-native-tool-calling-v2. Frozen snapshot Split Rows Direct speech Multi-call Maximum rendered tokens train 3,806 340 78 3,098 validation 530 52 4 2,874 test 633 102 22 2,925 Tokenizer audit: model: unsloth/Qwen3-4B-Instruct-2507 revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.2-Native-Strict.texttext-generation1K<n<10K0 likes53 downloads2mo agoHugging Face05Pradheep1647 /strictly-speaking Strictly Speaking Does a model's mathematical understanding hold up strictly speaking, at Lean-grade precision, or is it only right in the ordinary, looser sense that informal writing usually gets away with? Each row is derived from a real, already-formalized Lean 4 theorem (drawn from Pradheep1647/lean-verifier-formalizations). The theorem's informal statement is split into two pieces: the hypotheses/setup (prompt), and the conclusion that was elided from it (answer) - the… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/strictly-speaking.texttext-generationn<1K0 likes46 downloads5d agoHugging Face06eewer /qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed Qwen3 4B Thinking SFT v54 Processed Training View This dataset is the processed and filtered training view used by the v54 Qwen3-4B-Thinking SFT recipe. It starts from eewer/swerebench-traces-raw-source-targeted-limitations-compaction-full-20260616-2030 and uses the strict-passed raw2030 mini-swe aligned view. Rows are compressed JSONL.zst files under data/. Each row contains a top-level messages column, optional tools, and scalar source mapping fields such as source_uuid… See the full description on the dataset page: https://huggingface.co/datasets/eewer/qwen3-4b-thinking-sft-v54-raw2030-strictpassed-processed.text-generation1K<n<10K0 likes44 downloads3mo agoHugging Face07callofthenight1 /gaokao-sft-chinese-strict-abcd Gaokao SFT Chinese Balanced This is the strict balanced Chinese SFT dataset version. Only multiple-choice samples with explicit A/B/C/D option-level explanations are kept in this balanced release. Composition Total samples: 645 Train samples: 632 Validation samples: 13 Subject Counts { "biology": 199, "chemistry": 170, "english": 13, "geography": 34, "history": 118, "physics": 77, "politics": 34 } Fields id lang subject source… See the full description on the dataset page: https://huggingface.co/datasets/callofthenight1/gaokao-sft-chinese-strict-abcd.texttext-generationn<1K0 likes41 downloads5mo agoHugging Face08edithatogo /qwen3-hermes-strict-toolcall-synthetic-v4 Qwen3 Hermes Strict Tool-Call Synthetic V4 Registry status Registry ID: edithatogo/qwen3-hermes-strict-toolcall-synthetic-v4 Family: hermes Repository role: canonical_training_dataset Canonical dataset: edithatogo/qwen3-hermes-strict-toolcall-synthetic-v4 Operational status: active Rights status: apache-2.0-synthetic Authoritative catalog: edithatogo/dataset-estate-registry Origin and provenance Origin repository:… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/qwen3-hermes-strict-toolcall-synthetic-v4.texttext-generationn<1K0 likes33 downloads2mo agoHugging Face09tuxevil /Home-Assistant-Requests-V5.1-Native-Strict Home Assistant Requests V5.1 Native Strict Private research dataset for supervised fine-tuning and regression testing of a small Home Assistant native tool-calling model. Contract: ha-native-tool-calling-v2. Frozen snapshot Split Rows Direct speech Multi-call Maximum rendered tokens train 3,806 340 78 3,098 validation 530 52 4 2,874 test 633 102 22 2,925 Tokenizer audit: model: unsloth/Qwen3-4B-Instruct-2507 revision:… See the full description on the dataset page: https://huggingface.co/datasets/tuxevil/Home-Assistant-Requests-V5.1-Native-Strict.texttext-generation1K<n<10K0 likes32 downloads2mo agoHugging Face10kilicai /turkish-sft-clean-v5-strict-pruned Clean Turkish SFT v5 Strict Pruned This is the strictest cleaned version. It prioritizes correctness over row count: content errors, mislabeled rows, unsafe/refusal-contaminated examples, exact duplicates, and near-duplicates in template-heavy categories were deleted rather than padded. What was fixed Removed category-contaminated rows with task-specific validators. Removed exact duplicates by normalized SHA-256 fingerprint. Removed near-duplicates in formal-writing… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-clean-v5-strict-pruned.texttext-generation10K<n<100K0 likes22 downloads4mo agoHugging Face11tuandunghcmut /deepresearch-9k-tool-calling-strictgated DeepResearch-9K — Tool Calling Format (Strict) Strict converted version of artillerywu/DeepResearch-9K. Key difference from the standard version: When an assistant message contains tool_calls, the content field is null. <think> reasoning blocks are dropped from tool-calling turns. Dataset Summary Property Value Source artillerywu/DeepResearch-9K Samples 3,974 Tool search Format OpenAI-compatible messages + tools_json Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/deepresearch-9k-tool-calling-strict.texttext-generation1K<n<10K0 likes17 downloads7mo agoHugging Face12ksu-help /babylm2026-strict-small-custom-corpora BabyLM 2026 Strict-Small — custom training corpora Datasheet for the custom (teacher-generated) corpora used to train our BabyLM 2026 Strict-Small submissions. Each corpus stays within the 10M-word Strict-Small budget (~9.98M words seen per condition). Prepared following the Datasheets for Datasets framework (Gebru et al., 2021). Contents Path What Words real_base/babylm10m_clean.txt Cleaned BabyLM Strict-Small English base corpus (see cleaning below)… See the full description on the dataset page: https://huggingface.co/datasets/ksu-help/babylm2026-strict-small-custom-corpora.texttext-generation1M<n<10M0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.