CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes3.8k downloads1mo agoHugging Face02VibrantVista /TTCW-Based-Review TTCW Creative Writing Evaluation Dataset If you use this dataset in your research, please cite our paper — it helps support ongoing academic work. Citation details are at the bottom of this page. Dataset Description Summary A supervised fine-tuning (SFT) dataset for training LLMs to act as creative writing evaluators. Each example contains a creative story and four message-format columns representing different evaluation objectives — from… See the full description on the dataset page: https://huggingface.co/datasets/VibrantVista/TTCW-Based-Review.tabulartext-generation100K<n<1M2 likes2k downloads4mo agoHugging Face03Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes1.4k downloads1mo agoHugging Face04LLaMAX /BenchMAX_Rule-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios. We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.texttext-generation1K<n<10K1 likes623 downloads2y agoHugging Face05ftajwar /maxrl_qwen3_4B_base_polaris_rollouts MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts) Every training rollout from an online RL run, with exact token ids, sampling log-probs, and raw rewards — usable as a replay buffer to study off-policy RL for LLM reasoning completely offline. The run: Qwen3-4B-Base trained with the maxRL advantage estimator (A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt; maxRL paper) and a pure REINFORCE loss (L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.tabulartext-generation1M<n<10M0 likes598 downloads2mo agoHugging Face06Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes584 downloads1mo agoHugging Face07KingNish /reasoning-base-20k Dataset Card for Reasoning Base 20k Dataset Details Dataset Description This dataset is designed to train a reasoning model. That can think through complex problems before providing a response, similar to how a human would. The dataset includes a wide range of problems from various domains (science, coding, math, etc.), each with a detailed chain of thought (COT) and the correct answer. The goal is to enable the model to learn and refine its reasoning process… See the full description on the dataset page: https://huggingface.co/datasets/KingNish/reasoning-base-20k.texttext-generation10K<n<100K232 likes517 downloads1y agoHugging Face08BananaMind /BananaMind-Base-Bench-1.1gated BananaMind Base Bench 1.1 BananaMind Base Bench 1.1 is an English text-completion benchmark for base causal language models. It contains 350 individually authored examples across seven categories and reports one fixed-scale Overall Elo score. This is not an instruction-following benchmark. Models receive plain text followed by four possible continuations. The official runner selects the continuation with the highest mean conditional token log-probability. It does not use a chat… See the full description on the dataset page: https://huggingface.co/datasets/BananaMind/BananaMind-Base-Bench-1.1.tabulartext-generationn<1K11 likes316 downloads1mo agoHugging Face09JWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes315 downloads2mo agoHugging Face10liodon-ai /nanochat-calendar-arithmetic-base10 nanochat Base-10 Calendar Arithmetic A deterministic, base-10 arithmetic corpus scoped to three cyclic calendar units: hour-of-day (mod 24), day-of-week (mod 7), and month-of-year (mod 12). Companion to Yujivus/nanochat-climbmix-arithmetic-base10, built the same way but scoped to real modular calendar units instead of free-integer add/sub/mul/div/mod. Every example is a single line — question and answer collapsed into one equation, no exposed reasoning: 23:00 + 18965h = 04:00… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/nanochat-calendar-arithmetic-base10.texttext-generation100K<n<1M0 likes268 downloads27d agoHugging Face11nhagar /dclm-baseline-1.0-parquet_urls Dataset Card for dclm-baseline-1.0-parquet_urls This dataset provides the URLs and top-level domains associated with training records in mlfoundations/dclm-baseline-1.0-parquet. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dclm-baseline-1.0-parquet_urls.texttext-generation1B<n<10B0 likes265 downloads1y agoHugging Face12ranjithraj /cancer-knowledge-base Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation The only open CC-BY-4.0 oncology knowledge base that combines: 110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this. A provable 152-question MCQ benchmark — every answer derives from this KB's own structured data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.tabularquestion-answering10K<n<100K0 likes260 downloads1mo agoHugging Face13LLaMAX /BenchMAX_Model-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Model-based is a dataset of BenchMAX, sourcing from m-ArenaHard, which evaluates the instruction following capability via model-based judgment. We extend the original dataset to include languages that are not supported by m-ArenaHard through… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Model-based.texttext-generation1K<n<10K0 likes200 downloads2y agoHugging Face14BaseIntelligence /deepagent DeepAgent Hard, Docker-verifiable software-engineering benchmarks from real merged PRs DeepAgent ships real_pr Harbor hardness packs: live-mined multi-file pull requests, clone@SHA agent images, held-out verifier tests, and Docker dual-truth (solution reward = 1, null reward = 0). Primary product work runs through the deepagent CLI in the GitHub monorepo. Surface Ref Role HF stable pin this dataset revision main Current product on Hub (N=9) HF automation… See the full description on the dataset page: https://huggingface.co/datasets/BaseIntelligence/deepagent.tabulartext-generationn<1K0 likes178 downloads2mo agoHugging Face15RX5950XT /silicon-based-girlfriend-v2-dataset 矽基女友 v2 · 繁中角色扮演合成語料 繁體中文(臺灣)角色扮演的合成對話語料,2,109 筆、38,460 輪、角色輪合計 1,350 萬字元。 訓練出來的模型見 RX5950XT/silicon-based-girlfriend-v2-GGUF。 ⚠️ 全部是模型合成的資料,不是真人對話。 內容包含成人向角色扮演,不適合未成年人。 僅供研究用途。所有角色皆為虛構成年人。 內容 檔案 內容 sharegpt_dataset.json 2,109 筆多輪對話,ShareGPT 格式(id / system / conversations) grpo_prompts.json 648 題 GRPO 用的提示,與 SFT 語料零重疊 holdout_ids.json 100 筆 holdout ID 清單,這些已從 SFT 訓練集排除,供驗收用 general_probes.json 64 題通用能力探針(5 類),用來檢測微調後有無退化… See the full description on the dataset page: https://huggingface.co/datasets/RX5950XT/silicon-based-girlfriend-v2-dataset.texttext-generation1K<n<10K1 likes158 downloads7d agoHugging Face16F555 /qwen3.5-2b-base-blind-spots Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision) Model Tested Field Value Model Qwen/Qwen3.5-2B-Base Parameters 2.27 B (2,274 M per HF metadata) Architecture Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder Type Pre-trained base model (not instruction-tuned) Context 262 144 tokens Modalities Text + Vision (early-fusion multimodal) Key Contributions Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.imagetext-generationn<1K0 likes145 downloads6mo agoHugging Face17jizej /Competence-Based-Evaluation Competence-Based Evaluation (Invariance Benchmark) A benchmark for testing whether language models give the same answer to semantically equivalent reformulations of a logical-ordering question. Given a set of pairwise constraints (e.g. Alice is in front of Bob), a model should answer transitive-closure queries (Is Carol in front of Dave?) consistently whether the constraints are stated using a relation or its inverse. Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.textquestion-answering100K<n<1M0 likes128 downloads5mo agoHugging Face18Basepair /T2T-Centromere-Regulatory T2T Centromere Regulatory Curated and released by Basepair | Follow updates on X: @BasepairSci. Dataset Summary The T2T Centromere Regulatory is the first comprehensive, base-pair resolution mapping of cryptic transcriptional switches and secondary structural elements across the newly sequenced Telomere-to-Telomere (T2T-CHM13 v2.0 / hs1) human centromeres. For decades, centromeric alpha-satellite DNA (~100–200 Mb across human chromosomes) was considered… See the full description on the dataset page: https://huggingface.co/datasets/Basepair/T2T-Centromere-Regulatory.tabulartabular-classification1K<n<10K2 likes94 downloads13d agoHugging Face19Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes85 downloads2y agoHugging Face20disco-jack-basement /byob-pd-book-corpus Dataset Card for BYOB LM on Steroids - Public-Domain Book Training Corpus A multilingual, public-domain corpus of literary, philosophical, and scientific works, published for character-level language-model pre-training. It ships in two forms: raw per-author .txt files (one file per author) and a pre-tokenized, memmap-ready cache (train.bin / val.bin / meta.json) for fast training. Dataset Sources Project Gutenberg (https://www.gutenberg.org) - all tiers Standard… See the full description on the dataset page: https://huggingface.co/datasets/disco-jack-basement/byob-pd-book-corpus.texttext-generation100M<n<1B1 likes85 downloads3mo agoHugging Face21nuhmanpk /dev-knowledge-base Dev Knowledge Base (Programming Documentation Dataset) A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems. Do Follow me on Github: https://github.com/nuhmanpk Overview This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as: Programming languages Frameworks (frontend, backend) DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.tabularquestion-answering100K<n<1M1 likes80 downloads6mo agoHugging Face22liodon-ai /math-dow-mod-synthetic-v1-base6 Math + Cyclic-Time Synthetic Dataset Synthetic dataset for training a small (~10M-100M param), task-specialized LLM on arithmetic (addition, multiplication), cyclic time arithmetic (days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder probes — generalization-focused rather than memorization, following on from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and hours are all instances of the same underlying cyclic/modular-addition structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1-base6.texttext-generation10K<n<100K0 likes80 downloads2mo agoHugging Face23TerenceLau /nanoJEPA-base nanoJEPA EN/ZH Ultra-FineWeb Dataset This is a small pretraining dataset package for nanoJEPA. It is built by streaming openbmb/Ultra-FineWeb split en and/or zh. Files train.jsonl: {"text": "...", "source": "...", "dataset": "...", "language": "en|zh", "score": 0.0} valid.jsonl: same schema as train.jsonl test.jsonl: same schema as train.jsonl Generation Command uv run python data/build_hf_dataset.py \ --out-dir dataset/nanojepa-small \ --languages en,zh… See the full description on the dataset page: https://huggingface.co/datasets/TerenceLau/nanoJEPA-base.texttext-generation1M<n<10M0 likes77 downloads4mo agoHugging Face24liodon-ai /math-dow-mod-synthetic-v1-base7 Math + Cyclic-Time Synthetic Dataset Synthetic dataset for training a small (~10M-100M param), task-specialized LLM on arithmetic (addition, multiplication), cyclic time arithmetic (days-of-week, months, 12-hour and 24-hour clock), and mod-k remainder probes — generalization-focused rather than memorization, following on from the T2/T5/T10 modular-circuit discussion. Days-of-week, months, and hours are all instances of the same underlying cyclic/modular-addition structure… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/math-dow-mod-synthetic-v1-base7.texttext-generation10K<n<100K0 likes76 downloads2mo agoHugging Face25RickyDeSkywalker /GAR_baseDataset GAR Base Dataset This repository contains the base dataset used in the paper GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving. GAR (Generative Adversarial Reinforcement learning) is a comprehensive RL training framework that jointly trains a problem composer and a solver in an adversarial loop. This dataset serves as the starting point for the implicit curriculum learning mechanism, which aligns task difficulty with the prover's evolving capability to… See the full description on the dataset page: https://huggingface.co/datasets/RickyDeSkywalker/GAR_baseDataset.texttext-generation100K<n<1M0 likes75 downloads7mo agoHugging Face26newtype-2038 /pkm-agent-baseline-v2 PKM Agent Baseline — 500 + 50 Scenarios + Six-Grader Artifacts (v2) Two deterministically generated, Korean-language benchmarks for evaluating multi-tool Personal Knowledge Management (PKM) agents over Notion, Gmail, and Google Calendar, plus the Six-Grader Ensemble scoring artifacts (100-scenario reference subset + per-scenario six-metric scores for vanilla and LoRA models). Released alongside the preprint: Vault-Grounded 4B Agent: A Hybrid Reasoning–Fact Architecture for Local… See the full description on the dataset page: https://huggingface.co/datasets/newtype-2038/pkm-agent-baseline-v2.tabulartext-generationn<1K0 likes75 downloads5mo agoHugging Face27karmiq /wikipedia-embeddings-cs-e5-baseThis dataset contains the Czech subset of the wikimedia/wikipedia dataset. Each page is divided into paragraphs, stored as a list in the chunks column. For every paragraph, embeddings are created using the intfloat/multilingual-e5-base model. Usage Load the dataset: from datasets import load_dataset ds = load_dataset("karmiq/wikipedia-embeddings-cs-e5-base", split="train") ds[1] { 'id': '1', 'url': 'https://cs.wikipedia.org/wiki/Astronomie', 'title': 'Astronomie'… See the full description on the dataset page: https://huggingface.co/datasets/karmiq/wikipedia-embeddings-cs-e5-base.texttext-generation100K<n<1M1 likes73 downloads3y agoHugging Face28nuvocare /MSD_manual_topics_user_base MSD_manual_topics_user_base This dataset has been built with the website https://www.msdmanuals.com/ provided by Merck & Co for the greater audience. The MSD manual is an essential source of knowledge for many topics related to symptoms, diseases, health and other related topics. The manual makes an extra effort to make it available both for professionals and patients by having two distinct version. The content, while being labelled the same, differs by the type of user in order to… See the full description on the dataset page: https://huggingface.co/datasets/nuvocare/MSD_manual_topics_user_base.texttext-classification100K<n<1M2 likes67 downloads2y agoHugging Face29Shaer-AI /ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits Ashaar Enhanced Description SFT Stratified Splits Source dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500 Target dataset: Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits This dataset publishes deterministic train / eval / test splits with a 94 / 3 / 3 policy. Split policy Primary stratification key: base_meter form length_bucket Length buckets: 1-3 4-6 7-10 11-20 Small groups fall back… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits.tabulartext-generation100K<n<1M0 likes67 downloads7d agoHugging Face30RickyDeSkywalker /GAR_baseDataset_NuminaMath GAR-Official This is the official repository for the paper GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving. GitHub Repository: RickySkywalker/GAR-Official Trained Models: GAR_Goedel-Prover-V2 GAR_DeepSeek-Prover-V2 Base Datasets: Original base dataset Base dataset under Numina-Math Introduction We introduce GAR: Generative Adversarial Reinforcement Learning, an RL training method that intends to solve inefficiency and suboptimal… See the full description on the dataset page: https://huggingface.co/datasets/RickyDeSkywalker/GAR_baseDataset_NuminaMath.texttext-generation100K<n<1M0 likes63 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.