CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iiis-lean /NuminaMath-LEAN-Sol NuminaMath-LEAN Cleaned with NL Solutions Dataset Summary This is a cleaned version of the NuminaMath-LEAN dataset, enhanced with natural language (NL) solutions matched from source datasets. The primary goal is to provide paired formal statements/proofs with natural language solutions for proof formalization and theorem proving research. The dataset matches problems from NuminaMath-LEAN with their corresponding natural language solutions from: olympiads-ref: A… See the full description on the dataset page: https://huggingface.co/datasets/iiis-lean/NuminaMath-LEAN-Sol.texttext-generation10K<n<100K0 likes845 downloads8mo agoHugging Face02dougalldeepmind /2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment. field value experiment LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.textn<1K0 likes360 downloads23d agoHugging Face03dougalldeepmind /2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture Qwen3.6-27B training bundle — 2026-08-04-qwen36-27b-1000ex-da350-numina650-train code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_1000ex_da350_numina650.yaml. field value experiment 1000-example mixture: 350 difficult-advice (all t1-t3 + t4 fill) + 650 NuminaMath-CoT; lr 4e-5, 1 epoch date_generated 2026-08-03 constitution constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture.text1K<n<10K1 likes140 downloads29d agoHugging Face04dougalldeepmind /2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers 499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5. Source Examples Tokens Share Marker NuminaMath-CoT 611 333,351 66.9% no No Robots 271 82,239 16.5% yes TULU3 119 82,445 16.5% yes Total 1,001 499,595 390 marked Derived from qwen3.6-27b-mixture-500k-numina-heavy by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.texttext-generation1K<n<10K0 likes139 downloads23d agoHugging Face05nyu-dice-lab /lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private Dataset Card for Evaluation run of penfever/Llama-3-8B-NuminaCoT Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-NuminaCoT The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private.tabular100K<n<1M0 likes131 downloads2y agoHugging Face06dougalldeepmind /2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture Qwen3.6-27B training bundle — 1,000 examples (250 synthdoc_v2 t1-t3 + 750 NuminaMath-CoT) RunPod training bundle: code.tar.gz (the trainer, src/, configs/) plus mixture.jsonl. The pod pulls this, untars it, copies the jsonl to data/, and runs configs/train_1000ex_da250_numina750.yaml. field value experiment 1-epoch assistant-only-loss LoRA SFT of Qwen3.6-27B on a 1,000-example mixture that is 25% difficult-advice by example count date_generated 2026-08-03… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture.text1K<n<10K0 likes120 downloads29d agoHugging Face07dougalldeepmind /2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armA-1000ex-numina666-tulu334-train code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies the jsonl to data/mixture.jsonl, and runs configs/train_armA_1000ex_numina666_tulu334.yaml. field value experiment Arm A: 1,000 examples, no difficult-advice - 666 NuminaMath-CoT + 167 TULU3 + 167 No Robots date_generated 2026-08-03 constitution constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-arma-1000ex-numina-666-tulu-334-train-mixture.text1K<n<10K0 likes109 downloads29d agoHugging Face08tlb-22 /euclean-numina-geometry Euclean Numina-Geometry Euclean Numina-Geometry is a generated Lean 4 / Mathlib geometry formalization dataset released with the ICML 2026 paper: Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean GitHub repository: https://github.com/tlb-22/Euclean Dataset Description This release contains 183,796 Numina-derived geometry problems with generated Lean theorem statements. The formalizations were regenerated with Codex GPT-5.4 using… See the full description on the dataset page: https://huggingface.co/datasets/tlb-22/euclean-numina-geometry.text100K<n<1M1 likes103 downloads3mo agoHugging Face09pateltakshm /2000-Reasoning-AIME-Numinamathtext10K<n<100K0 likes84 downloads9mo agoHugging Face10ubowang /qwen_math_numina_80k_add_critique_0119tabular10K<n<100K0 likes78 downloads2y agoHugging Face11Azure99 /NuminaMath-1.5-Pro NuminaMath-1.5-Pro Dataset Overview NuminaMath-1.5-Pro targets post-training and verifiable reasoning scenarios. It applies strict filtering, judge-based consistency checks, and staged solution regeneration on top of the upstream NuminaMath-1.5 dataset. All data processing and synthesis for this dataset is executed with the BlossomData framework, covering the full pipeline—loading, filtering, judging, generation, retry, and export—with an emphasis on reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/Azure99/NuminaMath-1.5-Pro.texttext-generation10K<n<100K1 likes74 downloads1y agoHugging Face12ChuGyouk /AI-MO-NuminaMath-TIR-korean-240918 IMPORTANT NOTE This data is part of the progress. Current translation progress: 24.85% (2024-09-18 01:32 KST) I'm taking a short break due to personal reasons. I'll be back in a month. TODO-LIST Finish translation Translation I used gemini-1.5-pro-exp-0827. The prompt used for translation will be disclosed at the end. Dataset Card for NuminaMath CoT Dataset Summary Tool-integrated reasoning (TIR) plays a crucial role in this… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI-MO-NuminaMath-TIR-korean-240918.texttext-generation10K<n<100K5 likes72 downloads2y agoHugging Face13Lyric1010 /numina-0.2B-drop-entropy50-merge Dataset: numina-0.2B-drop-entropy50-merge This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_qwen3/numina-0.2B-drop-entropy50-merge/stage_1. textn<1K0 likes69 downloads10mo agoHugging Face14RickyDeSkywalker /GAR_baseDataset_NuminaMath GAR-Official This is the official repository for the paper GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving. GitHub Repository: RickySkywalker/GAR-Official Trained Models: GAR_Goedel-Prover-V2 GAR_DeepSeek-Prover-V2 Base Datasets: Original base dataset Base dataset under Numina-Math Introduction We introduce GAR: Generative Adversarial Reinforcement Learning, an RL training method that intends to solve inefficiency and suboptimal… See the full description on the dataset page: https://huggingface.co/datasets/RickyDeSkywalker/GAR_baseDataset_NuminaMath.texttext-generation100K<n<1M0 likes63 downloads7mo agoHugging Face15Lyric1010 /numina-0.02B-drop-merge-llama Dataset: numina-0.02B-drop-merge-llama This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/numina-0.02B-drop-merge/stage_1. textn<1K0 likes62 downloads10mo agoHugging Face16codezakh /NuminaMath-1.5-EFA-Subset📃 Paper This dataset contains EFAs inferred for a subset of NuminaMath_CoT, specifically the first 5,000 problems. These EFAs were inferred by this model, and the prompts used for training are linked in the model card. The dataset contains multiple EFA candidates for most of the first 5,000 problems in NuminaMath. Each row in the dataset is described by the Row class below: from pydantic import BaseModel class ProblemVariant(BaseModel): """Synthetic problem variants constructed by… See the full description on the dataset page: https://huggingface.co/datasets/codezakh/NuminaMath-1.5-EFA-Subset.texttext-generation10K<n<100K1 likes51 downloads1y agoHugging Face17NovaSky-AI /labeled_numina_difficultyWe also include data of labeling difficulty from NUMINA, in the following files: labeled_amc_aime_0_-1.json, labeled_math_0_-1.json, labeled_olympiads_0_-1.json. text100K<n<1M6 likes50 downloads2y agoHugging Face18Hoglet-33 /NuminaMath-500k500,000 samples from the AI-MO/NuminaMath-CoT dataset text100K<n<1M0 likes41 downloads6mo agoHugging Face19flatlander1024 /NuminaMath-CoT-filteredDataset that contains problems that appears in both QwQ-LongCoT-130K-cleaned and NuminaMath-CoT. There are approximately 100k problems where the solution is in plain-CoT manner. textquestion-answering100K<n<1M1 likes40 downloads2y agoHugging Face20LangAGI-Lab /train-rl-o1-mini-annotated-math-numina-22ktabular10K<n<100K1 likes37 downloads2y agoHugging Face21codex-master /openthoughts3_numinamath-1.5-pro_mixturetexttext-generation10K<n<100K0 likes29 downloads1mo agoHugging Face22codex-master /numina_smoltalk_mixturetexttext-generation100K<n<1M0 likes28 downloads2mo agoHugging Face23HanningZhang /Llama3.1-100K-N16-Numina-ORMtext1M<n<10M0 likes23 downloads2y agoHugging Face24rqzhang /NuminaMath-processed-instructtext100K<n<1M0 likes22 downloads1y agoHugging Face25llm-compe-2025-kato /numina-math-to-gsm8k numina-math-to-gsm8k 急遽必要になったgsm8k形式データ AI-MO/NuminaMath-1.5 (Apache 2.0)のフォーマット変形版 オリンピック以上のサブセットかつ、検証された問題及び解答かつ、合成していないもののみ使用 英語以外の言語がごく僅か混じっていたので取り除き \Boxed{} の回答形式が400件程度混じっていたのでまとめて取り除き text10K<n<100K0 likes17 downloads1y agoHugging Face26HanningZhang /Llama3-200K-N16-Numina-ORMtext1M<n<10M0 likes14 downloads2y agoHugging Face27LLMTeamAkiyama /cleaned_NuminaMath-RL-Verifiable_with_prooftabular100K<n<1M0 likes13 downloads1y agoHugging Face28open-llm-leaderboard /AI-MO__NuminaMath-7B-TIR-detailsgated Dataset Card for Evaluation run of AI-MO/NuminaMath-7B-TIR Dataset automatically created during the evaluation run of model AI-MO/NuminaMath-7B-TIR The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-MO__NuminaMath-7B-TIR-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face29HanningZhang /Numina-ORM-650Ktext100K<n<1M0 likes12 downloads2y agoHugging Face30FUfu99 /NuminaMath-CoT-Small-Hard-200ktext100K<n<1M0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.