CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JWei05 /DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for reinforcement-learning experiments. Difficulty is defined by how often the pretrained google/gemma-4-26B-A4B teacher solved each question across eight temperature-1 samples under the same rule-based grader used by the RL training pipeline. The Hub dataset has three configurations—easy, medium, and hard—and each configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.texttext-generation1K<n<10K0 likes341 downloads1mo agoHugging Face02KiteFishAI /arxiv-tex-corpus-mediumarxiv-tex-corpus-medium (15GB) Medium-scale LaTeX corpus from arXiv (math, CS, physics, statistics) 📄 Paper: https://arxiv.org/abs/2602.17288 📚 Overview arxiv-tex-corpus-medium (15GB) is a medium-sized version of the arXiv LaTeX corpus, containing structured LaTeX source content extracted from selected arXiv categories. This dataset is restricted to the following categories: math cs hep-th hep-ph quant-ph stat.ML stat.TH This version (~15GB) is intended for: Research… See the full description on the dataset page: https://huggingface.co/datasets/KiteFishAI/arxiv-tex-corpus-medium.texttext-generation100K<n<1M3 likes309 downloads7mo agoHugging Face03david-thrower /HelixLM-medium-1500.0Mt-2988750pt-20260528 david-thrower/HelixLM-medium-1500.0Mt-2988750pt-20260528 A 1.5 billion token corpus of high quality educational pretraining data Composition: A sample from randomly sampled shards of https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus (99%) Randomly chosen rows from: https://huggingface.co/datasets/open-web-math/open-web-math (1%) Composition based on Kye's (https://huggingface.co/kye) recommendations at… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/HelixLM-medium-1500.0Mt-2988750pt-20260528.texttext-generation1M<n<10M0 likes214 downloads4mo agoHugging Face04Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes206 downloads2y agoHugging Face05shaikat005 /medium-web-pentesting Medium Web Pentesting Articles Dataset Description A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body. This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.tabulartext-classificationn<1K0 likes107 downloads5mo agoHugging Face06Metaskepsis /Olympiads_medium Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 13284 Filtered size: 13240 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_medium.tabulartext-generation10K<n<100K1 likes94 downloads2y agoHugging Face07Metaskepsis /Numina_medium Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 37133 Filtered size: 37133 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina_medium.tabulartext-generation10K<n<100K0 likes81 downloads2y agoHugging Face08BEE-spoke-data /medium-articles-en Dataset Card for "medium-articles-en" fabiochiu/medium-articles filtered for en only and 100 GPT-4 tiktoken tokens or more. texttext-classification100K<n<1M2 likes63 downloads9mo agoHugging Face09devvrit /polaris_filtered_nemotron_medium_math_verifiable Polaris Filtered Nemotron Medium Sympy Verifiable (v2) This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using math verify-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset. Dataset Summary Total Original Samples: 2,424,392 Final Kept Samples: 357,790 (14.8%) Target Reasoning Length: 4k-8k… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_math_verifiable.texttext-generation100K<n<1M0 likes62 downloads9mo agoHugging Face10JWei05 /Nemotron-Math-v2-Medium-10k Nemotron-Math-v2-Medium-10k A lightweight 10,500-problem subset of nvidia/Nemotron-Math-v2 for long-horizon Python-TIR reinforcement learning. It contains 1,500 problems from each metadata.reason_high_with_tool.pass bucket 1 through 7. A deterministic seed-42 shuffle assigns 500 examples to validation and 10,000 to train. This Hugging Face release intentionally contains no teacher traces. The full messages/tools aggregation is retained as a separate local artifact.… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Nemotron-Math-v2-Medium-10k.tabulartext-generation10K<n<100K0 likes61 downloads1mo agoHugging Face11AxeML /MediumSetPT 📚 Dataset de Perguntas e Respostas por Tópico Este repositório contém um dataset com 40.000 amostras estruturadas para tarefas de Processamento de Linguagem Natural (PLN), com foco em perguntas temáticas e respostas desenvolvidas. 📁 Estrutura dos Dados Cada amostra é representada em formato JSON com os seguintes campos: id (string): Identificador único da amostra (UUID). topic (lista de strings): Lista com os tópicos abordados. prompts (lista de strings):… See the full description on the dataset page: https://huggingface.co/datasets/AxeML/MediumSetPT.texttext-generation10K<n<100K4 likes48 downloads1y agoHugging Face12devvrit /polaris_filtered_nemotron_medium_sympy_verifiable Polaris Filtered Nemotron Medium Sympy Verifiable This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using sympy-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset. Dataset Summary Total Original Samples: 2,500,820 Final Kept Samples: 263,123 (10.5%) Target Reasoning Length: Optimized for… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_sympy_verifiable.texttext-generation100K<n<1M0 likes46 downloads9mo agoHugging Face13maximuspowers /muat-pca-10-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods pca Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 8 to 10 Neurons per… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-pca-10-medium.texttext-generation10K<n<100K0 likes45 downloads10mo agoHugging Face14Aipresso /medium_512_1k_tokens_prompts Medium 512-1K Tokens Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ By using this dataset you agree to our Terms of Use. Overview 703 high-quality English prompts whose length lies between 512 and 1 000 tokens.Every prompt has been de-duplicated, cleaned and token-counted with the GPT-2 tokenizer. Statistics Rows Token range File size Format 703 512 – 1 000 2.9 MB CSV Use-cases Medium-context language-model fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/medium_512_1k_tokens_prompts.texttext-generationn<1K0 likes44 downloads11mo agoHugging Face15MicPie /unpredictable_rated-mediumThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes43 downloads4y agoHugging Face16crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes32 downloads1y agoHugging Face17hudsongouge /mediumish-small-agent-sft-preview-v0.1 mediumish-small-agent-sft-preview-v0.1 Procedurally generated ChatML SFT data for medium/small models, covering agentic tool use, anti-hallucination habits, grounded refusal, multi-step reasoning, and related epistemic behaviors. This build replaces the earlier 5,000-row preview. 6,976 rows, stratified across 20 domains (reliability/tool-use, bible study, hidden-assumption reasoning, advanced math, code repair against a documented spec, rulebook/policy simulation, ARC-style grid… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/mediumish-small-agent-sft-preview-v0.1.texttext-generation1K<n<10K1 likes31 downloads2mo agoHugging Face18yuhan-nlp /collabllm-medium-rl-grpo CollabLLM medium — RL (GRPO) train/validation split Inputs for GRPO training on the CollabLLM medium document-writing task, as used to produce the verl-grpo-medium-qwen3-4b-step{50,100,129} checkpoints. file rows size rl_train.parquet 2072 16M rl_validation.parquet — 2.2M Derived from the CollabLLM medium task (source articles: Kamaljp/medium_articles). Each row is a single-turn prompt that the training loop expands into a multi-turn conversation with a… See the full description on the dataset page: https://huggingface.co/datasets/yuhan-nlp/collabllm-medium-rl-grpo.texttext-generation1K<n<10K0 likes31 downloads2mo agoHugging Face19hadeelbkh /tokenized-IELTS-writing-task-2-evaluation-DialoGPT-mediumtexttext-generation1K<n<10K2 likes26 downloads1y agoHugging Face20BEE-spoke-data /falcon-refinedweb-1M_en_medium BEE-spoke-data/falcon-refinedweb-1M_en_medium A sample from falcon-refinedweb: more than 512 & less than 8192 gpt4 tiktoken tokens en only (via fasttext-langdetect) 1M samples GPT-4 tiktoken token count: token_count count 1000000.000000 mean 1197.179246 std 964.177338 min 513.000000 25% 653.000000 50% 871.000000 75% 1315.000000 max 8191.000000 Total count: 1197.18 M tokens texttext-generation1M<n<10M2 likes25 downloads9mo agoHugging Face21hudsongouge /mediumish-small-agent-sft-v3 mediumish-small-agent-sft-v3 50000 ChatML SFT rows assembled for a same-day training run. Field Value Release channel deadline_candidate production_sft false (Phase F / council not claimed) Format single column chatml (full multi-turn + tool traces) Tool-bearing rows 6545 Generated 2026-07-19 Source mix Track Rows curriculum 36864 reasoning_policy 5000 truth_seeker 4997 habit_lock 2500 agent_gym_live 639… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/mediumish-small-agent-sft-v3.texttext-generation10K<n<100K0 likes25 downloads2mo agoHugging Face22dzur658 /ping-technical-assistant-mediumNow 3x the size of Ping Technical Assitant Small! NOTE: A new LoRA will be trained on this data soon! Ping Technical Assistant Dataset Small This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-medium.texttext-generation1K<n<10K0 likes24 downloads7mo agoHugging Face23Simon-Liu /gemma-270m-medium-qa本資料集包含由 ** gemini-2.0-flash ** 生成的對話資料,採用 OpenAI Chat Messages 格式(.jsonl)。資料來源結合: Reference-free:由 seed 派生的單輪問答。 Reference-based:依據參考文本生成單輪問答。 檔案路徑:data/train.jsonl(選配:data/train.parquet) 結構說明 每列為一筆樣本:{"id": "...", "type": "...", "seed": "...", "context": "...", "messages": [{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]} type 欄位標示資料來源:reference_free 或 reference_based。 seed 欄位儲存 Reference-free 的原始 seed 指令,或 Reference-based 的參考文本片段。 context 欄位僅在… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/gemma-270m-medium-qa.texttext-generationn<1K0 likes22 downloads1y agoHugging Face24maximuspowers /muat-mean-std-fourier-5-pca-10-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std, pca, fourier Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-fourier-5-pca-10-medium.texttext-generation10K<n<100K0 likes22 downloads10mo agoHugging Face25maximuspowers /muat-fourier-5-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods fourier Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 6 to 8 Neurons… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-fourier-5-medium.texttext-generation10K<n<100K0 likes19 downloads10mo agoHugging Face26Jongbin-kr /VeriReason-reasoning-reproduced-1513_luna-medium VeriReason reasoning reproduced (Luna medium) This dataset is a deterministic sample of the locally updated reproduced VeriReason files. Source files: train (1).jsonl, validation (1).jsonl Sampling: independent shuffle then prefix slice, fixed seed 42 Splits: train 1513, validation 189 Original target dataset: Jongbin-kr/VeriReason-RTL-Coder_7b_reasoning_tb (train 1513 / validation 189) Schema: id, instruction, output, tb, tb_result Quality checks: valid JSON, unique IDs… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/VeriReason-reasoning-reproduced-1513_luna-medium.texttext-generation1K<n<10K0 likes18 downloads12h agoHugging Face27maximuspowers /muat-mean-std-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 6 to 8… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-medium.texttext-generation10K<n<100K0 likes17 downloads10mo agoHugging Face28JingweiNi /ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513 GPT-5.5 Medium Reannotation of Qwen3.5-Positive OCR2 Coding Steps This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and contains GPT-5.5 medium-reasoning reannotations for the 1,536 Qwen3.5-positive error steps. Summary Source dataset: JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 Source rows: 500 K2-Think Codeforces traces Source manifest-selected Qwen3.5 labels: 10,000 steps GPT-5.5 reannotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513.tabulartext-generationn<1K0 likes17 downloads4mo agoHugging Face29cs-giung /nemotron-math-v2-medium-mini Nemotron Math V2 Medium Mini A compact textual reasoning dataset derived from the medium split of nvidia/Nemotron-Math-v2 at immutable revision 8e793210e175b6406c752a870f585f62de98c0d3. Selection boundary This extraction intentionally excludes tool-use semantics: Accept exactly two messages with roles user then assistant. Reject records declaring tools, containing assistant tool calls, containing tool-result messages, or containing any additional turns. Require… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/nemotron-math-v2-medium-mini.texttext-generation10K<n<100K0 likes13 downloads1mo agoHugging Face30joppari /mn_business_benchmark_dataset_medium mn_business_benchmark_dataset_10000_diverse Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics, стратегийн 10000 мөртэй синтетик benchmark dataset. Schema id: 1-ээс 10000 хүртэлх дараалсан дугаар instruction: бизнесийн бодлогын өгүүлбэр input: хоосон string thinking: бодолт, томьёо, завсрын алхам output: эцсийн хариу topic: бизнесийн сэдэв difficulty: easy эсвэл medium image_svg: тухайн бодлогын энгийн SVG card дүрслэл Generated deterministically by… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_medium.texttext-generation10K<n<100K0 likes12 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.