CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SimVer-ano /simverse2026 SimVerse ⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review. A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.imagevisual-question-answering1K<n<10K0 likes472 downloads5mo agoHugging Face02while-ai /agent-simulations Agent Simulations Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets 53,971 synthetic agent trajectories generated by simulations across 34 agent types. The rows include successful and failed trajectories for supervised fine-tuning, preference work, reinforcement learning, and evaluation. NOTE: This is generated test and training data, not curated ground truth. Review and filter it for your application before training or… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.texttext-generation10K<n<100K0 likes388 downloads1d agoHugging Face03simutrade /simutrade-rag-sft-28k 📢 Domain & Email Migration Notice From May 6th, 2026, Simutrade will transition to new domains as simutrade.app will not be renewed: 🌐 Website: simutrade.faizath.com (formerly simutrade.app) ⚙️ API: simutrade-api.faizath.com (formerly api.simutrade.app) 📧 Email: contact@simutrade.faizath.com (formerly contact@simutrade.app) 🛰️ CDN: simutrade-cdn.faizath.com (formerly cdn.simutrade.app) 📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/simutrade/simutrade-rag-sft-28k.textquestion-answering10K<n<100K1 likes216 downloads1mo agoHugging Face04while-ai /tau2-simulated tau2 Simulated Training Set Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets The training set that took a base model from 5% to 30% on tau2-bench telecom, made from nothing but the agent's tool list and policy. If you build a customer-facing agent, you already have the two files this dataset was made from: the tools it can call and the policy it follows. The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.texttext-generation1K<n<10K0 likes133 downloads1d agoHugging Face05Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb_simple Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.texttext-generationn<1K0 likes84 downloads1y agoHugging Face06hasankursun /age-specific-text-simplification Age-Specific Text Simplification Dataset Dataset Description This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group. Dataset Summary Total Examples: 17,177 Training Split: 15,459 examples Validation Split: 1,718 examples Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.tabulartext-generation10K<n<100K3 likes80 downloads1y agoHugging Face07alexfromapex /simplemath-cot 🧮 SimpleMath-100k CoT A chain-of-thought (CoT) extension of the ProCreations/SimpleMath dataset. Every one of the 100 000 algebra / arithmetic problems is paired with a short, numbered reasoning trace (Step 1: … Step 2: …) that walks a language model from the problem statement to the known-correct answer. The traces in the Jupyter notebook are generated by Qwen3.8-27B and then post-processed to strip formatting noise, enforce sequential step numbering, and cap output at 1 000… See the full description on the dataset page: https://huggingface.co/datasets/alexfromapex/simplemath-cot.texttext-generationn<1K0 likes65 downloads19d agoHugging Face08ProCreations /simple-facts Simple Facts A dataset of simple, no BS, human collected, ethicly sourced facts. About 1000 examples. This dataset is growing, and every day I plan to add a few more facts. texttext-generation1K<n<10K4 likes62 downloads1y agoHugging Face09Simon-Liu /tw-finance-reasoning-instruct tw-finance-reasoning-instruct 台灣金融知識的繁體中文推理指令資料集。每一題都有完整的思考過程(think)與可查證的答案(output)。 欄位規格對齊 twinkle-ai/tw-reasoning-instruct-50k。 ⚠️ 使用限制:僅供研究,不得商業使用 本資料集以 CC BY-NC 4.0 授權釋出,僅供學術研究、模型能力探索與方法驗證之用。 不得作商業用途。 包含訓練用於對外營利的模型、包裝為付費產品或服務、 或作為商業交付物的一部分。若有商業需求,請自行重新建置資料並取得合規來源。 這不是財務、稅務、法律或投資建議。 資料中的稅率、費率、法規門檻依 2026 年 (民國 115 年)台灣公開資訊整理,但可能已經過時。任何實際決策前, 請以主管機關公告為準(財政部、金管會、勞動部、衛福部、中央銀行、全國法規資料庫)。 內容為程式化合成,非真實考題逐字收錄。 題目由計算器與模板生成, 並非任何證照考試的原始試題。… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/tw-finance-reasoning-instruct.texttext-generation1K<n<10K0 likes60 downloads15d agoHugging Face10rpisano /nemotron-cc-atomic-simplification-gemma4-31b nemotron-cc atomic-statement simplification (Gemma 4 31B-it) 2,000,000 records: source text from nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object statements. Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to <=8192 templated tokens. Fields id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.texttext-generation1M<n<10M0 likes59 downloads15d agoHugging Face11Simon-Liu /tw-finance-function-call-reasoning tw-finance-function-call-reasoning 台灣金融場景的繁體中文 function-calling + 推理鏈微調資料集。 欄位規格對齊 twinkle-ai/tw-function-call-reasoning-10k。 ⚠️ 使用限制:僅供研究,不得商業使用 本資料集以 CC BY-NC 4.0 授權釋出,僅供學術研究、模型能力探索與方法驗證之用。 請務必理解以下事項後再使用: 不得作商業用途。 包含但不限於:訓練用於對外營利的模型、包裝為付費產品或服務、 作為商業交付物的一部分。若有商業需求,請自行重新建置資料並取得合規來源。 這不是財務、稅務、法律或投資建議。 資料中的稅率、費率、法規門檻雖依 2026 年 (民國 115 年)台灣公開資訊整理,但可能已經過時或有誤。任何實際決策前, 請以主管機關公告為準(財政部、金管會、勞動部、衛福部、中央銀行、全國法規資料庫)。 內容為程式化合成,非真實考題逐字收錄。 題目由模板與參數取樣組合而成,… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/tw-finance-function-call-reasoning.texttext-generation1K<n<10K0 likes59 downloads15d agoHugging Face12Papersnake /ACG-SimpleQA ACG-SimpleQA 🌐 Website • 🤗 Hugging Face 中文 | English ACG-SimpleQA is an objective knowledge question-answering dataset focused on the Chinese ACG (Animation, Comic, Game) domain, containing 4242 auto-generated carefully designed QA samples. This benchmark aims to evaluate large language models' factual capabilities in the ACG culture domain, featuring Chinese language, diversity, high quality, static answers, and easy evaluation. 📢 Latest Updates… See the full description on the dataset page: https://huggingface.co/datasets/Papersnake/ACG-SimpleQA.texttext-generation1K<n<10K2 likes51 downloads1y agoHugging Face13Simon-Liu /kubectl-mcp-server-tool-call-reasoning-6k kubectl-mcp-server-tool-call-reasoning-6k MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。 語言:繁體中文 工具(來自 MCP server):install_helm_chart, upgrade_helm_chart, uninstall_helm_chart, helm_list, helm_status, helm_history, helm_get_values, helm_get_manifest, helm_get_notes, helm_get_hooks, helm_get_all, helm_show_chart, helm_show_values, helm_show_readme, helm_show_crds, helm_show_all, helm_search_repo, helm_search_hub, helm_repo_list… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/kubectl-mcp-server-tool-call-reasoning-6k.texttext-generation1K<n<10K0 likes50 downloads2mo agoHugging Face14thisisandreeeee /simple-llm-sft Simple LLM SFT Dataset This synthetic dataset contains 1,000 English prompt-response pairs for supervised fine-tuning. It was created to fine-tune Qwen/Qwen3.5-4B to give clear, direct, and technically correct answers in simple English. The writing guidance is inspired by ASD-STE100 Simplified Technical English. The dataset does not claim official ASD-STE100 compliance or certification. Dataset structure The default configuration contains: Split Examples… See the full description on the dataset page: https://huggingface.co/datasets/thisisandreeeee/simple-llm-sft.texttext-generation1K<n<10K0 likes45 downloads10d agoHugging Face15UWV /Leesplank_NL_wikipedia_simplificationsThe set contains 2.87M pragraphs of prompt/result combinations, where the prompt is a paragraph from Dutch Wikipedia and the result is a simplified text, which could include more than one paragraph. This dataset was created by UWV, as a part of project "Leesplank", an effort to generate datasets that are ethically and legally sound. The basis of this dataset was the wikipedia extract as a part of Gigacorpus (http://gigacorpus.nl/). The lines were fed one by one into GPT 4 1106 preview, where… See the full description on the dataset page: https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications.texttext-generation1M<n<10M6 likes42 downloads3y agoHugging Face16sapiens-technology /simple_bench 📊 Simple Bench Dataset A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.texttext-generationn<1K0 likes38 downloads5mo agoHugging Face17cahya /simple-wikipedia English Simple Wikipedia This is just a copy of english simple Wikipedia dataset that I converted to jsonl format for testing purpose when jsonl format is needed. Here is the link to download the jsonl file. texttext-generation100K<n<1M0 likes35 downloads2y agoHugging Face18Jotschi /visual_genome-simple-en Dataset Card for Visual Genome Annotations in Simple English This dataset contains captions that were rephrased into simple english so that a young child would understand it. Dataset Details Dataset Sources The processed Visual Genome captions in this repo are based on the following sources: 941425b651f50cdb1a6f0673eaab6260 vg_caption.json (https://storage.googleapis.com/sfr-vision-language-research/LAVIS/datasets/visual_genome/vg_caption.json) Visual… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/visual_genome-simple-en.texttext-generation100K<n<1M0 likes27 downloads3y agoHugging Face19Simonc-44 /Cygnis-Identity-SFT Cygnis Identity Dataset Présentation Ce dépôt contient le jeu de données d'entraînement initial pour l'identité de l'intelligence artificielle Cygnis. Ce dataset est conçu pour le réglage fin (fine-tuning) supervisé afin d'établir les fondements comportementaux et l'identité de l'agent. Spécifications du Dataset Le jeu de données est composé de paires d'instructions visant à définir l'origine, le concepteur et la nature du système. Format : JSONL / Hugging… See the full description on the dataset page: https://huggingface.co/datasets/Simonc-44/Cygnis-Identity-SFT.texttext-generationn<1K0 likes27 downloads6mo agoHugging Face20sagecodes /simple-python-grpo simple-python-grpo A curated set of simple Python function problems for GRPO / RLVR fine-tuning. Each row has a natural-language description, a function signature, and 3 auto-verified test assertions (generated by running a reference implementation, so every test is correct by construction). The reference is NOT included — the model must generate the body and is rewarded when the tests pass. Fields: name, prompt, func_prompt, tests (newline-separated asserts), setup_code. Built… See the full description on the dataset page: https://huggingface.co/datasets/sagecodes/simple-python-grpo.texttext-generationn<1K0 likes27 downloads3mo agoHugging Face21LiteMind /Simple-agent-traces 📱 Simple Agent Traces – Tiny Tool‑Calling Conversations for Small Models Simple Agent Traces is a compact, hand‑picked dataset of 605 real‑world tool‑calling conversations, each carefully truncated to ≤8,192 tokens (using the SmolLM2‑360M tokenizer).It is purpose‑built for training and fine‑tuning tiny language models (≤500M) that must run on‑device – smartphones, edge devices, or any environment with strict memory and latency constraints. 🧹 No chain‑of‑thought, no fluff.Every… See the full description on the dataset page: https://huggingface.co/datasets/LiteMind/Simple-agent-traces.tabulartext-generationn<1K3 likes26 downloads4mo agoHugging Face22Ibisbill /Semantic_similarity_deduplicated_reasoning_data_english Semantic_similarity_deduplicated_reasoning_data_english 数据集描述 Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category 文件结构 semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式) 数据格式 数据集包含以下字段: question: str quality: int difficulty: int topic: str validity: int 使用方法 方法1: 使用datasets库 from datasets import load_dataset #… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face23simulatexp /SIMXP-26052026-METASYN001 SIMXP-26052026-METASYN001 Multi-Omics Agent Memory Simulation — Metabolic Syndrome TCA Cycle Biomarker Study This dataset supports the experiment described in the article "Does Your Research Agent Remember? Six Months of Multi-Omics Team Knowledge vs. None — A Controlled Comparison" and demonstrates the etchmem memory system for autonomous AI research agents. It contains the full event log, synthesized knowledge export, and fine-tuning pairs from a simulated six-month plasma… See the full description on the dataset page: https://huggingface.co/datasets/simulatexp/SIMXP-26052026-METASYN001.texttext-generationn<1K0 likes22 downloads4mo agoHugging Face24simpissa /countdown-qwen3-0.6b Countdown Qwen3-0.6B Pass@10 Buckets Countdown arithmetic problems filtered by observed local Qwen/Qwen3-0.6B success rate over 10 rollouts per problem. Each problem asks for an arithmetic expression that reaches a target using each listed source number at most once. The final answer should be inside \boxed{...}. Canonical solutions are provided, but any verifier-valid expression is accepted. Subsets subset source bucket count observed successes out of 10… See the full description on the dataset page: https://huggingface.co/datasets/simpissa/countdown-qwen3-0.6b.tabulartext-generation1K<n<10K0 likes22 downloads4mo agoHugging Face25Simon-Liu /twinkle_hub_finetune_dataset twinkle_hub_finetune_dataset MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。 語言:繁體中文 工具(來自 MCP server):search_datasets, get_dataset, query_rows, materialize_dataset, search_patents, get_patent_body, search_exam, search_exam_questions, get_exam_paper, search_teacher_exam, search_teacher_exam_questions, get_teacher_exam_paper, search_teacher_recruit, search_teacher_recruit_questions, get_teacher_recruit_paper, search_taiwan_md… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/twinkle_hub_finetune_dataset.texttext-generation1K<n<10K1 likes22 downloads2mo agoHugging Face26SIMBA9657 /haddas-instruct-ti haddas-instruct-ti Instruction-tuning pairs synthesized from real Eritrean newspaper articles: summarize, translate, classify topic, extract keywords. Inputs are Tigrinya, outputs are English (or topic label). Source Derived from the Haddas Eritrea newspaper archive: 63 PDF issues processed by the haddas-eritrea pipeline (extract -> clean -> segment -> translate -> label). Generated: 2026-04-26 12:21 UTC Row count: 6720 Schema: id, task, instruction, input, output, topic… See the full description on the dataset page: https://huggingface.co/datasets/SIMBA9657/haddas-instruct-ti.texttext-generation1K<n<10K0 likes21 downloads5mo agoHugging Face27Sadou /medical-reports-simplification-dataset 🏥 Medical Reports Simplification Dataset 📋 Description Dataset créé avec Gemini 2.5 Pro (Preview) pour entraîner des modèles à simplifier les rapports médicaux complexes en explications compréhensibles pour les patients. 🎯 Objectif : Démocratiser l'accès à l'information médicale en rendant les rapports techniques accessibles au grand public. 🔧 Génération du Dataset Génération : Gemini 2.5 Pro (Preview) Validation : Contrôle qualité automatisé… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medical-reports-simplification-dataset.texttext-generationn<1K0 likes19 downloads1y agoHugging Face28joppari /mn_business_benchmark_dataset_simple mn_business_benchmark_dataset_2000_diverse Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset. Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан. Schema id: 1-ээс 2000 хүртэлх дараалсан дугаар instruction: бизнесийн бодлогын өгүүлбэр input: хоосон string thinking: бодолт… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_simple.texttext-generation1K<n<10K1 likes16 downloads4mo agoHugging Face29jmp1987 /Simon_Trove SimsonTrove – Racing-Planet Agentic Traces Synthetischer ReAct-Datensatz (Reason + Act) zur Zweitakt-Tuningberatung auf Basis des Racing-Planet.de Sortiments. Jeder Trace bildet die ingenieurmäßige Entscheidungs- kette eines Tuning-Mechanikers ab: Beobachtung → Hypothese → Action (Teilewahl) → Feedback → Iteration → Final Response. Schnellüberblick Property Value Anzahl Traces 300 Sprache Deutsch (technisch) Format JSON (Liste von Trace-Objekten) Avg.… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/Simon_Trove.texttext-generationn<1K2 likes13 downloads4mo agoHugging Face30yugi5 /simple-text-generation-basic Simple Text Generation Basic Dataset This dataset contains very simple text samples designed for testing and basic text generation tasks. Dataset Structure Each row contains a single field: text: a plain English sentence Example {"text": "Artificial intelligence is transforming the world."} texttext-generationn<1K0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.