CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simplescaling /aime24_nofiguresThe 30 problems from AIME 2024 only with the ASY code for figures when it is necessary to solve the problem. Figure code that is not core to the problem was excluded. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime24_nofigures.textn<1K2 likes13k downloads1y agoHugging Face02OpenStellarTeam /Chinese-SimpleQA Overview 🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper • 📊 Leaderboard Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, our benchmark covers 6 major topics with 99 diverse subtopics. Please visit our website or check our paper for more details.… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SimpleQA.textquestion-answering1K<n<10K38 likes2.6k downloads2y agoHugging Face03allenai /SimpleToM SimpleToM Dataset and Evaluation data The SimpleToM dataset of stories with associated questions are described in the paper "SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs" Associated evaluation data for the models analyzed in the paper can be found in the separate dataset: SimpleToM-eval-data. Question sets There are three question sets in the SimpleToM dataset: mental-state-qa questions about information awareness… See the full description on the dataset page: https://huggingface.co/datasets/allenai/SimpleToM.text1K<n<10K11 likes2.3k downloads7mo agoHugging Face04simplelex /ATO-Australian-Tax-Rulings-and-Guidance ATO Rulings & Guidance — Australian Tax Law, Structured for AI 67,000+ Australian Taxation Office documents as RAG-ready NDJSON/CSV — Edited Private Advice, public rulings and determinations, ATO Interpretative Decisions, practical compliance guidelines, taxpayer alerts, decision impact statements, practice statements and legislative instruments. Every document parsed into structured, typed fields for legal RAG, LLM fine-tuning, and tax research automation. Machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/simplelex/ATO-Australian-Tax-Rulings-and-Guidance.text10K<n<100K1 likes1.7k downloads1d agoHugging Face05simplescaling /aime25_nofigurestextn<1K0 likes408 downloads2y agoHugging Face06SimpleFunctions /sf-index-history SimpleFunctions Index History Time series of the SF Index: a four-number summary of prediction-market consensus — disagreement (0-100), geo-risk (0-100), breadth (-1..+1), and activity (0-100) — computed every 15 minutes from ~50K markets. Flat JSONL for easy charting / analysis. License and Use This dataset is released under Creative Commons Attribution 4.0 International (CC-BY-4.0; https://creativecommons.org/licenses/by/4.0/). You may use it freely for personal… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/sf-index-history.tabular10K<n<100K0 likes407 downloads9h agoHugging Face07simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes384 downloads2d agoHugging Face08simplescaling /aime_nofiguresThe 90 problems from AIME 2022, 2023, 2024 only with the ASY code for figures when it is necessary to solve the problem. Figure code that is not core to the problem was excluded. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime_nofigures.textn<1K1 likes349 downloads2y agoHugging Face09simpleG2023 /chinese-clean-energy-battery-open-intelligence 🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.tabulartext-retrieval1K<n<10K0 likes278 downloads2d agoHugging Face10simpleG2023 /chinese-ai-and-robotics-open-intelligence 🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes256 downloads2d agoHugging Face11simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes240 downloads2d agoHugging Face12deep-analysis-research /simple-evalstext100K<n<1M0 likes223 downloads10mo agoHugging Face13alibaba-pai /SimpleQA-Bench SimpleQA-Bench Tags: factuality, EN, ZH, short-form-answer, human-label Copyright: © 2024 alibaba-pai Source.OpenAI's SimpleQA: Blog & Paper / Data & simple-evals ProjectOpenStellarTeam's Chinese-SimpleQA: Blog & Paper, Data@HF Factuality is a complicated topic because it is hard to measure—evaluating the factuality of any given arbitrary claim is challenging, and language models can generate long completions that contain dozens of factual claims. In SimpleQA, we will focus on… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/SimpleQA-Bench.text1K<n<10K3 likes178 downloads2y agoHugging Face14simplescaling /openaimathThe 500 problems from MATH that were used in "Let's verify step-by-step" and "s1: Simple test-time scaling" are the test set; rest is regular train set. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/openaimath.text10K<n<100K5 likes174 downloads1y agoHugging Face15alfredplpl /simple-zundamon シンプルずんだもんデータセット はじめに ずんだもんの設定が詰まったシンプルなデータセットです。 作者がインターネットで調べたり、運営の人からもらったデータから作成しました。 キャラクターLLMを作るための動作確認にお使いください。 ただし、可能な限り動作確認でもライセンスをよく読んでください。 他の用途はライセンスをよく読んでください。 各種フォーマット ChatGPT: zmn.jsonl axolotlでの設定例 以下のようにデータセット周りを設定してください。 # データセットの設定 datasets: - path: alfredplpl/simple-zundamon # 使用するデータセット(Hugging Face上のデータセット名) type: chat_template # 会話形式のデータセットを使用 field_messages: messages #… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/simple-zundamon.textn<1K16 likes163 downloads6mo agoHugging Face16simplelex /Australian-Tax-Legislation-and-Amendment-History Australian Tax Legislation & Amendment History The full text of every section of the 22 principal Acts the Australian Taxation Office administers — Income Tax Assessment Act 1997, Income Tax Assessment Act 1936, the GST Act, FBTAA, the Taxation Administration Act 1953, the superannuation and fuel tax Acts and more — each section joined to its complete amendment history: which Act changed it, which schedule item, and when it commenced. Sourced from the Federal Register of… See the full description on the dataset page: https://huggingface.co/datasets/simplelex/Australian-Tax-Legislation-and-Amendment-History.text10K<n<100K1 likes131 downloads14d agoHugging Face17simplescaling /aime24_figuresThe 30 problems from AIME 2024 with all ASY code for figures. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}, year={2025}, eprint={2501.19393}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime24_figures.textn<1K0 likes120 downloads1y agoHugging Face18embedding-data /simple-wiki Dataset Card for "simple-wiki" Dataset Summary This dataset contains pairs of equivalent sentences obtained from Wikipedia. Supported Tasks Sentence Transformers training; useful for semantic search and sentence similarity. Languages English. Dataset Structure Each example in the dataset contains pairs of equivalent sentences and is formatted as a dictionary with the key "set" and a list with the sentences as "value". {"set":… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/simple-wiki.textsentence-similarity100K<n<1M11 likes99 downloads4y agoHugging Face19isp-uv-es /SimpleS2 SimpleS2 To load the data: import json import pickle # Read cube with open('cubo1_pickle', 'rb') as file: data = pickle.load(file).to_dataset(dim='band') # Read metadata with open('cubo1.json') as f: meta = json.load(f) Citation This dataset is related to the paper: arXiv:2506.19656tabularn<1K0 likes98 downloads1y agoHugging Face20simplescaling /aime_figuresThe 90 problems from AIME 2022, 2023, 2024 with all ASY code for figures. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}, year={2025}, eprint={2501.19393}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime_figures.textn<1K0 likes89 downloads2y agoHugging Face21alexfromapex /simplemath-cot 🧮 SimpleMath-100k CoT A chain-of-thought (CoT) extension of the ProCreations/SimpleMath dataset. Every one of the 100 000 algebra / arithmetic problems is paired with a short, numbered reasoning trace (Step 1: … Step 2: …) that walks a language model from the problem statement to the known-correct answer. The traces in the Jupyter notebook are generated by Qwen3.8-27B and then post-processed to strip formatting noise, enforce sequential step numbering, and cap output at 1 000… See the full description on the dataset page: https://huggingface.co/datasets/alexfromapex/simplemath-cot.texttext-generationn<1K0 likes85 downloads17d agoHugging Face22thisisandreeeee /simple-llm-sft Simple LLM SFT Dataset This synthetic dataset contains 1,000 English prompt-response pairs for supervised fine-tuning. It was created to fine-tune Qwen/Qwen3.5-4B to give clear, direct, and technically correct answers in simple English. The writing guidance is inspired by ASD-STE100 Simplified Technical English. The dataset does not claim official ASD-STE100 compliance or certification. Dataset structure The default configuration contains: Split Examples… See the full description on the dataset page: https://huggingface.co/datasets/thisisandreeeee/simple-llm-sft.texttext-generation1K<n<10K0 likes73 downloads8d agoHugging Face23billhdzhao /SimpleQuestions-3000text1K<n<10K1 likes66 downloads1y agoHugging Face24simplex-ai-inc /LiteResearcher-Corpustext10M<n<100M0 likes65 downloads4mo agoHugging Face25ProCreations /simple-facts Simple Facts A dataset of simple, no BS, human collected, ethicly sourced facts. About 1000 examples. This dataset is growing, and every day I plan to add a few more facts. texttext-generation1K<n<10K4 likes63 downloads1y agoHugging Face26naos-ku /SimpleMCQ SimpleMCQ Dataset Summary SimpleMCQ is a collection of multiple-choice question sets in the "fill-in-the-blank" format. Each item supplies a question sentence that contains a single blank ({}), a list of discrete answer options, and the index of the correct choice. The dataset is organized into four subsets—KR-200m, KR-200s, P-100, and P-20—and does not contain predefined splits such as train, validation, or test. Original paper is "Applying Relation Extraction and Graph… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/SimpleMCQ.textmultiple-choicen<1K0 likes63 downloads10mo agoHugging Face27RUC-AIBOX /0.8k-data-SimpleDeepSearchertabularn<1K5 likes61 downloads1y agoHugging Face28simplecloud /VidChain-Datatabular10K<n<100K1 likes59 downloads2y agoHugging Face29LLMTeamAkiyama /nvidia_openmathinstruct-2-simple-processed元データ https://huggingface.co/datasets/nvidia/OpenMathInstruct-2 tabular10M<n<100M0 likes57 downloads1y agoHugging Face30Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb_simple Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.texttext-generationn<1K0 likes56 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.