CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simplescaling /aime24_nofiguresThe 30 problems from AIME 2024 only with the ASY code for figures when it is necessary to solve the problem. Figure code that is not core to the problem was excluded. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime24_nofigures.textn<1K2 likes13k downloads1y agoHugging Face02OpenStellarTeam /Chinese-SimpleQA Overview 🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper • 📊 Leaderboard Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, our benchmark covers 6 major topics with 99 diverse subtopics. Please visit our website or check our paper for more details.… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SimpleQA.textquestion-answering1K<n<10K38 likes2.4k downloads2y agoHugging Face03allenai /SimpleToM SimpleToM Dataset and Evaluation data The SimpleToM dataset of stories with associated questions are described in the paper "SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs" Associated evaluation data for the models analyzed in the paper can be found in the separate dataset: SimpleToM-eval-data. Question sets There are three question sets in the SimpleToM dataset: mental-state-qa questions about information awareness… See the full description on the dataset page: https://huggingface.co/datasets/allenai/SimpleToM.text1K<n<10K11 likes2.4k downloads7mo agoHugging Face04simplelex /ATO-Australian-Tax-Rulings-and-Guidance ATO Rulings & Guidance — Australian Tax Law, Structured for AI 67,000+ Australian Taxation Office documents as RAG-ready NDJSON/CSV — Edited Private Advice, public rulings and determinations, ATO Interpretative Decisions, practical compliance guidelines, taxpayer alerts, decision impact statements, practice statements and legislative instruments. Every document parsed into structured, typed fields for legal RAG, LLM fine-tuning, and tax research automation. Machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/simplelex/ATO-Australian-Tax-Rulings-and-Guidance.text10K<n<100K1 likes1.6k downloads12h agoHugging Face05k19862217 /simpsons_script_linestext10K<n<100K0 likes1.1k downloads3y agoHugging Face06embedding-data /simple-wiki Dataset Card for "simple-wiki" Dataset Summary This dataset contains pairs of equivalent sentences obtained from Wikipedia. Supported Tasks Sentence Transformers training; useful for semantic search and sentence similarity. Languages English. Dataset Structure Each example in the dataset contains pairs of equivalent sentences and is formatted as a dictionary with the key "set" and a list with the sentences as "value". {"set":… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/simple-wiki.textsentence-similarity100K<n<1M11 likes497 downloads4y agoHugging Face07simplescaling /aime25_nofigurestextn<1K0 likes440 downloads2y agoHugging Face08simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes419 downloads5h agoHugging Face09SimpleFunctions /sf-index-history SimpleFunctions Index History Time series of the SF Index: a four-number summary of prediction-market consensus — disagreement (0-100), geo-risk (0-100), breadth (-1..+1), and activity (0-100) — computed every 15 minutes from ~50K markets. Flat JSONL for easy charting / analysis. License and Use This dataset is released under Creative Commons Attribution 4.0 International (CC-BY-4.0; https://creativecommons.org/licenses/by/4.0/). You may use it freely for personal… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/sf-index-history.tabular10K<n<100K0 likes414 downloads18h agoHugging Face10simplescaling /aime_nofiguresThe 90 problems from AIME 2022, 2023, 2024 only with the ASY code for figures when it is necessary to solve the problem. Figure code that is not core to the problem was excluded. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime_nofigures.textn<1K1 likes352 downloads2y agoHugging Face11simpleG2023 /chinese-clean-energy-battery-open-intelligence 🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.tabulartext-retrieval1K<n<10K0 likes320 downloads5h agoHugging Face12simpleG2023 /chinese-ai-and-robotics-open-intelligence 🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes295 downloads5h agoHugging Face13simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes282 downloads5h agoHugging Face14LLukas22 /nq-simplified Dataset Card for "nq" Dataset Summary This is a modified version of the original Natural Questions (nq) dataset for qa tasks. The original is availabe here. Each sample was preprocessed into a squadlike format. The context was shortened from an entire wikipedia article into the passage containing the answer. Dataset Structure Data Instances An example of 'train' looks as follows. { "context": "The 2017 Major League Baseball All - Star Game was… See the full description on the dataset page: https://huggingface.co/datasets/LLukas22/nq-simplified.textquestion-answering100K<n<1M4 likes279 downloads3y agoHugging Face15deep-analysis-research /simple-evalstext100K<n<1M0 likes222 downloads11mo agoHugging Face16alibaba-pai /SimpleQA-Bench SimpleQA-Bench Tags: factuality, EN, ZH, short-form-answer, human-label Copyright: © 2024 alibaba-pai Source.OpenAI's SimpleQA: Blog & Paper / Data & simple-evals ProjectOpenStellarTeam's Chinese-SimpleQA: Blog & Paper, Data@HF Factuality is a complicated topic because it is hard to measure—evaluating the factuality of any given arbitrary claim is challenging, and language models can generate long completions that contain dozens of factual claims. In SimpleQA, we will focus on… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/SimpleQA-Bench.text1K<n<10K3 likes189 downloads2y agoHugging Face17simplescaling /openaimathThe 500 problems from MATH that were used in "Let's verify step-by-step" and "s1: Simple test-time scaling" are the test set; rest is regular train set. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/openaimath.text10K<n<100K5 likes187 downloads1y agoHugging Face18alfredplpl /simple-zundamon シンプルずんだもんデータセット はじめに ずんだもんの設定が詰まったシンプルなデータセットです。 作者がインターネットで調べたり、運営の人からもらったデータから作成しました。 キャラクターLLMを作るための動作確認にお使いください。 ただし、可能な限り動作確認でもライセンスをよく読んでください。 他の用途はライセンスをよく読んでください。 各種フォーマット ChatGPT: zmn.jsonl axolotlでの設定例 以下のようにデータセット周りを設定してください。 # データセットの設定 datasets: - path: alfredplpl/simple-zundamon # 使用するデータセット(Hugging Face上のデータセット名) type: chat_template # 会話形式のデータセットを使用 field_messages: messages #… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/simple-zundamon.textn<1K16 likes168 downloads6mo agoHugging Face19simplelex /Australian-Tax-Legislation-and-Amendment-History Australian Tax Legislation & Amendment History The full text of every section of the 22 principal Acts the Australian Taxation Office administers — Income Tax Assessment Act 1997, Income Tax Assessment Act 1936, the GST Act, FBTAA, the Taxation Administration Act 1953, the superannuation and fuel tax Acts and more — each section joined to its complete amendment history: which Act changed it, which schedule item, and when it commenced. Sourced from the Federal Register of… See the full description on the dataset page: https://huggingface.co/datasets/simplelex/Australian-Tax-Legislation-and-Amendment-History.text10K<n<100K1 likes147 downloads2d agoHugging Face20simplescaling /aime24_figuresThe 30 problems from AIME 2024 with all ASY code for figures. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}, year={2025}, eprint={2501.19393}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime24_figures.textn<1K0 likes113 downloads1y agoHugging Face21vulturuldemare /Estonian-Text-Simplification Estonian Text Simplification This repository contains resources and models for Estonian text simplification, including datasets and pre-trained models. Files Dataset Files simplification_training_set.json: A dataset used to fine-tune LLaMA 3.1 for text simplification. src: The source of the data. original: The original sentence to be simplified. simpl_lex: A lexical simplification (may be empty). simpl_final: The final simplified sentence.… See the full description on the dataset page: https://huggingface.co/datasets/vulturuldemare/Estonian-Text-Simplification.text10K<n<100K0 likes105 downloads2y agoHugging Face22Nellyw888 /VeriReason-RTL-Coder_7b_reasoning_tb_simple Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason and our page Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log 2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple Project Description This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.texttext-generationn<1K0 likes89 downloads1y agoHugging Face23simplescaling /aime_figuresThe 90 problems from AIME 2022, 2023, 2024 with all ASY code for figures. Citation Information @misc{muennighoff2025s1simpletesttimescaling, title={s1: Simple test-time scaling}, author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto}, year={2025}, eprint={2501.19393}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/simplescaling/aime_figures.textn<1K0 likes85 downloads2y agoHugging Face24hasankursun /age-specific-text-simplification Age-Specific Text Simplification Dataset Dataset Description This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group. Dataset Summary Total Examples: 17,177 Training Split: 15,459 examples Validation Split: 1,718 examples Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.tabulartext-generation10K<n<100K3 likes77 downloads1y agoHugging Face25ProCreations /simple-facts Simple Facts A dataset of simple, no BS, human collected, ethicly sourced facts. About 1000 examples. This dataset is growing, and every day I plan to add a few more facts. texttext-generation1K<n<10K4 likes74 downloads1y agoHugging Face26isp-uv-es /SimpleS2 SimpleS2 To load the data: import json import pickle # Read cube with open('cubo1_pickle', 'rb') as file: data = pickle.load(file).to_dataset(dim='band') # Read metadata with open('cubo1.json') as f: meta = json.load(f) Citation This dataset is related to the paper: arXiv:2506.19656tabularn<1K0 likes71 downloads1y agoHugging Face27naos-ku /SimpleMCQ SimpleMCQ Dataset Summary SimpleMCQ is a collection of multiple-choice question sets in the "fill-in-the-blank" format. Each item supplies a question sentence that contains a single blank ({}), a list of discrete answer options, and the index of the correct choice. The dataset is organized into four subsets—KR-200m, KR-200s, P-100, and P-20—and does not contain predefined splits such as train, validation, or test. Original paper is "Applying Relation Extraction and Graph… See the full description on the dataset page: https://huggingface.co/datasets/naos-ku/SimpleMCQ.textmultiple-choicen<1K0 likes65 downloads10mo agoHugging Face28alexfromapex /simplemath-cot 🧮 SimpleMath-100k CoT A chain-of-thought (CoT) extension of the ProCreations/SimpleMath dataset. Every one of the 100 000 algebra / arithmetic problems is paired with a short, numbered reasoning trace (Step 1: … Step 2: …) that walks a language model from the problem statement to the known-correct answer. The traces in the Jupyter notebook are generated by Qwen3.8-27B and then post-processed to strip formatting noise, enforce sequential step numbering, and cap output at 1 000… See the full description on the dataset page: https://huggingface.co/datasets/alexfromapex/simplemath-cot.texttext-generationn<1K0 likes63 downloads21d agoHugging Face29RUC-AIBOX /0.8k-data-SimpleDeepSearchertabularn<1K5 likes60 downloads1y agoHugging Face30billhdzhao /SimpleQuestions-3000text1K<n<10K1 likes60 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.