CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenStellarTeam /Chinese-SimpleQA Overview 🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper • 📊 Leaderboard Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, our benchmark covers 6 major topics with 99 diverse subtopics. Please visit our website or check our paper for more details.… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SimpleQA.textquestion-answering1K<n<10K38 likes2.4k downloads2y agoHugging Face02simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes410 downloads1h agoHugging Face03simpleG2023 /chinese-clean-energy-battery-open-intelligence 🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.tabulartext-retrieval1K<n<10K0 likes313 downloads55m agoHugging Face04simpleG2023 /chinese-ai-and-robotics-open-intelligence 🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes284 downloads46m agoHugging Face05simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes274 downloads1h agoHugging Face06alexfromapex /simplemath-cot 🧮 SimpleMath-100k CoT A chain-of-thought (CoT) extension of the ProCreations/SimpleMath dataset. Every one of the 100 000 algebra / arithmetic problems is paired with a short, numbered reasoning trace (Step 1: … Step 2: …) that walks a language model from the problem statement to the known-correct answer. The traces in the Jupyter notebook are generated by Qwen3.8-27B and then post-processed to strip formatting noise, enforce sequential step numbering, and cap output at 1 000… See the full description on the dataset page: https://huggingface.co/datasets/alexfromapex/simplemath-cot.texttext-generationn<1K0 likes63 downloads20d agoHugging Face07MichaelAnthony /echidna-simplerag-massive echidna-simplerag-massive Echidna — large SimpleRAG assistant dataset. Contents simplerag_massive.jsonl (120 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the Echidna RAG assistant (Michael Anthony Falabella). textquestion-answeringn<1K0 likes43 downloads29d agoHugging Face08MichaelAnthony /snowfox-simplerag-phase5 snowfox-simplerag-phase5 SnowFox — SimpleRAG phase 5 (messages). Contents simplerag_snowfox_phase5_messages.jsonl (185 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for SnowFox / SimpleRAG (Michael Anthony Falabella). textquestion-answeringn<1K0 likes43 downloads29d agoHugging Face09MichaelAnthony /echidna-simplerag-massive-combined echidna-simplerag-massive-combined Echidna — combined SimpleRAG massive dataset. Contents simplerag_massive_combined.jsonl (152 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the Echidna RAG assistant (Michael Anthony Falabella). textquestion-answeringn<1K0 likes41 downloads29d agoHugging Face10Impulse2000 /simple_bench_public-20-12-2024 Simple Bench Public Where Everyday Human Reasoning Still Surpasses Frontier Models. Dataset Details Dataset Description "[...] A multiple-choice text benchmark for LLMs where individuals with unspecialized (high school) knowledge outperform SOTA models. SimpleBench includes over 200 questions covering spatio-temporal reasoning, social intelligence, and what we call linguistic adversarial robustness (or trick questions). For the vast majority of text-based… See the full description on the dataset page: https://huggingface.co/datasets/Impulse2000/simple_bench_public-20-12-2024.textquestion-answeringn<1K0 likes39 downloads1y agoHugging Face11sapiens-technology /simple_bench 📊 Simple Bench Dataset A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.texttext-generationn<1K0 likes39 downloads5mo agoHugging Face12MichaelAnthony /echidna-round3-simplerag echidna-round3-simplerag Echidna — round 3 SimpleRAG-specific extraction examples. Contents round3_simplerag.jsonl (18 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the Echidna RAG assistant (Michael Anthony Falabella). textquestion-answeringn<1K0 likes37 downloads29d agoHugging Face13MichaelAnthony /snowfox-simplerag-phase4 snowfox-simplerag-phase4 SnowFox — SimpleRAG phase 4 (messages). Contents simplerag_snowfox_phase4_messages.jsonl (185 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for SnowFox / SimpleRAG (Michael Anthony Falabella). textquestion-answeringn<1K0 likes34 downloads29d agoHugging Face14MisterAI /SimpleSmallFrenchQA Présentation Dépôt de datasets de type "Question/Réponse" (QR/QA) en Français. Ces jeux de données sont conçus pour l'entraînement et l'évaluation de modèles de traitement du langage naturel (NLP). Description des Datasets Contenu Le dépôt contient plusieurs jeux de données : Questions/Réponses Générales : Questions sur des sujets variés. Réponses factuelles et prouvées. Questions/Réponses d'Évaluation : Questions conçues pour tester la compréhension et la… See the full description on the dataset page: https://huggingface.co/datasets/MisterAI/SimpleSmallFrenchQA.textquestion-answeringn<1K1 likes24 downloads2y agoHugging Face15joppari /mn_business_benchmark_dataset_simple mn_business_benchmark_dataset_2000_diverse Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset. Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан. Schema id: 1-ээс 2000 хүртэлх дараалсан дугаар instruction: бизнесийн бодлогын өгүүлбэр input: хоосон string thinking: бодолт… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_simple.texttext-generation1K<n<10K1 likes16 downloads4mo agoHugging Face16giangnt /TVM-SimpleQAgated Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details textquestion-answering1K<n<10K0 likes13 downloads10mo agoHugging Face17JImyai123 /jimy-simple-datasettexttext-generationn<1K0 likes8 downloads10mo agoHugging Face18PhatNguyen41 /vietnamese-qa-simple Vietnamese Simple Question Answering Dataset A small Vietnamese QA dataset designed for chatbot and QA model training. Data Fields question answer Intended Use Chatbots Question answering systems textquestion-answeringn<1K0 likes4 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.