CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SimpleStories /SimpleStories 📘📕 SimpleStories 📙📗 SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository. If you'd like to commission other languages or story formats, feel free to send mail. When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories.tabulartext-generation1M<n<10M39 likes2.7k downloads9mo agoHugging Face02Xuhui /sim-posttrain HUMANUAL Posttraining Data Posttraining data for user simulation, derived from the train splits of the HUMANUAL benchmark datasets. Datasets HUMANUAL (posttraining) Config Rows Description news 48,618 News article comment responses politics 45,429 Political discussion responses opinion 37,791 Reddit AITA / opinion thread responses book 34,170 Book review responses chat 23,141 Casual chat responses email 6,377 Email reply responses… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sim-posttrain.tabulartext-generation1M<n<10M1 likes1.7k downloads5mo agoHugging Face03SimVer-ano /simverse2026 SimVerse ⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review. A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.imagevisual-question-answering1K<n<10K0 likes472 downloads5mo agoHugging Face04The-CoLab /multilingual-textarena-SimpleTak-v0-train-v2 TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train-v2.tabulartext-generation100K<n<1M0 likes192 downloads3mo agoHugging Face05hugfaceguy0001 /simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing. tabulartext-classificationn<1K0 likes152 downloads3y agoHugging Face06SimpleStories /SimpleStories-JA 📘📕 SimpleStories 📙📗 このデータセットは、gpt-4o-miniによって生成された短編小説で出来ているデータセットです。生成方法や、自分で物語を生成する方法については、こちらのリポジトリをご覧ください。 他の言語や物語形式の制作を希望される場合は、メールにてお問い合わせください。 SimpleStoriesは、EldenとLiによるTinyStoriesの改良版です。 特徴 物語の注釈情報(theme、topic、styleなど) 多様性の高さ 2024年のモデルによって生成 NLPのデータが用意しているためフィルタリングしやすい 以下の言語版が利用可能: 英語 日本語 他にも追加予定 This dataset is a collection of short stories generated by gpt-4o-mini (+ other models, soon). To see how this dataset was generated, or to generate some stories… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories-JA.tabulartext-generation1M<n<10M1 likes147 downloads2y agoHugging Face07The-CoLab /multilingual-textarena-SimpleTak-v0-train TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train.tabulartext-generation1M<n<10M0 likes145 downloads1mo agoHugging Face08minnesotanlp /lawflow-reasoning-simulation LawFlow: Collecting and Simulating Lawyers' Thought Processes Debarati Das, Khanh Chi Le*, Ritik Parkar*, Karin De Langis, Brendan Madson, Chad Berryman, Robin Willis, Daniel Moses, Brett McDonnell†, Daniel Schwarcz†, Dongyeop Kang† Minnesota NLP, University of Minnesota Twin Cities *equal contribution, †senior advisors Arxiv Project Page Dataset Summary and Purpose LawFlow: Collecting and Simulating Lawyers' Thought Processes The purpose of this dataset is aim… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/lawflow-reasoning-simulation.tabulartext-generationn<1K2 likes92 downloads1y agoHugging Face09SimPPL /sakhi Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.tabularquestion-answering1K<n<10K2 likes89 downloads5mo agoHugging Face10trillionlabs /SimScholar-SFT S3 SFT Trajectories Complete ReAct trajectories for scientific-literature search. Code · S3 collection · Source corpus This dataset contains 14,633 single- and two-hop tool-use trajectories. In each trajectory, a policy searches and reads a fixed scientific corpus through nine tools, then submits an answer with a correctness label. The messages column uses OpenAI tool-calling chat format. At a glance Question type Rows Correct Incorrect Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.tabularquestion-answering10K<n<100K0 likes87 downloads2mo agoHugging Face11Neura-parse /quantum-simulation-chemistry-materials Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.tabulartext-generation100K<n<1M0 likes83 downloads3mo agoHugging Face12hasankursun /age-specific-text-simplification Age-Specific Text Simplification Dataset Dataset Description This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group. Dataset Summary Total Examples: 17,177 Training Split: 15,459 examples Validation Split: 1,718 examples Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.tabulartext-generation10K<n<100K3 likes80 downloads1y agoHugging Face13duoduoyeah /SimpleStories 📘📕 SimpleStories 📙📗 SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository. If you'd like to commission other languages or story formats, feel free to send mail. When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/duoduoyeah/SimpleStories.tabulartext-generation1M<n<10M0 likes70 downloads9mo agoHugging Face14mj33 /SimCoPilot Dataset Card for Dataset Name SimCoPilot is a benchmark for evaluating LLMs to perform as a "copilot"-style, interactive coding assistant. Dataset Details Dataset Description SimCoPilot is a benchmark for evaluating LLMs to perform as a "copilot"-style, interactive coding assistant, testing their ability to add and complete code in complex real-world software environments and analyzing how LLMs manage different code dependencies and logic complexities. Curated… See the full description on the dataset page: https://huggingface.co/datasets/mj33/SimCoPilot.tabulartext-generation1K<n<10K1 likes50 downloads1y agoHugging Face15simone-papicchio /bird Dataset This dataset is a polished version of the BIRD dataset. It was introduced in the paper Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL. It has been used to train the reasoning Text2SQL model simone-papicchio/Think2SQL-7B. Please refer to the paper for further details. License: CC BY-SA 4.0 Citation @misc{papicchio2025think2sqlreinforcellmreasoning, title={Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL}, author={Simone… See the full description on the dataset page: https://huggingface.co/datasets/simone-papicchio/bird.tabulartext-generation10K<n<100K0 likes50 downloads8mo agoHugging Face16Aratako /iterative-dpo-data-for-SimPO-iter2 iterative-dpo-data-for-SimPO-iter2 概要 合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kを元に以下のような手順で作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter1を用いて、temperature=1で回答を5回生成 5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施 1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置 全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-SimPO-iter2.tabulartext-generation10K<n<100K1 likes40 downloads2y agoHugging Face17Ayaka /ORCHESTRA-simple-1M ORCHESTRA-simple-1M GitHub: nk2028/ORCHESTRA-dataset 中文簡介 ORCHESTRA (cOmpRehensive Classical cHinESe poeTRy dAtaset) 是一個全面的古典中文詩歌的數據集,數據來自搜韻網。本數據集由 nk2028 進行格式轉換並發佈,希望透過公開高品質的古典中文詩歌數據,促進對古典中文詩歌及古典中文自然語言處理的研究。 ORCHESTRA-simple 是 ORCHESTRA 數據集的簡化格式,僅保留 id, title, group_index, type, dynasty, author, content 這 7 個欄位,而去除其他欄位,以簡化使用。 本資料集可用於大型語言模型的訓練。如欲作其他用途,請向數據提供者搜韻網諮詢。 English Introduction ORCHESTRA (cOmpRehensive Classical cHinESe poeTRy dAtaset) is a comprehensive dataset of classical… See the full description on the dataset page: https://huggingface.co/datasets/Ayaka/ORCHESTRA-simple-1M.tabulartext-generation1M<n<10M6 likes36 downloads3y agoHugging Face18jgchaicoski /datasus_sim Dataset Card for [DATASUS SIM] This dataset is a large-scale collection of information around deaths registered by the Brazilian public health care system. Dataset Details Check the Data Dictionary attached to the project. Dataset Sources Repository: [Link to HF Repo or GitHub] Paper: [Optional: Link to ArXiv or Journal] Demo: [Optional: Link to Space or Web App] Uses Direct Use Pre-training: Training Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/jgchaicoski/datasus_sim.documenttext-generation10M<n<100M0 likes28 downloads7mo agoHugging Face19LiteMind /Simple-agent-traces 📱 Simple Agent Traces – Tiny Tool‑Calling Conversations for Small Models Simple Agent Traces is a compact, hand‑picked dataset of 605 real‑world tool‑calling conversations, each carefully truncated to ≤8,192 tokens (using the SmolLM2‑360M tokenizer).It is purpose‑built for training and fine‑tuning tiny language models (≤500M) that must run on‑device – smartphones, edge devices, or any environment with strict memory and latency constraints. 🧹 No chain‑of‑thought, no fluff.Every… See the full description on the dataset page: https://huggingface.co/datasets/LiteMind/Simple-agent-traces.tabulartext-generationn<1K3 likes26 downloads4mo agoHugging Face20simpissa /countdown-qwen3-0.6b Countdown Qwen3-0.6B Pass@10 Buckets Countdown arithmetic problems filtered by observed local Qwen/Qwen3-0.6B success rate over 10 rollouts per problem. Each problem asks for an arithmetic expression that reaches a target using each listed source number at most once. The final answer should be inside \boxed{...}. Canonical solutions are provided, but any verifier-valid expression is accepted. Subsets subset source bucket count observed successes out of 10… See the full description on the dataset page: https://huggingface.co/datasets/simpissa/countdown-qwen3-0.6b.tabulartext-generation1K<n<10K0 likes22 downloads4mo agoHugging Face21jmp1987 /simson-assembly-dependency-graph 🔗 Simson Assembly Dependency Graph 13 „Braucht-auch"-Ketten für Simson-Reparaturen und Tuning. Was das ist Jeder Eintrag beschreibt, was man zusätzlich braucht wenn man ein bestimmtes Teil einbaut: Pflicht-Teile (mandatory_with): Ohne diese geht es nicht Empfohlene Teile (recommended_with): Sinnvoll, aber optional Inkompatible Teile: Was NICHT gleichzeitig verbaut werden kann Upgrade-Pfad: Was als nächstes Sinn macht Geschätzte Arbeitszeit & Skill-Level Kosten:… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-assembly-dependency-graph.tabularquestion-answeringn<1K0 likes18 downloads4mo agoHugging Face22simonzimmo /Complete-FABLE.5-traces-2M Complete FABLE.5 Traces 2M Full FABLE.5 / Mythos corpus restored, with session-limit answer rows removed. Dataset Viewer | Parquet | Raw JSONL.gz This dataset is a post-closure compilation of all available FABLE.5 / Mythos trace datasets found on Hugging Face during the curation pass after the closure of Fable and Mythos. It is deduplicated at the normalized-row level and keeps row-level provenance through first_source_dataset, first_source_config, first_source_split… See the full description on the dataset page: https://huggingface.co/datasets/simonzimmo/Complete-FABLE.5-traces-2M.tabulartext-generation1M<n<10M1 likes17 downloads3mo agoHugging Face23jmp1987 /simson-youtube-tutorials 📺 Simson YouTube Tutorial Metadata 20 kuratierte YouTube-Tutorial-Einträge für Simson-Moped Reparatur, Tuning und Restaurierung. Inhalt Strukturierte Metadaten der wichtigsten Simson-Tutorial-Videos auf YouTube: Kanal-Typen: DIY-Werkstatt, Tuning-Spezialist, Restaurierungs-Kanal, Enthusiasten-Kanal, Dokumentation Topics: Motor, Zündung, Vergaser, Elektrik, Tuning, Restaurierung, Fahrwerk, Geschichte, Wartung Fahrzeuge: S50, S51, S70, KR51/1, KR51/2 (Schwalbe)… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-youtube-tutorials.tabulartext-generationn<1K0 likes15 downloads4mo agoHugging Face24jmp1987 /simson-unified-knowledge-graph 🧠 Simson Unified Knowledge Graph 173 Nodes × 348 Edges – der Klebstoff zwischen allen Simson-Datasets. Was das ist Ein maschinenlesbarer Graph, der alle 6 Datasets miteinander verknüpft: Dataset Status Nodes racing-planet-simson-traces Diagnose-Traces 15 simson-forum-qa-pairs Forum-Wissen 30 simson-repair-manual Technische Daten 14 racing-planet-product-catalog Teilekatalog 37 simson-youtube-tutorials Video-Tutorials 20… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-unified-knowledge-graph.tabularquestion-answeringn<1K0 likes11 downloads4mo agoHugging Face25deepaksamuel-cuk /drich-simhitstabulartext-generation10K<n<100K0 likes7 downloads7mo agoHugging Face26SIMBA9657 /haddas-tigrinya-corpus haddas-tigrinya-corpus Monolingual Tigrinya newspaper text segmented into article bodies. Suitable for continued pretraining or causal language modeling of Tigrinya LLMs. Source Derived from the Haddas Eritrea newspaper archive: 63 PDF issues processed by the haddas-eritrea pipeline (extract -> clean -> segment -> translate -> label). Generated: 2026-04-26 12:21 UTC Row count: 2653 Schema: id, text, char_count, topic, issue_date, source_pdf, page_start, page_end… See the full description on the dataset page: https://huggingface.co/datasets/SIMBA9657/haddas-tigrinya-corpus.tabulartext-generation1K<n<10K0 likes7 downloads5mo agoHugging Face27transZ /controlled_text_simplygatedControlled text simplification, targetting at different audiences. A dataset for a group project. tabulartext-generation10K<n<100K0 likes2 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.