CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bertievidgen /SimpleSafetyTeststexttext-generationn<1K12 likes3.3k downloads3y agoHugging Face02SimpleStories /SimpleStories 📘📕 SimpleStories 📙📗 SimpleStories is a dataset of >2 million model-generated short stories. It was made to train small, interpretable language models on it. The generation process is open-source: To see how the dataset was generated, or to generate some stories yourself, head over to this repository. If you'd like to commission other languages or story formats, feel free to send mail. When using SimpleStories in your work, please cite the SimpleStories paper:… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories.tabulartext-generation1M<n<10M39 likes2.8k downloads9mo agoHugging Face03Xuhui /sim-posttrain HUMANUAL Posttraining Data Posttraining data for user simulation, derived from the train splits of the HUMANUAL benchmark datasets. Datasets HUMANUAL (posttraining) Config Rows Description news 48,618 News article comment responses politics 45,429 Political discussion responses opinion 37,791 Reddit AITA / opinion thread responses book 34,170 Book review responses chat 23,141 Casual chat responses email 6,377 Email reply responses… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sim-posttrain.tabulartext-generation1M<n<10M1 likes1.7k downloads5mo agoHugging Face04pszemraj /simple_wikipedia simple wikipedia the 'simple' split of Wikipedia, from Sept 1 2023. The train split contains about 65M tokens, Pulled via: dataset = load_dataset( "wikipedia", language="simple", date="20230901", beam_runner="DirectRunner" ) stats train split general info <class 'pandas.core.frame.DataFrame'> RangeIndex: 226242 entries, 0 to 226241 Data columns (total 4 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 id… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/simple_wikipedia.texttext-generation100K<n<1M11 likes1.3k downloads9mo agoHugging Face05fblgit /simple-math Simple Math: 2+2=4 -1=3 (LoLo: Learning Only Logical Operations) Just like my teacher gave me homework, i thought maybe we can also add some of these basics on the trainings of our models. It was created with very simple code that is in the repo, if you add more complex operations and so.. please share the code :D thank you Current Code Version: 20240127.fblgit (A modification over @win10 for progressive and DPO operation) Does it Works? 34BEAGLES… See the full description on the dataset page: https://huggingface.co/datasets/fblgit/simple-math.texttext-generation100K<n<1M19 likes617 downloads3y agoHugging Face06SimVer-ano /simverse2026 SimVerse ⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review. A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.imagevisual-question-answering1K<n<10K0 likes475 downloads5mo agoHugging Face07Similoluwa /african-multilingual-tokenizer-challenge African Multilingual Tokenizer Challenge dataset The frozen public corpus for the African Multilingual Tokenizer Challenge. It contains one balanced multilingual train split and one balanced validation split. Split Per language Total Train 40,000 240,000 Validation 4,000 24,000 Languages are English (en), French (fr), Hausa (ha), Swahili (sw), Yoruba (yo) and Amharic (am). Official public-test and private-test text are deliberately absent from this repository.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/african-multilingual-tokenizer-challenge.texttext-generation100K<n<1M0 likes416 downloads23d agoHugging Face08while-ai /agent-simulations Agent Simulations Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets 53,971 synthetic agent trajectories generated by simulations across 34 agent types. The rows include successful and failed trajectories for supervised fine-tuning, preference work, reinforcement learning, and evaluation. NOTE: This is generated test and training data, not curated ground truth. Review and filter it for your application before training or… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.texttext-generation10K<n<100K0 likes381 downloads2d agoHugging Face09ellamind /simpleqa-verified-multilingual SimpleQA Verified Multilingual Multilingual translations of SimpleQA Verified, a 1,000-prompt factuality benchmark from Google DeepMind that evaluates short-form parametric knowledge (facts stored in model weights). Source: google/simpleqa-verified (eval split, 1,000 examples) Languages Config Language Examples ces Czech 100 dan Danish 100 deu German 1,000 fra French 100 ita Italian 100 nld Dutch 100 pol Polish 100 spa Spanish 100 More to… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/simpleqa-verified-multilingual.textquestion-answering1K<n<10K1 likes373 downloads7mo agoHugging Face10Orange /simplequestions-sparqltotext Dataset Card for SimpleQuestions-SPARQLtoText Dataset Summary Special version of SimpleQuestions with SPARQL queries formatted for the SPARQL-to-Text task. JSON fields The original version of SimpleQuestions is a raw text file listing triples and the natural language question. A JSON version has been generated and augmented with the following fields: rdf_subject, rdf_property, rdf_object: triple in the Wikidata format (IDs) nl_subject, nl_property, nl_object:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/simplequestions-sparqltotext.textquestion-answering10K<n<100K2 likes271 downloads3y agoHugging Face11simplex-ai-inc /LiteResearcher-SFT-Data LiteResearcher — SFT Cold-Start Data Distilled deep-research trajectories used for the SFT cold-start of LiteResearcher-4B This dataset contains the 68,231 multi-turn deep-research trajectories used to train the SFT cold-start checkpoint that RL (GRPO+TIS) is later launched from — the "68.2 K distilled deep-research trajectories" referenced in the paper and in LiteResearcher-Data. Each row is a complete ReAct-style episode: a research question, the model's interleaved… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-SFT-Data.textquestion-answering10K<n<100K0 likes262 downloads2mo agoHugging Face12simutrade /simutrade-rag-sft-28k 📢 Domain & Email Migration Notice From May 6th, 2026, Simutrade will transition to new domains as simutrade.app will not be renewed: 🌐 Website: simutrade.faizath.com (formerly simutrade.app) ⚙️ API: simutrade-api.faizath.com (formerly api.simutrade.app) 📧 Email: contact@simutrade.faizath.com (formerly contact@simutrade.app) 🛰️ CDN: simutrade-cdn.faizath.com (formerly cdn.simutrade.app) 📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/simutrade/simutrade-rag-sft-28k.textquestion-answering10K<n<100K1 likes206 downloads1mo agoHugging Face13mkurman /simplescaling-s1K-R1 Dataset Card: s1k R1 Dataset Description The s1k R1 dataset is a fork of the simplescaling/s1K dataset. It contains a collection of conversations where the assistant's messages have been enhanced to include Chain of Thought (CoT) reasoning within <think> ... </think> tags, followed by the final answer. This modification aims to improve the interpretability and reasoning capabilities of AI models by providing explicit thought processes in the responses. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/simplescaling-s1K-R1.texttext-generation1K<n<10K0 likes202 downloads2y agoHugging Face14The-CoLab /multilingual-textarena-SimpleTak-v0-train-v2 TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train-v2.tabulartext-generation100K<n<1M0 likes193 downloads3mo agoHugging Face15Lots-of-LoRAs /task1347_glue_sts-b_similarity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1347_glue_sts-b_similarity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1347_glue_sts-b_similarity_classification.texttext-generation1K<n<10K0 likes181 downloads2y agoHugging Face16hugfaceguy0001 /simpsons_infoThe information of all episodes of the cartoon show "The Simpsons" from wikipedia. Some (mainly in recent 32, 33, 34 seasons) plot missing. tabulartext-classificationn<1K0 likes157 downloads3y agoHugging Face17SimpleStories /SimpleStories-JA 📘📕 SimpleStories 📙📗 このデータセットは、gpt-4o-miniによって生成された短編小説で出来ているデータセットです。生成方法や、自分で物語を生成する方法については、こちらのリポジトリをご覧ください。 他の言語や物語形式の制作を希望される場合は、メールにてお問い合わせください。 SimpleStoriesは、EldenとLiによるTinyStoriesの改良版です。 特徴 物語の注釈情報(theme、topic、styleなど) 多様性の高さ 2024年のモデルによって生成 NLPのデータが用意しているためフィルタリングしやすい 以下の言語版が利用可能: 英語 日本語 他にも追加予定 This dataset is a collection of short stories generated by gpt-4o-mini (+ other models, soon). To see how this dataset was generated, or to generate some stories… See the full description on the dataset page: https://huggingface.co/datasets/SimpleStories/SimpleStories-JA.tabulartext-generation1M<n<10M1 likes155 downloads2y agoHugging Face18The-CoLab /multilingual-textarena-SimpleTak-v0-train TextArena Language Trajectories This dataset contains language-conditioned TextArena trajectory data. Each dataset configuration corresponds to a different model, experiment group, or source folder. Available configurations: gemma4-e4b-it qwen3-4b ministral3-3b-instruct Usage Install the datasets library: pip install datasets Load a specific configuration: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/The-CoLab/multilingual-textarena-SimpleTak-v0-train.tabulartext-generation1M<n<10M0 likes146 downloads1mo agoHugging Face19while-ai /tau2-simulated tau2 Simulated Training Set Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets The training set that took a base model from 5% to 30% on tau2-bench telecom, made from nothing but the agent's tool list and policy. If you build a customer-facing agent, you already have the two files this dataset was made from: the tools it can call and the policy it follows. The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.texttext-generation1K<n<10K0 likes135 downloads2d agoHugging Face20pszemraj /simple_wikipedia_LM Dataset Card for "simple_wikipedia_LM" A filtered/edited version of pszemraj/simple_wikipedia that removes headings/contents that appear in the text column without any relevant text for them (at least in the simple split). import re def split_on_headings(text): headings = ["References", "Related pages", "Other websites", "Further reading"] for heading in headings: parts = re.split( r"^\s*" + re.escape(heading) + r".*$", text, flags=re.MULTILINE… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/simple_wikipedia_LM.texttext-generation100K<n<1M13 likes133 downloads9mo agoHugging Face21simmo /CanlIICaseSummaries Canadian Case Law Summaries A database of (currently, still growing) >600 case law summaries generated by GPT 4 for random case law in Ontario or Canada textsummarizationn<1K3 likes105 downloads3y agoHugging Face22ProCreations /SimpleMath 🧮 SimpleMath 100K SimpleMath 100K is a high-quality synthetic dataset of 100,000 basic arithmetic problems — no noise, no tricks, just clean and accurate math. ✅ Purpose This was made for small AI models — not to struggle with complex math, but to get simple math right every time. 📦 Contents 75,000 numeric problems, evenly split: 18,750 addition (456 + 789 =) 18,750 subtraction (900 - 345 =) 18,750 multiplication (12 x 15 =) 18,750 division (144 / 12 =)… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/SimpleMath.texttext-generation100K<n<1M8 likes103 downloads1y agoHugging Face23Paillat /simple-wiki-article Simple-wiki-article This dataset is directly derived from rahular/simple-wikipedia. Various techniques were used to detect which strings were article titles, and separate the original dataset into articles, with a somewhat good accuracy. Article content strings were merged and split with \n, and <br> was replaced with \n to make the dataset more usable. texttext-generation100K<n<1M1 likes102 downloads1y agoHugging Face24pszemraj /simplepile-lite Dataset Card for "simplepile-lite" Interleaved dataset using 'first exhausted' strategy. Counts: DatasetDict({ train: Dataset({ features: ['text'], num_rows: 452432 }) validation: Dataset({ features: ['text'], num_rows: 1000 }) test: Dataset({ features: ['text'], num_rows: 11908 }) }) token counts - train using GPTNeoX Tokenizer: token_count count 452432 mean 868.642 std 4791.71… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/simplepile-lite.textfill-mask100K<n<1M1 likes97 downloads9mo agoHugging Face25minnesotanlp /lawflow-reasoning-simulation LawFlow: Collecting and Simulating Lawyers' Thought Processes Debarati Das, Khanh Chi Le*, Ritik Parkar*, Karin De Langis, Brendan Madson, Chad Berryman, Robin Willis, Daniel Moses, Brett McDonnell†, Daniel Schwarcz†, Dongyeop Kang† Minnesota NLP, University of Minnesota Twin Cities *equal contribution, †senior advisors Arxiv Project Page Dataset Summary and Purpose LawFlow: Collecting and Simulating Lawyers' Thought Processes The purpose of this dataset is aim… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/lawflow-reasoning-simulation.tabulartext-generationn<1K2 likes97 downloads1y agoHugging Face26SimPPL /sakhi Sakhi: A Community-Validated Multilingual Maternal-Health Benchmark Sakhi is a benchmark for evaluating large language models on maternal and reproductive-health questions in three languages spoken in low-resource settings: English, Hindi, and Marathi. It was built around a deployed WhatsApp-based maternal-health chatbot reaching rural mothers in Hindi- and Marathi-speaking districts of India, with a three-channel review pipeline: practising Indian doctors, Accredited Social Health… See the full description on the dataset page: https://huggingface.co/datasets/SimPPL/sakhi.tabularquestion-answering1K<n<10K2 likes92 downloads5mo agoHugging Face27simplex-ai-inc /LiteResearcher-Data LiteResearcher — RL Training Data Companion training data for the LiteResearcher paper A low-cost, scalable Agentic RL training framework for deep-research agents. This dataset contains the two-stage curriculum of question–answer prompts used to train LiteResearcher-4B with on-policy GRPO+TIS, fully against a local search / browse environment. Both stages share the same validation set. What this is not: the underlying webpage corpus (~32 M records, used by the local… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-Data.textquestion-answering10K<n<100K2 likes89 downloads4mo agoHugging Face28Lots-of-LoRAs /task146_afs_argument_similarity_gun_control Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task146_afs_argument_similarity_gun_control Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task146_afs_argument_similarity_gun_control.texttext-generation1K<n<10K0 likes87 downloads2y agoHugging Face29Lots-of-LoRAs /task934_turk_simplification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task934_turk_simplification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task934_turk_simplification.texttext-generation1K<n<10K0 likes84 downloads2y agoHugging Face30trillionlabs /SimScholar-SFT S3 SFT Trajectories Complete ReAct trajectories for scientific-literature search. Code · S3 collection · Source corpus This dataset contains 14,633 single- and two-hop tool-use trajectories. In each trajectory, a policy searches and reads a fixed scientific corpus through nine tools, then submits an answer with a correctness label. The messages column uses OpenAI tool-calling chat format. At a glance Question type Rows Correct Incorrect Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.tabularquestion-answering10K<n<100K0 likes84 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.