CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ScienceOne-AI /S1-Omni-Corpus-10K S1-Omni-Corpus-10K An open-source scientific multimodal reasoning dataset subset for S1-Omni 🧬 Model Introduction S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences. S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.image10K<n<100K1 likes920 downloads2mo agoHugging Face02ScienceOne-AI /S1-DeepResearch-15k S1-DeepResearch-15k Dataset Overview The S1-DeepResearch dataset is a curated collection of approximately 15k samples designed to improve deep research capabilities of large language models. The dataset includes two types of tasks: Verifiable tasks (labeled as "Closed-ended Multi-hop Resolution") Open-ended tasks (labeled as "Open-ended Exploration") Dataset Composition The dataset is organized into five core capability dimensions: Long-chain complex… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k.text10K<n<100K12 likes590 downloads5mo agoHugging Face03cua-ai /cua-s1-forms cua-s1-forms (dataset) Synthetic + real training/eval data for cua-ai/cua-s1-forms, a jev-like one-pass option scorer for GUI form filling behind cua-driver. Generator source: cua_s1/synth.py in https://github.com/trycua/cua/tree/main/libs/cua-s1. Files file rows source train.jsonl ~150k synthetic validation.jsonl ~18k synthetic test.jsonl ~20k synthetic, form-signature-disjoint from train/validation demo.jsonl 196 real: 3 real JevBrowser form… See the full description on the dataset page: https://huggingface.co/datasets/cua-ai/cua-s1-forms.texttext-classification100K<n<1M14 likes530 downloads5d agoHugging Face04CohereLabs /fusion-synth-data-s1kx Offline Synthetic Data (s1K-X) for: Making, not taking, the Best-of-N Content This data contains completions for the s1K-X training split prompts from 5 different teacher models and 2 aggregations: Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images. gemma3-27b: GEMMA3-27B-IT kimik2: KIMI-K2-INSTRUCT qwen3: QWEN3-235B… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-s1kx.texttext-generation10K<n<100K1 likes199 downloads1y agoHugging Face05semran1 /dclm-fix-dedup-s1text1M<n<10M0 likes171 downloads1y agoHugging Face06Silin1590 /Multilingual-S1text100K<n<1M0 likes171 downloads7mo agoHugging Face07mPLUG /UI_S1_dataset Introduction This repository contains the dataset example for UI-S1-7B, presented in UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning. Project page: https://github.com/X-PLUG/MobileAgent/tree/main/UI-S1 text1K<n<10K7 likes133 downloads11mo agoHugging Face08soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2400-s1-token-cachetabularn<1K0 likes127 downloads1mo agoHugging Face09Jackrong /Natural-Reasoning-gpt-oss-120B-S1 Dataset Card: Natural-Reasoning-gpt-oss-120B-S1 📜 Dataset Overview This is a meticulously curated instruction fine-tuning dataset designed specifically for efficient knowledge distillation tasks. Built upon the first 100,000 questions from the large-scale reasoning corpus facebook/natural_reasoning (s1, I will process the remaining parts later), it aims to transfer the advanced, multi-step reasoning capabilities of the teacher model gpt-oss-120-high to a student model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Natural-Reasoning-gpt-oss-120B-S1.texttext-generation10K<n<100K24 likes103 downloads11mo agoHugging Face10Abner0803 /Amazon-Beauty-S1This dataset is derived from Amazon Reviews'23 [1] Beauty category. The split is standard leave-one-out: the last item is the test target, the second-to-last is the validation target, and everything before that is training. The training portion is expanded by sliding window — every prefix becomes one example — so a user with a sequence of length $L$ contributes $L-3$ training rows with histories of length $1 \dots L-3$, one validation row and one test row. Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Abner0803/Amazon-Beauty-S1.texttext-generation10K<n<100K0 likes83 downloads20d agoHugging Face11baesad /s1K-1.1-deepseek-cot s1K-1.1 (DeepSeek-R1 traces) — format cho SegmentSelectiveSFT Chuyen doi tu simplescaling/s1K-1.1 bang prepare_s1k.py (default flags) trong repo SegmentSelectiveSFT. Moi dong jsonl co 3 truong: Truong Nguon question question solution deepseek_thinking_trajectory (long-CoT trace cua R1) answer \\boxed{...} cuoi cung trong trace, fallback ve solution cua s1K Giu 934 / 1000 mau — bo cac mau khong co trace, khong co dap an, hoac dap an dai hon 200 ky tu. from… See the full description on the dataset page: https://huggingface.co/datasets/baesad/s1K-1.1-deepseek-cot.texttext-generationn<1K0 likes76 downloads2d agoHugging Face12AngoHF /ANGO-S1ANGO is A Novel Generation-Oriented Chinese LLM evaluation benchmark. We introduces the format of single-question multiple-keypoints dataset for the first time, which include 171 keypoints accumulated in 4 hierarchical levels and 9 difficulty categories. The data were exclusively obtained from the Administrative Proficiency Test, which serves as a significant component of the Chinese civil service examination. We will apply a seasonal system for the leaderboard, updating them every two months.… See the full description on the dataset page: https://huggingface.co/datasets/AngoHF/ANGO-S1.tabularquestion-answering1K<n<10K3 likes74 downloads3y agoHugging Face13open-llm-leaderboard /ruizhe1217__sft-s1-qwen-0.5b-detailsgated Dataset Card for Evaluation run of ruizhe1217/sft-s1-qwen-0.5b Dataset automatically created during the evaluation run of model ruizhe1217/sft-s1-qwen-0.5b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ruizhe1217__sft-s1-qwen-0.5b-details.tabular10K<n<100K0 likes68 downloads2y agoHugging Face14tokyotech-llm /s1-test-time-scaling-synth-public s1-test-time-scaling-synth: Japanese and English Reinforcement Learning Dataset Derived from the s1 Simple Test-Time Scaling Dataset This repository contains s1-test-time-scaling-synth, a reinforcement learning dataset in Japanese and English.This dataset is built upon the supervised fine-tuning dataset simplescaling/data_ablation_full59K (hereafter, the "original dataset"), originally developed in "s1: Simple test-time scaling" [Muennighoff+, EMNLP25]. The original dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/s1-test-time-scaling-synth-public.texttext-generation10K<n<100K0 likes68 downloads7mo agoHugging Face15geraldaton20 /matmoe-s1-datatextn<1K0 likes67 downloads23d agoHugging Face16adraganov /arch-entity-matrix-lpi-260904T0010-plain_italy_n1000_e3_s1text1K<n<10K0 likes67 downloads19d agoHugging Face17soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2308-s1-token-cachetabularn<1K0 likes65 downloads1mo agoHugging Face18jayzou3773 /less-is-moe-s1-calibration-128 Less-is-MoE S1K calibration data — 128 full-length samples This repository contains the exact 128 S1K rows selected for Less-is-MoE full-model pruning. The selection reproduces the released loader: source: yentinglin/s1K-1.1-trl-format revision: 58a01564d278477da20ead1bcf1cde8e31f36251 split: train order: Dataset.shuffle(seed=1234) samples: first 128 nonempty messages rows sequence-length limit: none truncation: disabled padding: disabled calibration.jsonl stores every… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128.tabulartext-generationn<1K0 likes55 downloads5d agoHugging Face19soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-1667-s1-token-cachetabularn<1K0 likes47 downloads1mo agoHugging Face20adraganov /arch-code-scale-lpi-260904T0500-code_block_only_superintelligence_n1000_e3_s1_a0.5tabularn<1K0 likes47 downloads19d agoHugging Face21soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-0000-s1-token-cachetabularn<1K0 likes42 downloads1mo agoHugging Face22lwl-uestc /S1_QFFT 📘 S1–QFFT S1–QFFT is a question-free version of the original simplescaling/s1K-1.1 dataset, designed for QFFT training workflows. 🔍 Description This dataset discards the original questions and any system instructions, keeping only the reasoning completions as supervision. It is especially useful for models that aim to learn when and how to think, rather than just how to answer. The dataset is fully converted into a format compatible with LLaMA-Factory training.… See the full description on the dataset page: https://huggingface.co/datasets/lwl-uestc/S1_QFFT.texttext-generation1K<n<10K1 likes34 downloads1y agoHugging Face23tmobley96 /black_mirror_scripts_S1-5Black Mirror Scripts Dataset (Seasons 1-5) This dataset, titled 'black_mirror_scripts_S1-5.csv', contains the meticulously compiled transcripts of the critically acclaimed anthology series Black Mirror, covering Seasons 1 through 5. Each entry in this dataset is categorized by unique identifiers including Script ID, Title, Scene, Dialogue, and Timestamp, making it an ideal resource for natural language processing tasks, script analysis, sentiment analysis, and more. Dataset Composition Our… See the full description on the dataset page: https://huggingface.co/datasets/tmobley96/black_mirror_scripts_S1-5.text10K<n<100K2 likes31 downloads3y agoHugging Face24chrisvnz /s1-1aTesting a dataset conversion text1K<n<10K0 likes29 downloads2y agoHugging Face25s1lv3rj1nx /openjev-healthcare-router Healthcare router: a typed-decision benchmark with real headroom Synthetic pharmacy messages, routed by ten questions answered together: one intent choice, five multi-label noul topic flags, and four safety gates. 988 items across 18 tiers, 450 in test. Built for OpenJev to compare against a commercial typed-decision API, and deliberately built to be hard. Why it exists The first version of this task was useless. A commercial API scored 1.000 on four of its tiers… See the full description on the dataset page: https://huggingface.co/datasets/s1lv3rj1nx/openjev-healthcare-router.texttext-classificationn<1K0 likes29 downloads2d agoHugging Face26InfiX-ai /s1K-1.1-850This data is obtained by simplescaling/s1K-1.1. Compared with the original simplescaling/s1K-1.1 data, our filtered data uses less data and achieves better results. What we did Text Embedding Generation: We use all-MiniLM-L6-v2 (from SentenceTransformers library) to generate "input" embeddings. Dimensionality reduction: We use UMAP approach which preserves local and global data structures. n_components=2, n_neighbors=15, min_dist=0.1 Data Sparsification (Dense Points… See the full description on the dataset page: https://huggingface.co/datasets/InfiX-ai/s1K-1.1-850.textn<1K2 likes27 downloads2y agoHugging Face27CircularBalls /tt641-v3-seedsweep-s13 TT639G Recombined Tiny Assistant v1 Recombines isolated proof rungs: TT638D code behavior + dyadic/Mercy proof upstream TT639E2 context-copy behavior TT639F3 task-routing behavior simple rule/Q&A behavior Blocking dense gates: seen_combined_pass upstream_regression_pass mixed_heldout_pass anti_collision_pass Do not run dyadic/Mercy compare unless all four gates pass. text100K<n<1M0 likes26 downloads3mo agoHugging Face28vivifix /flock-off-s1-text-2-sqltextn<1K0 likes25 downloads11mo agoHugging Face29flock-io /flock-off-s1-competitiontext1K<n<10K0 likes25 downloads8mo agoHugging Face30FIRSTACCOUNT69 /ssrf-oauth-leak-s14 OAuth Secret Leak Test textn<1K0 likes24 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.