CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asd567557275 /zhtw-roleplay-space-grimoire Space Grimoire RP Corpus (Traditional Chinese) Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0. 中文說明在下方 Dataset Summary Source text 283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.tabulartext-generation10K<n<100K1 likes128 downloads12d agoHugging Face02spade-rl /SPADE-Grounding-Corpus-ToolUse-15K SPADE grounding corpus: tool use (15k) Reference documents the SPADE Environment Designer is grounded on when generating multi-turn tool-use environments. 15,552 source files drawn from nvidia/Nemotron-Pretraining-Code-v3. Documents 15,552 Setting tool_use Fields text (the document), metadata (source provenance) Each generation prompt embeds one sampled document, so the environments a Designer writes stay anchored to a real concept or technique rather than… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-ToolUse-15K.texttext-generation10K<n<100K1 likes113 downloads1mo agoHugging Face03Thinking-Space /OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format. Data Creation and Cleaning This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.texttext-generation100K<n<1M3 likes107 downloads5mo agoHugging Face04Anoopsingh53 /isro-space-ocean-dataset ISRO Multimodal Space & Ocean Telemetry Dataset Official open-source scientific dataset curated for the National Space Day 2026 Hackathon and ISRO/IN-SPACe research submissions. Dataset Structure rain.jsonl: 1,204 high-precision instruction-tuning pairs mapping 6-band multispectral satellite telemetry (Coastal, Blue, Green, Red, NIR, SWIR) to atmospheric composition (O2 %, N2 %, Water Vapor g/m3) and oceanographic parameters (SST deg C, Salinity PSU).… See the full description on the dataset page: https://huggingface.co/datasets/Anoopsingh53/isro-space-ocean-dataset.texttabular-to-text1K<n<10K1 likes106 downloads1mo agoHugging Face05spade-rl /SPADE-Grounding-Corpus-Games-15K SPADE grounding corpus — games (15k) Reference documents the SPADE proposer is grounded on when generating cognitive-skill game environments. 15,000 documents: 10k drawn from a mathematics corpus and 5k from a science corpus. Documents 15,000 Setting games Fields Field Description text The document, exactly as embedded in the generation prompt metadata domain (mathematics / science) and url (source provenance) Each generation… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-Games-15K.texttext-generation10K<n<100K0 likes100 downloads29d agoHugging Face06CCCCCC /SPaR Dataset Card for SPaR Data Summary To enhance the instruction-following abilities of language models, we present SPaR, a self-play framework designed for continuous, autonomous improvement. SPaR focuses on generating high-quality preference pairs by minimizing interfering factors. We release an SFT dataset containing 8,000 samples curated using gpt-4o-mini. In addition, we provide DPO datasets derived from llama-3-8b-instruct and mistral-7b-instruct. Please refer to our… See the full description on the dataset page: https://huggingface.co/datasets/CCCCCC/SPaR.texttext-generation100K<n<1M8 likes99 downloads2y agoHugging Face07sparrowaisolutions /ember-dataset Ember Dataset Ember Dataset is a large-scale instruction-style dataset designed for training language models focused on creative writing, poetry generation, storytelling, and conversational responses. The dataset combines several well-known open instruction datasets and creative writing sources into a unified instruction–response format suitable for fine-tuning small and medium language models. The dataset is released by SparrowAISolutions. Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/sparrowaisolutions/ember-dataset.texttext-generation100K<n<1M0 likes82 downloads7mo agoHugging Face08Dude311 /spark-math-audit-20260911 Spark-X2.5: solving and auditing misleading worked solutions Status: experiment running; not a completed competition entry yet. Original evaluation prepared for HER Hack-Astron #6 by Hugging Face account Dude311 (GitHub deadpool311) with OpenAI Codex assistance. Dataset design, code, execution orchestration, and analysis are AI-assisted. Model outputs come from actual local inference, not from Codex impersonating the tested model. No human review of the model's reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/Dude311/spark-math-audit-20260911.texttext-generationn<1K0 likes69 downloads13d agoHugging Face09fabsssss /ssao-space-instruct ssao-space-instruct Instruction data teaching a language model to write valid RDF Turtle in the Space Situational Awareness Ontology (SSAO) for real space objects, and to judge proposed catalogue-to-ontology alignments using instance evidence. 1,245 examples: 1,072 train, 74 validation, 99 test. Built by Tesseract Academy. The construction principle No example asserts anything a validator cannot check. Every Turtle target is generated from a real CelesTrak SATCAT… See the full description on the dataset page: https://huggingface.co/datasets/fabsssss/ssao-space-instruct.texttext-generation1K<n<10K0 likes67 downloads2mo agoHugging Face10benjleite /FairytaleQA-translated-spanish Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.textquestion-answering10K<n<100K0 likes59 downloads1y agoHugging Face11spandyie /amadablam-dpo-preferences Ama Dablam DPO Preference Data Preference pairs used to DPO-tune Ama Dablam, a 322M trilingual (Nepali/Maithili/Bhojpuri) language model, across all three languages and three writing systems (Devanagari, IAST, phonetic romanization). See the technical report §9 for full methodology. Splits split rows purpose train 14,152 DPO Stage 2 preference-optimization training validation 744 preference-accuracy / forgetting evaluation warmup 3,203 Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/spandyie/amadablam-dpo-preferences.texttext-generation10K<n<100K0 likes54 downloads20d agoHugging Face12drewoodward /spanglish-sentences Spanglish Sentences A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models. Data format Each line of spanglish_sentences.jsonl is a JSON object with two fields: field description sentence A Spanglish utterance (mixed Spanish / English, or monolingual in either language). english_translation The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.texttranslation10K<n<100K0 likes51 downloads5mo agoHugging Face13osteele /mental-spaces Mental Spaces Corpus Version: 0.1.0 The Mental Spaces Corpus is a controlled suite of natural-language stimuli for testing whether language models keep base-space and alternative-space discourse targets separate. It is designed for probing, causal interventions, and behavioral readouts in mental-space constructions such as counterfactuals, belief contexts, and depictive spaces, including nested belief and nested depictive spaces. This release is a stimulus suite for controlled… See the full description on the dataset page: https://huggingface.co/datasets/osteele/mental-spaces.texttext-generation1K<n<10K0 likes49 downloads2mo agoHugging Face14SALT-NLP /SparkMe-SyntheticUsers SparkMe-SyntheticUsers Synthetic user profiles for evaluating AI interview systems, released alongside the SparkMe. Each profile represents a simulated workforce participant with demographic metadata, a shuffled list of persona facts, and structured ground-truth interview notes across 10 topics covering the impact of AI in the workplace. Dataset Description The 200 profiles were generated from WorkBank worker seed data using SparkMe's user agent pipeline. Each user has:… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SparkMe-SyntheticUsers.texttext-generationn<1K1 likes48 downloads6mo agoHugging Face15Spakie /DeepSeek-V4-Pro-distilled DeepSeek-V4-Pro-distilled 17,670 general-purpose instruction-following examples distilled from DeepSeek-V4-Pro, fact-checked and patched using GPT-5.5 Thinking. Pipeline Distillation — responses generated via DeepSeek-V4-Pro API Fact-checking — GPT-5.5 Thinking with web search reviewed all examples for factual errors and hallucinations Format Standard chat format, compatible with most SFT frameworks. Each row is one JSON object with a messages array:… See the full description on the dataset page: https://huggingface.co/datasets/Spakie/DeepSeek-V4-Pro-distilled.texttext-generation10K<n<100K1 likes48 downloads4mo agoHugging Face16Spakie /DeepSeek-V4-Pro-distill-V2 DeepSeek-V4-Pro-distill-V2 39,830 general-purpose chat and instruction-following examples distilled from DeepSeek-V4-Pro, fact-checked and patched using GPT-5.5 Thinking. Pipeline Distillation — responses generated via DeepSeek-V4-Pro API Fact-checking — GPT-5.5 Thinking with web search inside Codex reviewed all examples for factual errors, hallucinations and syntax/runtime erros in code examples Format Standard chat format, compatible with most… See the full description on the dataset page: https://huggingface.co/datasets/Spakie/DeepSeek-V4-Pro-distill-V2.texttext-generation10K<n<100K1 likes46 downloads3mo agoHugging Face17hybrid-diff-ar /stack-v2-sparse-classes-10k Stack v2 Sparse Python Classes 10k This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 9,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.tabulartext-generation10K<n<100K0 likes41 downloads5mo agoHugging Face18hybrid-diff-ar /stack-v2-sparse-classes-75kplus Stack v2 Sparse Python Classes 75kplus This is a frozen snapshot with 75829 samples for Diffusion + Autoregressive hybrid code generation experiments. Splits train.jsonl: 74829 val.jsonl: 500 test.jsonl: 500 all.jsonl: 75829 Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-75kplus.tabulartext-generation10K<n<100K0 likes33 downloads5mo agoHugging Face19rgjj30 /spanish-programmatic-seo-services-dataset Spanish Programmatic SEO & Services Dataset (1,249 Tracks) Este dataset de alta densidad contiene 1,249 trayectorias de agentes sintéticos diseñadas específicamente para el entrenamiento (fine-tuning) de modelos de lenguaje (LLMs) en tareas de razonamiento local, intenciones de búsqueda transaccionales y generación de estructuras SEO avanzadas para el mercado de España. Estructura del Dataset Cada registro sigue el formato de instrucción tuning estándar… See the full description on the dataset page: https://huggingface.co/datasets/rgjj30/spanish-programmatic-seo-services-dataset.texttext-generation1K<n<10K0 likes32 downloads13d agoHugging Face20hugoramallo /legal-ai-act-spanish-sft-7k⚠️ Legal and Liability Disclaimer This dataset is provided for research and educational purposes only. It does not constitute legal advice, nor does it represent an official or authoritative interpretation of Regulation (EU) 2024/1689 (EU AI Act). The content is synthetically generated and may contain errors, omissions, or hallucinations. Under no circumstances should this dataset be used as a basis for legal, compliance, or regulatory decision-making. The authors disclaim any liability for… See the full description on the dataset page: https://huggingface.co/datasets/hugoramallo/legal-ai-act-spanish-sft-7k.textquestion-answering1K<n<10K0 likes30 downloads6mo agoHugging Face21hybrid-diff-ar /stack-v2-sparse-classes-36k Stack v2 Sparse Python Classes 36k This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 35,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.tabulartext-generation10K<n<100K0 likes27 downloads5mo agoHugging Face22TokenHaven /FineWeb-Edu-Spanish High Quality Spanish Corpus This dataset contains a sample of a large collection of high-quality Spanish text data with their metadata. To access the full data please visit Token Haven Creation The dataset was created by filtering all English common crawl data for high-quality text using the FineWeb-Edu classifier with education score of 4 or higher over 5. The data is source from the v1.0.0 of the HuggingFaceFW/fineweb-edu dataset which corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/TokenHaven/FineWeb-Edu-Spanish.texttext-generationn<1K0 likes21 downloads1y agoHugging Face23spacekat99 /General_Conversation_Mixed_Datasettextquestion-answering1K<n<10K0 likes21 downloads4mo agoHugging Face24fineset-io /state-space-models-papers State Space Models & Mamba Papers — FineSet A research-paper dataset on State Space Models & Mamba Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on State Space Models & Mamba Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/state-space-models-papers.tabulartext-classificationn<1K0 likes14 downloads3mo agoHugging Face25Luisr-ecu /spanglish Spanglish Sentences A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models. Data format Each line of spanglish_sentences.jsonl is a JSON object with two fields: field description sentence A Spanglish utterance (mixed Spanish / English, or monolingual in either language). english_translation The English translation. When the source… See the full description on the dataset page: https://huggingface.co/datasets/Luisr-ecu/spanglish.texttranslation10K<n<100K0 likes13 downloads1mo agoHugging Face26que-app /cuban-spanish-sample que Cuban Spanish Conversational Sample (v0.3) A synthetic sample that demonstrates the schema of the que conversational dataset for Cuban Spanish (es-CU). It accompanies the que white paper and shows prospective research partners what a que record looks like. This is synthetic demonstration data, not a collected corpus. What this is These records were constructed to the production schema to illustrate its shape. They stand in for data that has yet to be collected… See the full description on the dataset page: https://huggingface.co/datasets/que-app/cuban-spanish-sample.texttext-generationn<1K0 likes7 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.