datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.SPADE-Grounding-Corpus-ToolUse-15K
SPADE grounding corpus: tool use (15k)
Reference documents the SPADE Environment Designer is grounded on when generating multi-turn tool-use environments. 15,552 source files drawn from nvidia/Nemotron-Pretraining-Code-v3.
Documents
15,552
Setting
tool_use
Fields
text (the document), metadata (source provenance)
Each generation prompt embeds one sampled document, so the environments a Designer
writes stay anchored to a real concept or technique rather than… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-ToolUse-15K.OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B
OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format.
Data Creation and Cleaning
This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.isro-space-ocean-dataset
ISRO Multimodal Space & Ocean Telemetry Dataset
Official open-source scientific dataset curated for the National Space Day 2026 Hackathon and ISRO/IN-SPACe research submissions.
Dataset Structure
rain.jsonl: 1,204 high-precision instruction-tuning pairs mapping 6-band multispectral satellite telemetry (Coastal, Blue, Green, Red, NIR, SWIR) to atmospheric composition (O2 %, N2 %, Water Vapor g/m3) and oceanographic parameters (SST deg C, Salinity PSU).… See the full description on the dataset page: https://huggingface.co/datasets/Anoopsingh53/isro-space-ocean-dataset.SPADE-Grounding-Corpus-Games-15K
SPADE grounding corpus — games (15k)
Reference documents the SPADE proposer is grounded on when generating cognitive-skill
game environments. 15,000 documents: 10k drawn from a mathematics corpus and 5k from a
science corpus.
Documents
15,000
Setting
games
Fields
Field
Description
text
The document, exactly as embedded in the generation prompt
metadata
domain (mathematics / science) and url (source provenance)
Each generation… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-Games-15K.SPaR
Dataset Card for SPaR
Data Summary
To enhance the instruction-following abilities of language models, we present SPaR, a self-play framework designed for continuous, autonomous improvement. SPaR focuses on generating high-quality preference pairs by minimizing interfering factors.
We release an SFT dataset containing 8,000 samples curated using gpt-4o-mini. In addition, we provide DPO datasets derived from llama-3-8b-instruct and mistral-7b-instruct.
Please refer to our… See the full description on the dataset page: https://huggingface.co/datasets/CCCCCC/SPaR.ember-dataset
Ember Dataset
Ember Dataset is a large-scale instruction-style dataset designed for training language models focused on creative writing, poetry generation, storytelling, and conversational responses.
The dataset combines several well-known open instruction datasets and creative writing sources into a unified instruction–response format suitable for fine-tuning small and medium language models.
The dataset is released by SparrowAISolutions.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/sparrowaisolutions/ember-dataset.spark-math-audit-20260911
Spark-X2.5: solving and auditing misleading worked solutions
Status: experiment running; not a completed competition entry yet.
Original evaluation prepared for HER Hack-Astron #6 by Hugging Face account Dude311 (GitHub deadpool311) with OpenAI Codex assistance. Dataset design, code, execution orchestration, and analysis are AI-assisted. Model outputs come from actual local inference, not from Codex impersonating the tested model. No human review of the model's reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/Dude311/spark-math-audit-20260911.ssao-space-instruct
ssao-space-instruct
Instruction data teaching a language model to write valid RDF Turtle in the
Space Situational Awareness Ontology (SSAO) for real space objects, and to
judge proposed catalogue-to-ontology alignments using instance evidence.
1,245 examples: 1,072 train, 74 validation, 99 test. Built by
Tesseract Academy.
The construction principle
No example asserts anything a validator cannot check. Every Turtle target is
generated from a real CelesTrak SATCAT… See the full description on the dataset page: https://huggingface.co/datasets/fabsssss/ssao-space-instruct.FairytaleQA-translated-spanish
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Spanish machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-spanish.amadablam-dpo-preferences
Ama Dablam DPO Preference Data
Preference pairs used to DPO-tune Ama Dablam,
a 322M trilingual (Nepali/Maithili/Bhojpuri) language model, across all three languages
and three writing systems (Devanagari, IAST, phonetic romanization). See the
technical report §9 for full
methodology.
Splits
split
rows
purpose
train
14,152
DPO Stage 2 preference-optimization training
validation
744
preference-accuracy / forgetting evaluation
warmup
3,203
Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/spandyie/amadablam-dpo-preferences.spanglish-sentences
Spanglish Sentences
A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models.
Data format
Each line of spanglish_sentences.jsonl is a JSON object with two fields:
field
description
sentence
A Spanglish utterance (mixed Spanish / English, or monolingual in either language).
english_translation
The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.mental-spaces
Mental Spaces Corpus
Version: 0.1.0
The Mental Spaces Corpus is a controlled suite of natural-language stimuli for testing
whether language models keep base-space and alternative-space discourse targets
separate. It is designed for probing, causal interventions, and behavioral readouts in
mental-space constructions such as counterfactuals, belief contexts, and depictive
spaces, including nested belief and nested depictive spaces.
This release is a stimulus suite for controlled… See the full description on the dataset page: https://huggingface.co/datasets/osteele/mental-spaces.SparkMe-SyntheticUsers
SparkMe-SyntheticUsers
Synthetic user profiles for evaluating AI interview systems, released alongside the SparkMe. Each profile represents a simulated workforce participant with demographic metadata, a shuffled list of persona facts, and structured ground-truth interview notes across 10 topics covering the impact of AI in the workplace.
Dataset Description
The 200 profiles were generated from WorkBank worker seed data using SparkMe's user agent pipeline. Each user has:… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SparkMe-SyntheticUsers.DeepSeek-V4-Pro-distilled
DeepSeek-V4-Pro-distilled
17,670 general-purpose instruction-following examples distilled from DeepSeek-V4-Pro, fact-checked and patched using GPT-5.5 Thinking.
Pipeline
Distillation — responses generated via DeepSeek-V4-Pro API
Fact-checking — GPT-5.5 Thinking with web search reviewed all examples for factual errors and hallucinations
Format
Standard chat format, compatible with most SFT frameworks. Each row is one JSON object with a messages array:… See the full description on the dataset page: https://huggingface.co/datasets/Spakie/DeepSeek-V4-Pro-distilled.DeepSeek-V4-Pro-distill-V2
DeepSeek-V4-Pro-distill-V2
39,830 general-purpose chat and instruction-following examples distilled from DeepSeek-V4-Pro, fact-checked and patched using GPT-5.5 Thinking.
Pipeline
Distillation — responses generated via DeepSeek-V4-Pro API
Fact-checking — GPT-5.5 Thinking with web search inside Codex reviewed all examples for factual errors, hallucinations and syntax/runtime erros in code examples
Format
Standard chat format, compatible with most… See the full description on the dataset page: https://huggingface.co/datasets/Spakie/DeepSeek-V4-Pro-distill-V2.stack-v2-sparse-classes-10k
Stack v2 Sparse Python Classes 10k
This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 9,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.stack-v2-sparse-classes-75kplus
Stack v2 Sparse Python Classes 75kplus
This is a frozen snapshot with 75829 samples for Diffusion + Autoregressive hybrid code generation experiments.
Splits
train.jsonl: 74829
val.jsonl: 500
test.jsonl: 500
all.jsonl: 75829
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-75kplus.spanish-programmatic-seo-services-dataset
Spanish Programmatic SEO & Services Dataset (1,249 Tracks)
Este dataset de alta densidad contiene 1,249 trayectorias de agentes sintéticos diseñadas específicamente para el entrenamiento (fine-tuning) de modelos de lenguaje (LLMs) en tareas de razonamiento local, intenciones de búsqueda transaccionales y generación de estructuras SEO avanzadas para el mercado de España.
Estructura del Dataset
Cada registro sigue el formato de instrucción tuning estándar… See the full description on the dataset page: https://huggingface.co/datasets/rgjj30/spanish-programmatic-seo-services-dataset.legal-ai-act-spanish-sft-7k⚠️ Legal and Liability Disclaimer
This dataset is provided for research and educational purposes only.
It does not constitute legal advice, nor does it represent an official or authoritative interpretation of Regulation (EU) 2024/1689 (EU AI Act).
The content is synthetically generated and may contain errors, omissions, or hallucinations.
Under no circumstances should this dataset be used as a basis for legal, compliance, or regulatory decision-making.
The authors disclaim any liability for… See the full description on the dataset page: https://huggingface.co/datasets/hugoramallo/legal-ai-act-spanish-sft-7k.stack-v2-sparse-classes-36k
Stack v2 Sparse Python Classes 36k
This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 35,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.FineWeb-Edu-Spanish
High Quality Spanish Corpus
This dataset contains a sample of a large collection of high-quality Spanish text data with their metadata.
To access the full data please visit Token Haven
Creation
The dataset was created by filtering all English common crawl data for high-quality text using the FineWeb-Edu classifier with education score of 4 or higher over 5.
The data is source from the v1.0.0 of the HuggingFaceFW/fineweb-edu dataset which corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/TokenHaven/FineWeb-Edu-Spanish.General_Conversation_Mixed_Datasetstate-space-models-papers
State Space Models & Mamba Papers — FineSet
A research-paper dataset on State Space Models & Mamba Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on State Space Models & Mamba Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/state-space-models-papers.spanglish
Spanglish Sentences
A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models.
Data format
Each line of spanglish_sentences.jsonl is a JSON object with two fields:
field
description
sentence
A Spanglish utterance (mixed Spanish / English, or monolingual in either language).
english_translation
The English translation. When the source… See the full description on the dataset page: https://huggingface.co/datasets/Luisr-ecu/spanglish.cuban-spanish-sample
que Cuban Spanish Conversational Sample (v0.3)
A synthetic sample that demonstrates the schema of the que conversational dataset for Cuban Spanish (es-CU). It accompanies the que white paper and shows prospective research partners what a que record looks like. This is synthetic demonstration data, not a collected corpus.
What this is
These records were constructed to the production schema to illustrate its shape. They stand in for data that has yet to be collected… See the full description on the dataset page: https://huggingface.co/datasets/que-app/cuban-spanish-sample.
