datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStories-Multilingual
Novelist: TinyStories Multilingual Edition
Dataset Summary
The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes.
The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.Novelist
Dataset Card for Novelist
Dataset Summary
Novelist is a synthetic creative-writing and narrative-reasoning dataset designed for long-context fiction systems, scene planners, continuity-aware story models, multilingual literary translation, and child-safe TinyStories generation. The dataset mixes direct prose, explicit reasoning traces, quality-only judge outputs, full-book artifacts, and multilingual translation outputs inside a single narrative training ecosystem.
This… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Novelist.dx7-patches-and-prompts
Yamaha DX7 Synthesizer Patches with AI-Generated Prompts
Dataset Description
This is a comprehensive, multi-task dataset designed for fine-tuning language models to understand and generate synthesizer patches for the Yamaha DX7.
The dataset contains over 20,000 examples across three distinct but related tasks, making it ideal for creating models that can not only generate patches but also understand and reason about their structure and validity.
How the Data Was… See the full description on the dataset page: https://huggingface.co/datasets/ccerati/dx7-patches-and-prompts.pega-constellation-dx-sft
Pega Constellation DX Components — Training Dataset
This dataset teaches a code model how to write Pega Constellation DX components — the custom React/TypeScript components that extend the Pega Platform UI.
It comes from the open-source constellation-ui-gallery repo, which Pega itself maintains as a reference for DX component authors. We pinned a specific commit so the dataset is reproducible.
What is in it
55 Pega DX components, each broken into 9 different… See the full description on the dataset page: https://huggingface.co/datasets/WeekendNoobs/pega-constellation-dx-sft.Novelist-CoT
Novelist Announcement.
Check out the new and biggest Novelist project [https://huggingface.co/datasets/Dxniz/Novelist]. With 5 thinking modes, agentic writing and long book support. Now only limitation is your imagination.
Novelist-CoT Creative Writing Dataset
Overview
Novelist-CoT is a long-form creative writing dataset designed for supervised fine-tuning and style-focused narrative generation.The dataset consolidates multiple generation pipelines into a… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Novelist-CoT.new_azerbaijan
Azerbaijani News Dataset
Description
Introducing new_azerbaijan, a rich and diverse dataset consisting of 30k [30,812] Azerbaijani news articles, meticulously collected using well-crafted heuristics. This dataset spans a wide range of common news topics, including war, government, politics, education, health, the environment, economy, business, fashion, entertainment, sports, and even unique or unconventional events.
The dataset is designed for both abstractive and… See the full description on the dataset page: https://huggingface.co/datasets/DxeiZ/new_azerbaijan.
