datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStories-Multilingual
Novelist: TinyStories Multilingual Edition
Dataset Summary
The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes.
The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.Novelist
Dataset Card for Novelist
Dataset Summary
Novelist is a synthetic creative-writing and narrative-reasoning dataset designed for long-context fiction systems, scene planners, continuity-aware story models, multilingual literary translation, and child-safe TinyStories generation. The dataset mixes direct prose, explicit reasoning traces, quality-only judge outputs, full-book artifacts, and multilingual translation outputs inside a single narrative training ecosystem.
This… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Novelist.new_azerbaijan
Azerbaijani News Dataset
Description
Introducing new_azerbaijan, a rich and diverse dataset consisting of 30k [30,812] Azerbaijani news articles, meticulously collected using well-crafted heuristics. This dataset spans a wide range of common news topics, including war, government, politics, education, health, the environment, economy, business, fashion, entertainment, sports, and even unique or unconventional events.
The dataset is designed for both abstractive and… See the full description on the dataset page: https://huggingface.co/datasets/DxeiZ/new_azerbaijan.
