datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GlotStoryBook
Dataset Description
Story Books for 180 ISO-639-3 codes.
The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset.
This dataset consists of 2 subsets:
default, which consists of 4 publishers:
asp: African Storybook
pb: Pratham Books
lcb: Little Cree Books
lida: LIDA Stories
nalibali, which comes from Nal'ibali stories.
Usage (HF Loader)
default:
from datasets import load_dataset
dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.Math-Glot-Cleaned
Math-Glot-Cleaned
Math-Glot-Cleaned is a filtered and structured version of the NVIDIA AceReason-1.1-SFT dataset, focusing specifically on high-quality math and code-related question-answer pairs. This version is intended for use in training and evaluating reasoning-focused large language models, particularly for math problem solving and structured generation tasks.
Dataset Overview
Source: Derived from NVIDIA’s AceReason-1.1-SFT
Size: 22,232 examples
Format:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Glot-Cleaned.GlotCC-V1_malMath-Glot-Cleaned-10K
Math-Glot-Cleaned-10K
Math-Glot-Cleaned-10K is a curated subset of 10,000 math and code reasoning samples, extracted and cleaned from the original NVIDIA AceReason-1.1-SFT dataset. This refined version focuses exclusively on high-quality, structured mathematical prompts paired with chain-of-thought style reasoning.
Dataset Summary
Source: Derived from NVIDIA's AceReason-1.1-SFT dataset
Total Entries: 10,000
Format: Text-to-text (input → output)
Modality: Text… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Math-Glot-Cleaned-10K.
