datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEA-Dataset
SEA-Dataset by Kreasof AI
The SEA-Dataset is a large-scale, multilingual, and instruction-based dataset curated by Kreasof AI. It combines over 34 high-quality, publicly available datasets, with a significant focus on enhancing the representation of Southeast Asian (SEA) languages. This dataset is designed for training and fine-tuning large language models (LLMs) to be more capable in a variety of domains including reasoning, mathematics, coding, and multilingual tasks, while also… See the full description on the dataset page: https://huggingface.co/datasets/kreasof-ai/SEA-Dataset.ECA-Zero
ECA-Zero: Elementary Cellular Automata Reasoning Dataset
A "CIFAR-for-Reasoning" dataset designed to test sequence models on Deduction, Induction, and Abduction tasks with strict Chain-of-Thought supervision.
This dataset implements the paradigm proposed in "Absolute Zero: Reinforced Self-play Reasoning with Zero Data" (Zhao et al., 2025), but adapted for the deterministic environment of Elementary Cellular Automata (Rule 0-255). It isolates algorithmic reasoning capabilities… See the full description on the dataset page: https://huggingface.co/datasets/kreasof-ai/ECA-Zero.
