datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UltraData-Math
UltraData-Math
🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README
UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models.
🆕 What's New… See the full description on the dataset page: https://huggingface.co/datasets/fishingguy/UltraData-Math.zygai_fishing_without_borders_2004
🎣 ZygAI – Fishing Without Borders 2004
A structured dataset of Lithuanian lakes and ichthyofauna (2004 edition)Dataset by ZygAI Research
🧭 About This Dataset
This dataset is an official extension of the ZygAI project’s✅ Atostogos Lietuvos kaime 2004 dataset.
The extended section, titled “Žvejyba be sienų 2004” (“Fishing Without Borders 2004”), contains detailed freshwater fishing information from the same publication.
It includes:
Lake names (LT + EN)
Municipal… See the full description on the dataset page: https://huggingface.co/datasets/ZygAI/zygai_fishing_without_borders_2004.
