datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MegaMath
MegaMath: Pushing the Limits of Open Math Copora
Megamath is part of TxT360, curated by LLM360 Team.
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens.
MegaMath is curated via the following three efforts:
Revisiting web data:
We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MegaMath.dlgenai-nppe-datasetdigenai-nppe-datasetdigenai-nppe-datasetnpb_data_appindian-english-nptel-v0dlgenai-nppe2-datasetOMat24_train_aimd_from_PBE_1000_npt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 1000 npt. ColabFit, 2025. https://doi.org/10.60732/25f16f85
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jqrkc9e7cgmh_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_1000_npt.NPTEL
BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages
Overview
BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English.
This repository consists of parallel data for Speech Translation from NPTEL, a subset of BhasaAnuvaad.
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/NPTEL.NPTEL_TECH_ENG_50NPset-2-Python-Edu
NPset-2 (Python-Edu)
A normalized semi-synthetic Python dataset for training small language models on code logic without the overhead of raw code syntax.
Why
Small language models trained on natural language corpora develop latent representations of logical constructs -- iteration, conditionals, data flow, function composition -- yet struggle to apply this reasoning to source code, where syntactic overhead (delimiters, indentation conventions, language-specific idioms)… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/NPset-2-Python-Edu.OMat24_train_aimd_from_PBE_3000_npt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 3000 npt. ColabFit, 2025. https://doi.org/10.60732/edd12490
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6xvvh8yl7rfd_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_3000_npt.fund-holdings-nport
US Fund and ETF Holdings — Form N-PORT
Every position of every US mutual fund and ETF, monthly, including the
bonds, loans, asset-backed paper and derivatives that 13F does not report at
all.
145 630 731 positions · 17 513 funds · 341 049 monthly reports ·
2019-09-30 to 2026-05-31
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
Why this and not 13F
13F is what everyone uses because it is what everyone knows about. It… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/fund-holdings-nport.indian-english-nptel-testMathNet
Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation
This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1.
Quick start
from datasets import load_dataset
# Default: all problems
ds = load_dataset("ShadenA/MathNet", split="train")
# Or a specific country / competition-body config… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MathNet.nphard_tsp2dolmino-mix-1124-OLMo-2-0425-1B-tokenizer-npyNPset-python
NPset
A normalized semi-sythetic Python dataset for training small language models on code logic without the overhead of raw code syntax.
Why
Small language models trained on natural language corpora develop latent representations of logical constructs -- iteration, conditionals, data flow, function composition -- yet struggle to apply this reasoning to source code, where syntactic overhead (delimiters, indentation conventions, language-specific idioms) occupies a… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/NPset-python.GAPS-nptel
GAPS: Golden-Aligned Parallel Speech Corpus
Overview
GAPS (Golden-Aligned Parallel Speech) is a multi-corpus dataset designed for foreign accent conversion.
The dataset provides parallel speech triplets consisting of:
Original non-native speech
Parallel native speech
Golden speaker speech — synthetic speech that preserves the non-native speaker’s timbre and timing (including pauses) while exhibiting native pronunciation
along with the corresponding text transcript.
GAPS… See the full description on the dataset page: https://huggingface.co/datasets/warisqr007/GAPS-nptel.URAG
URAG: Uncertainty-aware RAG Evaluation
Paper
URAG is a framework for evaluating RAG (Retrieval-Augmented Generation) systems with conformal prediction on multiple-choice QA benchmarks.
Running experiments (YAML only)
First, you need to clone our GitHub repo at: https://github.com/phuvinhnguyen/URAG
Install dependencies:
pip install -r requirements.txt
Run everything via a config file:
python cli.py --config <path/to/config.yaml>
Examples (LCA commit-message… See the full description on the dataset page: https://huggingface.co/datasets/npvinHnivqn/URAG.MsVRAG-Bench
MsVRAG-Bench — Multi-Step Video RAG Benchmark
MsVRAG-Bench is a structured benchmark for evaluating multi-step retrieval-augmented generation (RAG) over video. Each example pairs a set of ordered video segment clips with a natural-language question that requires reasoning over which segments are visible and which are missing, simulating real RAG pipelines where a retriever may not return all relevant evidence.
Quick start
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/npvinHnivqn/MsVRAG-Bench.dl-npppe2-datasetAraSeg-2026-Shared-Task-NP
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NP.finemath
📐 FineMath
What is it?
📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+) of mathematical educational content filtered from CommonCrawl. To curate this dataset, we trained a mathematical content classifier using annotations generated by LLama-3.1-70B-Instruct. We used the classifier to retain only the most educational mathematics content, focusing on clear explanations and step-by-step problem solving rather than… See the full description on the dataset page: https://huggingface.co/datasets/NP235/finemath.AraSeg-2026-Shared-Task-NoPnx-NP
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-NP.ai_art_np
Dataset Card for "ai_art_np"
More Information needed
npy_file_hsrNPM-Artifact-Explanation-Benchmark
NPM-Artifact-Explanation-Benchmark
English
NPM-Artifact-Explanation-Benchmark is a cross-category multimodal corpus and benchmark resource for Chinese cultural artifact understanding and explanation.
This release contains 28,826 cleaned artifact records derived from National Palace Museum source records' opendata (https://digitalarchive.npm.gov.tw/opendata/). Each record includes structured artifact metadata, image URLs, source record URLs, and human-written… See the full description on the dataset page: https://huggingface.co/datasets/shunanhe/NPM-Artifact-Explanation-Benchmark.tau-indian-banking
TauIndianBankBench: an Indian retail-banking benchmark for tool-using conversational agents
Data files for the indian_banking domain of the tau2-bench harness: a conversational agent
staffs the chat desk of a synthetic India-based retail bank. The code that runs and scores the
benchmark lives in the companion repository
npci/TauIndianBankBench;
this dataset holds only the domain data it loads.
Viewer note. The dataset viewer shows preview/tasks_preview.parquet, which holds the… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/tau-indian-banking.open-web-math-pro
📚 Open-Web-Math-Pro
ArXiv | Models | Code
Open-Web-Math-Pro is refined from open-web-math using the ProX refining framework.
It contains about 5B high quality math related tokens, ready for pre-training.
License
Open-Web-Math-Pro is based on open-web-math, which is made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU: https://commoncrawl.org/terms-of-use/. We do not alter the license of any of the underlying data.… See the full description on the dataset page: https://huggingface.co/datasets/NP235/open-web-math-pro.
