CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NP235 /MegaMath MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MegaMath.texttext-generation100M<n<1B0 likes8.9k downloads3mo agoHugging Face02Abhishek-A0 /dlgenai-nppe-datasettabularn<1K0 likes1.7k downloads6h agoHugging Face03pk1308 /digenai-nppe-datasettabular100K<n<1M0 likes1.3k downloads29d agoHugging Face04Vancheeswaran /digenai-nppe-datasettabularn<1K0 likes1.2k downloads29d agoHugging Face05Ramos-Ramos /npb_data_apptabular10M<n<100M1 likes707 downloads1h agoHugging Face06skbose /indian-english-nptel-v0audio100K<n<1M3 likes627 downloads2y agoHugging Face0722f2001542 /dlgenai-nppe2-datasettabularn<1K0 likes595 downloads6d agoHugging Face08colabfit /OMat24_train_aimd_from_PBE_1000_npt Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 1000 npt. ColabFit, 2025. https://doi.org/10.60732/25f16f85 This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_jqrkc9e7cgmh_0 Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_1000_npt.tabular10M<n<100M0 likes590 downloads11mo agoHugging Face09ai4bharat /NPTELgated BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This repository consists of parallel data for Speech Translation from NPTEL, a subset of BhasaAnuvaad. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/NPTEL.audio1M<n<10M9 likes394 downloads2y agoHugging Face10speechlab /NPTEL_TECH_ENG_50textautomatic-speech-recognition10K<n<100K0 likes386 downloads11mo agoHugging Face11AxiomicLabs /NPset-2-Python-Edu NPset-2 (Python-Edu) A normalized semi-synthetic Python dataset for training small language models on code logic without the overhead of raw code syntax. Why Small language models trained on natural language corpora develop latent representations of logical constructs -- iteration, conditionals, data flow, function composition -- yet struggle to apply this reasoning to source code, where syntactic overhead (delimiters, indentation conventions, language-specific idioms)… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/NPset-2-Python-Edu.texttext-generation1M<n<10M11 likes376 downloads5mo agoHugging Face12colabfit /OMat24_train_aimd_from_PBE_3000_npt Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 3000 npt. ColabFit, 2025. https://doi.org/10.60732/edd12490 This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_6xvvh8yl7rfd_0 Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_3000_npt.tabular1M<n<10M0 likes344 downloads11mo agoHugging Face13ZipLime /fund-holdings-nport US Fund and ETF Holdings — Form N-PORT Every position of every US mutual fund and ETF, monthly, including the bonds, loans, asset-backed paper and derivatives that 13F does not report at all. 145 630 731 positions · 17 513 funds · 341 049 monthly reports · 2019-09-30 to 2026-05-31 The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. Why this and not 13F 13F is what everyone uses because it is what everyone knows about. It… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/fund-holdings-nport.tabulartabular-regression100M<n<1B0 likes326 downloads18d agoHugging Face14skbose /indian-english-nptel-testaudio100K<n<1M1 likes323 downloads2y agoHugging Face15NP235 /MathNet Quick Start · Overview · Tasks · Comparison · Dataset Stats · Data Sources · Pipeline · Schema · License · Citation This is the official MathNet v0. A larger version v1 will be uploaded soon (more countires, problems and richer metadata). Schema is stable but field values may be revised in v1. Quick start from datasets import load_dataset # Default: all problems ds = load_dataset("ShadenA/MathNet", split="train") # Or a specific country / competition-body config… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MathNet.imagequestion-answering10K<n<100K2 likes293 downloads3mo agoHugging Face16GeoMeterData /nphard_tsp2tabularn<1K0 likes277 downloads2y agoHugging Face17zaaabik /dolmino-mix-1124-OLMo-2-0425-1B-tokenizer-npy100M<n<1B0 likes268 downloads1y agoHugging Face18AxiomicLabs /NPset-python NPset A normalized semi-sythetic Python dataset for training small language models on code logic without the overhead of raw code syntax. Why Small language models trained on natural language corpora develop latent representations of logical constructs -- iteration, conditionals, data flow, function composition -- yet struggle to apply this reasoning to source code, where syntactic overhead (delimiters, indentation conventions, language-specific idioms) occupies a… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/NPset-python.texttext-generation1M<n<10M5 likes218 downloads6mo agoHugging Face19warisqr007 /GAPS-nptel GAPS: Golden-Aligned Parallel Speech Corpus Overview GAPS (Golden-Aligned Parallel Speech) is a multi-corpus dataset designed for foreign accent conversion. The dataset provides parallel speech triplets consisting of: Original non-native speech Parallel native speech Golden speaker speech — synthetic speech that preserves the non-native speaker’s timbre and timing (including pauses) while exhibiting native pronunciation along with the corresponding text transcript. GAPS… See the full description on the dataset page: https://huggingface.co/datasets/warisqr007/GAPS-nptel.audioaudio-to-audio1M<n<10M0 likes206 downloads6mo agoHugging Face20npvinHnivqn /URAG URAG: Uncertainty-aware RAG Evaluation Paper URAG is a framework for evaluating RAG (Retrieval-Augmented Generation) systems with conformal prediction on multiple-choice QA benchmarks. Running experiments (YAML only) First, you need to clone our GitHub repo at: https://github.com/phuvinhnguyen/URAG Install dependencies: pip install -r requirements.txt Run everything via a config file: python cli.py --config <path/to/config.yaml> Examples (LCA commit-message… See the full description on the dataset page: https://huggingface.co/datasets/npvinHnivqn/URAG.text10K<n<100K2 likes188 downloads5mo agoHugging Face21npvinHnivqn /MsVRAG-Bench MsVRAG-Bench — Multi-Step Video RAG Benchmark MsVRAG-Bench is a structured benchmark for evaluating multi-step retrieval-augmented generation (RAG) over video. Each example pairs a set of ordered video segment clips with a natural-language question that requires reasoning over which segments are visible and which are missing, simulating real RAG pipelines where a retriever may not return all relevant evidence. Quick start from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/npvinHnivqn/MsVRAG-Bench.textvisual-question-answering10K<n<100K0 likes176 downloads5mo agoHugging Face22satyanshi431 /dl-npppe2-datasettabularn<1K0 likes175 downloads3d agoHugging Face23MBZUAI /AraSeg-2026-Shared-Task-NP Arabic Sentence Segmentation Shared Task 2026 For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit: https://www.araseg.aramlab.ai/ Dataset Summary AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NP.texttoken-classificationn<1K1 likes172 downloads4mo agoHugging Face24NP235 /finemath 📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+) of mathematical educational content filtered from CommonCrawl. To curate this dataset, we trained a mathematical content classifier using annotations generated by LLama-3.1-70B-Instruct. We used the classifier to retain only the most educational mathematics content, focusing on clear explanations and step-by-step problem solving rather than… See the full description on the dataset page: https://huggingface.co/datasets/NP235/finemath.tabular10M<n<100M0 likes171 downloads3mo agoHugging Face25MBZUAI /AraSeg-2026-Shared-Task-NoPnx-NP Arabic Sentence Segmentation Shared Task 2026 For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit: https://www.araseg.aramlab.ai/ Dataset Summary AraSeg is the first comprehensive benchmark for Arabic sentence segmentation. The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy. AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-NP.texttoken-classificationn<1K1 likes148 downloads4mo agoHugging Face26sj21867 /ai_art_np Dataset Card for "ai_art_np" More Information needed text1K<n<10K0 likes144 downloads2y agoHugging Face27Madnesss /npy_file_hsrtabular1K<n<10K0 likes130 downloads2y agoHugging Face28shunanhe /NPM-Artifact-Explanation-Benchmark NPM-Artifact-Explanation-Benchmark English NPM-Artifact-Explanation-Benchmark is a cross-category multimodal corpus and benchmark resource for Chinese cultural artifact understanding and explanation. This release contains 28,826 cleaned artifact records derived from National Palace Museum source records' opendata (https://digitalarchive.npm.gov.tw/opendata/). Each record includes structured artifact metadata, image URLs, source record URLs, and human-written… See the full description on the dataset page: https://huggingface.co/datasets/shunanhe/NPM-Artifact-Explanation-Benchmark.tabularimage-to-text10K<n<100K1 likes128 downloads11d agoHugging Face29NPCI /tau-indian-banking TauIndianBankBench: an Indian retail-banking benchmark for tool-using conversational agents Data files for the indian_banking domain of the tau2-bench harness: a conversational agent staffs the chat desk of a synthetic India-based retail bank. The code that runs and scores the benchmark lives in the companion repository npci/TauIndianBankBench; this dataset holds only the domain data it loads. Viewer note. The dataset viewer shows preview/tasks_preview.parquet, which holds the… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/tau-indian-banking.texttext-generation1K<n<10K1 likes124 downloads17d agoHugging Face30NP235 /open-web-math-pro 📚 Open-Web-Math-Pro ArXiv | Models | Code Open-Web-Math-Pro is refined from open-web-math using the ProX refining framework. It contains about 5B high quality math related tokens, ready for pre-training. License Open-Web-Math-Pro is based on open-web-math, which is made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU: https://commoncrawl.org/terms-of-use/. We do not alter the license of any of the underlying data.… See the full description on the dataset page: https://huggingface.co/datasets/NP235/open-web-math-pro.texttext-generation1M<n<10M0 likes113 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.