datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NotaGenX-opusThis dataset is generated by NotaGenX model.
Thanks to ElectricAlexis!
abc/
This folder contains pure ABC Notation files.
It is intended to include up to 1 million score pieces.
Sorry for sub directories splitting, but HuggingFace limits a single directory direct sub items number up to 10000.
You can rearrange them by mv abc/*/*/* ./abc/.
Dataset Distribution
Component frequency over all 993,183 .abc pieces in abc/ (parsed from the leading %Period / %Composer /… See the full description on the dataset page: https://huggingface.co/datasets/k-l-lambda/NotaGenX-opus.math-contests-2026
Math Contests 2026 (🔗 notadib/math-contests-2026)
197 problems from national olympiads and team-selection tests held January 2026 and onward — a held-out benchmark for math reasoning, sourced after the contests ran but before solutions were widely propagated, so they should not appear in any current LLM training data.
Excluded: any contest held in 2025 — BMO Round 1 (Nov 2025), USA TSTST, USA TST (Dec 2025) and Bundeswettbewerb Mathematik (Dec 2025) — kept strictly to events… See the full description on the dataset page: https://huggingface.co/datasets/notadib/math-contests-2026.cdnNASA-Power-Daily-Weather
NASA Power Weather Data over North, Central, and South America from 1984 to 2022
This dataset contains daily solar and meteorological data downloaded from the NASA Power API
Overview
The dataset includes solar and meteorological variables collected from January 1st, 1984, to December 31st, 2022.
We downloaded 28 variables directly and estimated an additional 3 from the collected data. The data spans a 5 x 8 grid covering
the United States, Central America, and South… See the full description on the dataset page: https://huggingface.co/datasets/notadib/NASA-Power-Daily-Weather.nota
Dataset Card for Nota
Dataset Summary
This data was created by the public institution Nota, which is part of the Danish Ministry of Culture. Nota has a library audiobooks and audiomagazines for people with reading or sight disabilities. Nota also produces a number of audiobooks and audiomagazines themselves.
The dataset consists of audio and associated transcriptions from Nota's audiomagazines "Inspiration" and "Radio/TV". All files related to one reading of one edition… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nota.AICamp-2023-Skin-Conditions-Datasetnotas-comunidade-ptbr
Community Notes BR: Enriched Portuguese Dataset for NLP and Text Mining
Community Notes BR is a curated Portuguese-language subset of X Community Notes enriched for Natural Language Processing (PLN), text mining, and research on collaborative misinformation moderation. It combines note-level metadata, topic and macrotheme labels, named entities, source-domain extraction, Matrix Factorization scoring fields, an optional Universal Dependencies syntax layer, and automatic… See the full description on the dataset page: https://huggingface.co/datasets/histlearn/notas-comunidade-ptbr.transcoda-random-notation-300k
Transcoda Random Notation 300k
Source: train: cminst/transcoda-random-notation-normalized-v1/uniform-v3-normalized-constant-spines-marks-300k; validation: fresh generation using the same normalized uniform-v3 generator recipe
This is a canonical standardized Transcoda dataset. All published splits use the standard Hugging Face Datasets Parquet layout under data/.
Canonical columns:
image
transcription
sample_id
source
metadata: JSON string preserving source/provenance fields… See the full description on the dataset page: https://huggingface.co/datasets/cminst/transcoda-random-notation-300k.ACYD
ACYD: Agricultural Crop Yield Dataset
Weekly, admin-2 level covariates for Argentina, Brazil, and Mexico (1979–2024), designed for crop yield prediction and benchmarking.
Code & processing pipeline: https://github.com/Neehan/amazing-crop-yield-datasets
The same scripts also support USA — only the --country flag changes. USA data is not included in this release but can be reproduced from the GitHub repo.
Files
This repository contains three directories per country… See the full description on the dataset page: https://huggingface.co/datasets/notadib/ACYD.towel_unfold_notape_v1_retry_20260914This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Woohi123/towel_unfold_notape_v1_retry_20260914.NOTA-datasetNotAllCodeIsEqual
NotAllCodeIsEqual
This dataset was created for the paper Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning.
It contains code fine-tuning datasets split by complexity metrics for studying the relationship between code complexity and reasoning capabilities.
We provide 2 types of dataset, that cover complementary settings:
CodeNet (solution-driven complexity):
The CodeNet splits contain the same programming problems across all complexity levels, but with… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/NotAllCodeIsEqual.usa-corn-belt-crop-yield
USA County Level Crop Yield Dataset
Dataset Summary
This dataset contains county level crop yield across 763 counties from 1984 till 2018 in the US Corn Belt. The data was originally collected in Khaki et al. 2020, then further processed, augmented dedup-ed in Hasan et al. 2026.
Here are the 9 unique states in the dataset:
Illinois
Indiana
Iowa
Kansas
Minnesota
Missouri
Nebraska
North Dakota
South Dakota
Each row of the CSV includes:
Weather: 6 weekly mean weather… See the full description on the dataset page: https://huggingface.co/datasets/notadib/usa-corn-belt-crop-yield.details_NotAiLOL__Boundary-Meta-Llama-3-2x8B-MoE
Dataset Card for Evaluation run of NotAiLOL/Boundary-Meta-Llama-3-2x8B-MoE
Dataset automatically created during the evaluation run of model NotAiLOL/Boundary-Meta-Llama-3-2x8B-MoE on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NotAiLOL__Boundary-Meta-Llama-3-2x8B-MoE.NuminaMath-CoT-Small-215k
Summary
This dataset is a scaled down version of the original AI-MO/NuminaMath-CoT dataset.
Source breakdown
Source
Number of Originial Samples
Number of Samples in This Dataset
aops_forum
30201
7548
amc_aime
4072
1017
cn_k12
276591
69138
gsm8k
7345
1835
math
7478
1869
olympiads
150581
37640
orca_math
153334
38328
synthetic_amc
62111
15527
synthetic_math
167895
41968
Total
859608
214870
towel_unfold_half_fold_bimanual_notape_v1_20260914This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Woohi123/towel_unfold_half_fold_bimanual_notape_v1_20260914.details_NotAiLOL__Yi-1.5-dolphin-9B
Dataset Card for Evaluation run of NotAiLOL/Yi-1.5-dolphin-9B
Dataset automatically created during the evaluation run of model NotAiLOL/Yi-1.5-dolphin-9B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_NotAiLOL__Yi-1.5-dolphin-9B.EDTBtranscoda-random-notation-shardsNuminaMath-CoT-Small-Hard-200k
Summary
This dataset is a scaled down version of the original AI-MO/NuminaMath-CoT dataset with more focus on hard math.
Source breakdown
Source
Number of Originial Samples
Number of Samples in This Dataset
aops_forum
30201
5000
amc_aime
4072
4070
cn_k12
276591
55310
gsm8k
7345
1000
math
7478
1000
olympiads
150581
37640
orca_math
153334
30662
synthetic_amc
62111
31054
synthetic_math
167895
33574
Total
859608
199310
details_NotAiLOL__Apollo-7b-orpo-Experimentaldetails_NotAiLOL__Athena-zephyr-7Blaouenan-notable-peopleLaouenan, M., Bhargava, P., Eymeoud, J.-B., Gergaud, O., Plique, G., & Wasmer, E. (2023). A Brief History of Human Time - Cross-verified Dataset. data.sciencespo. doi: 10.21410/7E4/RDAG3O
notabug-code
NotaBug Code Dataset
Dataset Description
This dataset was compiled from code repositories hosted on NotaBug.org, a free code hosting platform that emphasizes software freedom and privacy. NotaBug is built on a fully free software stack and is popular among free software advocates and privacy-conscious developers.
Dataset Summary
Statistic
Value
Total Files
12,622,961
Total Repositories
11,660
Total Size
12 GB (compressed Parquet)
Programming… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/notabug-code.afvoices-notag
AfVoices Top-20 Speakers without Tags
RobotsMali/afvoices-notag is a small experimental TTS-oriented selection derived from RobotsMali/afvoices, the African Next Voices Bambara speech corpus. It contains the 20 participants with the highest utterance counts and excludes transcripts containing semantic/acoustic annotation tags.
This is the dataset used for RobotsMali's first Bambara VITS experiments. It is not a high-quality studio TTS corpus: the source is spontaneous speech… See the full description on the dataset page: https://huggingface.co/datasets/RobotsMali/afvoices-notag.uploadscgu__notas_fiscais
Dataset Card: cgu_notas_fiscais
Data from electronic invoices for federal government purchases made available by
Comptroller General of the Union (Controladoria-Geral da União), which is a
Brazilian federal government agency responsible for oversight and transparency.
Dataset Details
Dataset Description
Curated by: Fred Guth (@fredguth)
Funded by: World Bank
Language(s) (NLP): pt-br
License: CC-BY 4.0
Dataset Sources
The source of this datasets… See the full description on the dataset page: https://huggingface.co/datasets/fredguth/cgu__notas_fiscais.synapse_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 84,
"total_frames": 25311,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:84"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/NotARoomba/synapse_5.so101_pick_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/nota-gmbh/so101_pick_place.pause63-echo.notag.r-1.0-k-8.statml-arxiv
