datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image_dummy\WxC-Bench
Dataset Card for WxC-Bench
WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks.
Dataset Details
WxC-Bench contains datasets for six key tasks:
Nonlocal Parameterization of Gravity Wave Momentum Flux
Prediction of Aviation Turbulence
Identifying Weather Analogs
Generation of Natural Language Weather Forecasts
Long-Term Precipitation Forecasting
Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.natural_questions
Dataset Card for Natural Questions
Dataset Summary
The NQ corpus contains questions from real users, and it requires QA systems to
read and comprehend an entire Wikipedia article that may or may not contain the
answer to the question. The inclusion of real user questions, and the
requirement that solutions should read an entire page to find the answer, cause
NQ to be a more realistic and challenging task than prior QA datasets.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.nbm-conus-analysis
NOAA NBM CONUS Daily Analysis (Zarr)
Daily best-estimate analysis derived from NOAA NBM (National Blend of Models)
CONUS forecasts, on the native ~2.5 km Lambert conformal grid (2345 x 1597).
Built by nbm-to-zarr, dynamical.org-style.
Variables: tmean / tmax / tmin (degC), precip (mm), srad (MJ/m2/day)
Construction: best estimate for day D = lead-day 1 of that day's 00z NBM init
Coverage: rolling backfill from 2020-10-01 (AWS NBM archive floor) to present
Layout: one standalone… See the full description on the dataset page: https://huggingface.co/datasets/nakas/nbm-conus-analysis.s2-naipAI2-S2-NAIP is a remote sensing dataset consisting of aligned NAIP, Sentinel-2, Sentinel-1, and Landsat images spanning the entire continental US.
Data is divided into tiles.
Each tile spans 512x512 pixels at 1.25 m/pixel in one of the 10 UTM projections covering the continental US.
At each tile, the following data is available:
National Agriculture Imagery Program (NAIP): an image from 2019-2021 at 1.25 m/pixel (512x512).
Sentinel-2 (L1C): between 16 and 32 images captured within a few… See the full description on the dataset page: https://huggingface.co/datasets/allenai/s2-naip.nawaqes-backup-v2Multitask-National-Speech-Corpus-v1Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus.
MNSC is a multitask speech understanding dataset derived and further annotated from IMDA NSC Corpus. It focuses on the knowledge of Singapore's local accent, localised terms, and code-switching.
ASR: Automatic Speech Recognition
SQA: Speech Question Answering
SDS: Spoken Dialogue Summarization
PQA: Paralinguistic Question Answering
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1.nara_revolutionary_war_pension_files_PDFs
Dataset Card for American Revolutionary War Pension Files - File-Level
Dataset Summary
A dataset derived from the National Archives and Records Administration (NARA) series Case Files of Pension and Bounty-Land Warrant Applications Based on American Revolutionary War Service (NARA Catalog Series, NAID 300022). This dataset provides a file-level representation of Revolutionary War pension records, aggregating individual page records into complete pension files… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files_PDFs.nano-banana-pro-prompts-datasets
🖼️ Nano Banana Pro Prompt Dataset
🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.NatureBench
Dataset Card for NatureBench
NatureBench is a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, spanning 6 scientific domains. It is designed to evaluate whether AI coding agents can move beyond reproduction toward discovery: each task asks an agent to solve a real scientific machine-learning problem and is scored against the source paper's reported state of the art.
📄 arXiv paper: https://arxiv.org/abs/2606.24530
💻 GitHub code… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/NatureBench.NatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.nam19955truongtien2002translations-raw
natgillin/translations-raw
Frozen, canonical raw bitext consolidated from upstream alvations/mtdata-raw* snapshots (since deleted). This is the read-only source-of-truth for downstream quality-filtering pipelines.
31,663 parquet files (1566.8 GB)
49 language pairs under data/<src-tgt>/
Schema: 5 columns — see below
Read-only for downstream pipelines. Do not delete or modify.
Schema
Each parquet has 5 columns:
column
type
description
source
string… See the full description on the dataset page: https://huggingface.co/datasets/natgillin/translations-raw.natural-instructionsPreprocessed version of Super-Natural-Instructions from https://github.com/allenai/natural-instructions/tree/master/splits. The same inputs may appear with different outputs, thus to avoid duplicate inputs, you can deduplicate by the id or the inputs field.
Train Tasks:
['task001_quoref_question_generation', 'task002_quoref_answer_generation', 'task022_cosmosqa_passage_inappropriate_binary', 'task023_cosmosqa_question_generation', 'task024_cosmosqa_answer_generation'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/natural-instructions.imagenet-sketch-datanaturalscenesdatasettruongtien2002hubbleRAEv2-data
RAEv2 Data
Pre-processed datasets and pretrained encoders for RAEv2: Improved Baselines with Representation Autoencoders. All rights to the original owners; per-subset attribution below.
Repo Structure
RAEv2-data/
|-- imagenet-256/ # ImageNet-1k at 256x256 (Arrow)
|-- blip3o-256/ # BLIP3o captioned images (WDS)
|-- render-text-256/ # Rendered-text images (WDS)
|-- scale-rae-256/ # Synthetic FLUX images (WDS)
|-- recon-256/ # Robot… See the full description on the dataset page: https://huggingface.co/datasets/nanovisionx/RAEv2-data.Nanochatnarrativeqa
Dataset Card for Narrative QA
Dataset Summary
NarrativeQA is an English-lanaguage dataset of stories and corresponding questions designed to test reading comprehension, especially on long documents.
Supported Tasks and Leaderboards
The dataset is used to test reading comprehension. There are 2 tasks proposed in the paper: "summaries only" and "stories only", depending on whether the human-generated summary or the full story text is used to answer the question.… See the full description on the dataset page: https://huggingface.co/datasets/deepmind/narrativeqa.traditionnals_arabic_shoes_splitruri-dataset-v2-ptWIP: 正式公開準備中
各データセットのライセンスは元データセットに従います。
cord-v2nawaqes-backupasr_dummySelf-supervised learning (SSL) has proven vital for advancing research in
natural language processing (NLP) and computer vision (CV). The paradigm
pretrains a shared model on large volumes of unlabeled data and achieves
state-of-the-art (SOTA) for various tasks with minimal adaptation. However, the
speech processing community lacks a similar setup to systematically explore the
paradigm. To bridge this gap, we introduce Speech processing Universal
PERformance Benchmark (SUPERB). SUPERB is a leaderboard to benchmark the
performance of a shared model across a wide range of speech processing tasks
with minimal architecture changes and labeled data. Among multiple usages of the
shared model, we especially focus on extracting the representation learned from
SSL due to its preferable re-usability. We present a simple framework to solve
SUPERB tasks by learning task-specialized lightweight prediction heads on top of
the frozen shared model. Our results demonstrate that the framework is promising
as SSL representations show competitive generalizability and accessibility
across SUPERB tasks. We release SUPERB as a challenge with a leaderboard and a
benchmark toolkit to fuel the research in representation learning and general
speech processing.
Note that in order to limit the required storage for preparing this dataset, the
audio is stored in the .flac format and is not converted to a float32 array. To
convert, the audio file to a float32 array, please make use of the `.map()`
function as follows:
```python
import soundfile as sf
def map_to_array(batch):
speech_array, _ = sf.read(batch["file"])
batch["speech"] = speech_array
return batch
dataset = dataset.map(map_to_array, remove_columns=["file"])
```zen_truyenModelNet40_Auto_aligned
ModelNet40 Auto Aligned
Auto-aligned version of the ModelNet40 3D CAD dataset. Each sample is an OFF mesh file organized by class and train/test split.
This dataset mirrors the layout of naderalfares/ModelNet40, but uses the auto-aligned meshes from the Princeton ModelNet release.
Dataset structure
modelnet40_auto_aligned/
{class}/
train/{class}_{id}.off
test/{class}_{id}.off
40 classes (airplane, bathtub, bed, …, xbox)
9,843 training meshes
2,468 test… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/ModelNet40_Auto_aligned.Multitask-National-Speech-Corpus-v1-extend
