datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WxC-Bench
Dataset Card for WxC-Bench
WxC-Bench primary goal is to provide a standardized benchmark for evaluating the performance of AI models in Atmospheric and Earth Sciences across various tasks.
Dataset Details
WxC-Bench contains datasets for six key tasks:
Nonlocal Parameterization of Gravity Wave Momentum Flux
Prediction of Aviation Turbulence
Identifying Weather Analogs
Generation of Natural Language Weather Forecasts
Long-Term Precipitation Forecasting
Hurricane Track and… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/WxC-Bench.MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset
The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html.
The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation.
The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.impulse-market-datafirst-impressions-v2
Dataset Card for First Impressions V2
The first impressions data set, comprises 10000 clips (average duration 15s) extracted from more than 3,000 different YouTube high-definition (HD) videos of people facing and speaking in English to a camera. The videos are split into training, validation and test sets with a 3:1:1 ratio. People in videos show different gender, age, nationality, and ethnicity.
Videos are labeled with personality traits variables. Amazon Mechanical Turk (AMT) was… See the full description on the dataset page: https://huggingface.co/datasets/yeray142/first-impressions-v2.self-improve-fragilitypane-binding-functions-attributionthe_stack_v2_python_repos_pretraining_dataset_imported_context-datasetimpossible_livecodebenchimpossible_swebenchimppres
Dataset Card for IMPPRES
Dataset Summary
Over >25k semiautomatically generated sentence pairs illustrating well-studied pragmatic inference types. IMPPRES is an NLI dataset following the format of SNLI (Bowman et al., 2015), MultiNLI (Williams et al., 2018) and XNLI (Conneau et al., 2018), which was created to evaluate how well trained NLI models recognize several classes of presuppositions and scalar implicatures.
Supported Tasks and Leaderboards
Natural… See the full description on the dataset page: https://huggingface.co/datasets/facebook/imppres.IMPACT
IMPACT v1.1
IMPACT is a synchronized five-view RGB-D dataset and benchmark for multi-granularity human procedural action understanding in industrial assembly. It contains 112 trials from 13 participants and 39.5 video hours across one egocentric and four exocentric views.
Project page
Benchmark code and task protocols
Google Drive mirror
Release Update
July 2026, v1.1. All 560 TAS-B annotation files were revalidated, and 117 of 560 RGB videos (20.9%) were… See the full description on the dataset page: https://huggingface.co/datasets/KratosWen/IMPACT.health_factPUBHEALTH is a comprehensive dataset for explainable automated fact-checking of
public health claims. Each instance in the PUBHEALTH dataset has an associated
veracity label (true, false, unproven, mixture). Furthermore each instance in the
dataset has an explanation text field. The explanation is a justification for which
the claim has been assigned a particular veracity label.
The dataset was created to explore fact-checking of difficult to verify claims i.e.,
those which require expertise from outside of the journalistics domain, in this case
biomedical and public health expertise.
It was also created in response to the lack of fact-checking datasets which provide
gold standard natural language explanations for verdicts/labels.
NOTE: There are missing labels in the dataset and we have replaced them with -1.theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.fm_impute_bench
Data format and usage
TODO
Dataset statistics
dataset
freq
# series
series length
# test windows
source
BDG2-Bear
1H
91
17,544
7,522
https://huggingface.co/datasets/Salesforce/lotsa_data
BDG2-Rat
1H
280
17,544
24,915
https://huggingface.co/datasets/Salesforce/lotsa_data
Borealis
1H
15
7,447
77
https://huggingface.co/datasets/Salesforce/lotsa_data
Covid19 Energy
1H
1
31,912
195
https://huggingface.co/datasets/Salesforce/lotsa_data
GFC12 Load
1H… See the full description on the dataset page: https://huggingface.co/datasets/taharnbl/fm_impute_bench.opm-ehri-dataarca-importaciones-argentina
ARCA Importaciones Argentina
Dataset público derivado de la Información Agregada de Comercio Exterior publicada por ARCA Argentina.
Cobertura
El objetivo histórico abarca todos los meses publicados por ARCA desde 02/2017 hasta 08/2026.
Dos granularidades reales de ARCA
ARCA no mantuvo el mismo formato durante todo el período. El pipeline detecta el encabezado de cada ZIP y conserva la semántica correcta:
data/items/YYYYMM.parquet: meses con… See the full description on the dataset page: https://huggingface.co/datasets/alexbozz1/arca-importaciones-argentina.ChatDoctor-HealthCareMagic-Output-Improved-GPT4.1design-patents-not-in-impact
US Design Patents Not Included in IMPACT (2008-2026)
Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are
absent from the AI4Patents/IMPACT dataset.
IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that
IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period,
plus 4,824 patents from years IMPACT does cover but did not include. There is no patent
overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.notch-beam-2d-impact
NotchBeam2D-Impact — StructBench canonical dataset
Download
One case, one file — fetch exactly what you need (pip install huggingface_hub):
from huggingface_hub import hf_hub_download, snapshot_download
# one case
path = hf_hub_download("StructBench/notch-beam-2d-impact",
filename="<case_id>.h5", repo_type="dataset")
# the full archive (resumable; cached under HF_HOME)
root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.copying_reasoning_task_improved
copying_reasoning_task_improved
Dataset Description
The Enhanced Copying Reasoning Task Dataset is designed to provide a rich resource for analyzing promotional texts and their key elements. This dataset includes a variety of question-and-answer formats, focusing on whether specific phrases are mentioned within the text. Its purpose is to assist in the training of models for natural language understanding tasks, particularly in identifying relevant information in… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/copying_reasoning_task_improved.AgiBotWorld-Beta_G1_task_765_Adjust_implantable_advertising
agibot_task_765
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 调整植入式广告
total_episodes: 1493
total_tasks: 1
size: 88G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├── observation.images.back_left_fisheye… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_765_Adjust_implantable_advertising.ImpactMesh-Fire
ImpactMesh-Fire
ImpactMesh is a large-scale multimodal, multitemporal dataset for flood and wildfire mapping, released by IBM, DLR, and the ESA Φ-lab.
It integrates Sentinel-1 SAR, Sentinel-2 optical, Copernicus DEM, and high-quality annotations from Copernicus EMS.
The technical report is released soon. You find the flood subset here: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Flood.
Features
Multimodal: SAR, optical, DEM
Multitemporal:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Fire.IMPACTpython-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.laion_improved_aesthetics_6.5plus_with_imagesbhl-impact-gt
FineBooks BHL IMPACT Ground Truth
2,165 page scans from six historical natural-history books, each paired with an expert, ~99.95%-accurate transcription and full page-layout ground truth. A benchmark for OCR, text recognition, and document layout analysis on real historical print.
This dataset is the basis of the BHL OCR Leaderboard, where open OCR models are scored against these transcriptions. As new OCR models are released, they are run through the same evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/finebooks/bhl-impact-gt.paper-impact-datasentry-impact-risk
NASA Sentry: Earth Impact Risk Assessment
Credit: NASA/Johns Hopkins APL
Part of a dataset collection on Hugging Face.
Dataset description
Near-Earth objects with non-zero Earth impact probability from NASA JPL Sentry system.
The Sentry system, operated by NASA's Center for Near-Earth Object Studies (CNEOS) at the Jet Propulsion Laboratory, continuously monitors the most current asteroid catalog for possibilities of future Earth impact. Objects are listed… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/sentry-impact-risk.ImpactMesh-Flood
ImpactMesh-Flood
ImpactMesh is a large-scale multimodal, multitemporal dataset for flood and wildfire mapping, released by IBM, DLR, and the ESA Φ-lab.
It integrates Sentinel-1 SAR, Sentinel-2 optical, Copernicus DEM, and high-quality annotations from Copernicus EMS.
The technical report is released soon. You find the wildfire subset here: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Fire.
Features
Multimodal: SAR, optical, DEM… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/ImpactMesh-Flood.libero-kinex-v0.6.0-vs-kinex-v0.6.0-improve-vs-codexEvaluation website · Data format
kinex (v0.6.0) vs kinex (v0.6.0-improve) vs codex · LIBERO Long · GPT-6 Astra / medium
45 planned episodes. Counts below are derived from the episode index.
Variant
Native successes / valid
Normal successes / valid
Interrupted
Missing
kinex (v0.6.0)
9/15
9/15
0
0
kinex (v0.6.0-improve)
9/15
9/15
0
0
codex
6/15
6/15
0
0
Native-valid counts include interrupted executions. Normal counts also require a finished execution and valid… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/libero-kinex-v0.6.0-vs-kinex-v0.6.0-improve-vs-codex.
