datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parti-prompts
Dataset Card for PartiPrompts (P2)
Dataset Summary
PartiPrompts (P2) is a rich set of over 1600 prompts in English that we release
as part of this work. P2 can be used to measure model capabilities across
various categories and challenge aspects.
P2 prompts can be simple, allowing us to gauge the progress from scaling. They
can also be complex, such as the following 67-word description we created for
Vincent van Gogh’s The Starry Night (1889):
Oil-on-canvas painting of a… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/parti-prompts.us-names-by-state
US Baby names
The SSA dataset with baby names:
https://www.ssa.gov/OACT/babynames/
Coniferest
We use this dataset in the active anomaly discovery Python package coniferest:
https://coniferest.snad.space/en/latest/notebooks/us-names.html
Update the data
Install Python packages: pip install requests aiohttp universal_pathlib pandas
Optionally: download https://www.ssa.gov/OACT/babynames/state/namesbystate.zip
./run.py PATH_OR_URL_TO_namesbystate.zip, path may be… See the full description on the dataset page: https://huggingface.co/datasets/snad-space/us-names-by-state.predictive-stock-datasetEnvironment-and-Natural-Resources-Indicators-For-African-Countries
Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.picotron_bench
Wrapup results:
compute mfu for each results
change status of jobs
Push to hub
add scripts reproductible
add topology
bandwidth etc
nan-nli
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
Natural Language Inference
Text Classification
Languages
en
Dataset Structure
Data Instances
Data Fields
premise:
hypothesis:
label:
Data Splits
Evaluation: 258 samples
Dataset Creation
Curation Rationale
Extracting samples corresponding to different linguistics constructions of… See the full description on the dataset page: https://huggingface.co/datasets/joey234/nan-nli.nangang_sports_centernaver-news-summarization-ko
Naver-News-KO: A Korean News Summarization Dataset
A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from
Naver News over a ten-day window in July 2022. It was originally built for a
Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023.
A technical report documenting the collection protocol, corpus statistics, contamination analysis, and
reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.genter-ajibawa-name-filled
GENTER Ajibawa Name-Filled
This dataset expands aieng-lab/genter-ajibawa by inserting concrete names into each template.
It provides nested Hugging Face configs with 1, 2, 5, or 10 names per gender and template.
For every template sentence, names are sampled independently from NAMEXACT (matching split; frequency-weighted), using K female and K male names in config nK.
It is intended for experiments that need concrete text rather than [NAME]/[MASK] placeholders while still… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/genter-ajibawa-name-filled.us-accidents
Dataset Card for US Accidents (2016 - 2021)
Dataset Summary
Description
This is a countrywide car accident dataset, which covers 49 states of the USA. The accident data are collected from February 2016 to Dec 2021, using multiple APIs that provide streaming traffic incident (or event) data. These APIs broadcast traffic data captured by a variety of entities, such as the US and state departments of transportation, law enforcement agencies, traffic cameras, and… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/us-accidents.narrow-model-safety-eval
Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset
Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.NACA_4_Digit_for_ML
NACA 4-Digit Airfoil CFD Dataset
Point-cloud CFD solutions for NACA 4-digit airfoils, generated with OpenFOAM v13 (k-ω SST). Intended for training surrogate models that predict steady-state flow fields from airfoil geometry and flow conditions.
Dataset Summary
~850 converged in-distribution cases across 50 distinct NACA 4-digit profiles
AoA range: −5° to +5°
Reynolds number range: 100,000 – 500,000
129 out-of-distribution (OOD) probe cases at high Re (1–2 × 10⁶)… See the full description on the dataset page: https://huggingface.co/datasets/Kokoslocke/NACA_4_Digit_for_ML.hungarian_national_hs_finals_exam
Testing Language Models on a Held-Out High School National Finals Exam
When xAI recently released Grok-1, they evaluated it on the 2023 Hungarian national high school finals in mathematics, which was published after the training data cutoff for all the models in their evaluation. While MATH and GSM8k are the standard benchmarks for evaluating the mathematical abilities of large language models, there are risks that modern models overfit to these datasets, either from training… See the full description on the dataset page: https://huggingface.co/datasets/keirp/hungarian_national_hs_finals_exam.nairobi-longitudinal-healthcare-utilization
Nairobi Longitudinal Healthcare Utilization
Synthetic longitudinal population data for predicting healthcare use across time.
100% synthetic. No real patient records. Every resident, household, insurance record,
employment state, education history, encounter, diagnosis, prescription, condition, and
healthcare outcome in this release is synthetic. No real individual's data was used to produce it.
What is this?
This dataset is a balanced longitudinal… See the full description on the dataset page: https://huggingface.co/datasets/stripeddonkey-data/nairobi-longitudinal-healthcare-utilization.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.russian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades
Data Statement for SHADES
How to use this document:
Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.nairobi-tertiary-graduate-employment-outcomes
Nairobi Tertiary Graduate Employment Outcomes
Synthetic graduate-employment data for predicting how a tertiary completion translates into a job.
100% synthetic. No real student or worker records. Every citizen, enrollment, household,
and employment episode in this dataset is synthetic. No real individual's data was used to
produce this release.
What is this?
This dataset contains 13,105 certified completed tertiary-education episodes from a synthetic
Nairobi… See the full description on the dataset page: https://huggingface.co/datasets/stripeddonkey-data/nairobi-tertiary-graduate-employment-outcomes.nasdaq_datanasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/penikmatrumput/nasa-cmapss-rul.harbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.muslim-names-dataset
Muslim Names Dataset
A comprehensive collection of Muslim names with meanings scraped from muslimnames.com. Contains 14,585 names with English names, Arabic names, meanings, and gender classifications.
Dataset Contents
This dataset contains ~14,585 Muslim names with the following information:
English name: Name in English/Latin script
Arabic name: Name in Arabic script
Meaning: Definition and meaning of the name
Gender: Classification as male or female
Files… See the full description on the dataset page: https://huggingface.co/datasets/takiuddinahmed/muslim-names-dataset.German_Names_Central_And_Eastern_EuropeThis dataset contains German exonyms for various places in modern day Poland, Czech Republic, Latvia, Lithuania and Estonia.
Exonym : - A placename that is used by people who are not locals. For example, Prague is the Eng. exonym of Czech capital Praha, or Cologne is an exonym for German city Köln.
Due to extensive historical German rule and presence over large chunks of modern day Poland and Czech republic, these two countries populate the dataset the most.
SpatialConsistency-Navigationsurya-bench-ar-segmentation
A Dataset of Binary Maps of Active Regions with Polarity Inversion Lines
Dataset Summary
This dataset provides hourly binary segmentation maps (4096×4096 resolution) derived from Solar Dynamics Observatory (SDO) / Helioseismic and Magnetic Imager (HMI) line-of-sight magnetograms. The maps highlight regions containing Active Regions (ARs) and Polarity Inversion Lines (PILs). The dataset spans observations from May 13, 2010 to December 31, 2024 and is intended for image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-ar-segmentation.naive-physics-ironing-v0.2
nAIve physics — Ironing Pilot v0.2 + Interaction Analysis v0.3
Visual Preview
Original RGB demonstration — IRON_009
▶ Watch IRON_009 original RGB demonstration
v0.3 interaction analysis — IRON_009
▶ Watch IRON_009 analysed interaction video
Raw → analysed: the first video is the original RGB demonstration; the second shows the v0.3 garment semantics, tool tracking, and temporal interaction analysis derived from the same episode.
A… See the full description on the dataset page: https://huggingface.co/datasets/CaramelCoffee19/naive-physics-ironing-v0.2.gleif_nameNASA_Nearest_Earth_Objects_1910-2024CONTEXT:
There are many dangerous bodies in space, one of them is N.E.O. - "Nearest Earth Objects". Some such bodies really pose a danger to the planet Earth, NASA classifies them as "is_hazardous". This dataset contains ALL NASA observations of similar objects from 1910 to 2024!!!
There are 338,199 records of N.E.O. in the Dataset!
Try to predict "is_hazardous" as accurately as possible! (otherwise we will not be ready for an asteroid attack)
SOURCES:
NASA Open API: https://api.nasa.gov/… See the full description on the dataset page: https://huggingface.co/datasets/IvanSher/NASA_Nearest_Earth_Objects_1910-2024.ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.
