datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
srtm30m-ozt2-v2
SRTM 30m OZT2 Elevation Tiles
This dataset contains SRTM 30-meter resolution elevation data encoded in the OZT2 tile format.
Format
OZT2 is a high-performance elevation tile format:
Compression: ~93% smaller than Terrarium PNG
Prediction: Gradient-based prediction (left neighbor + vertical gradient)
Quantization: Adaptive bit-depth (8/10/12/16-bit per channel)
Codec: Zstd q3 (30× faster encode than Brotli, same decode speed)
Each tile is 256×256 pixels in Web… See the full description on the dataset page: https://huggingface.co/datasets/aliasfox/srtm30m-ozt2-v2.TFUScapes
tags:
- medical
pretty_name: tfuscapes
task_categories:
- other
language:
- en
size_categories:
- 1K<n<10K
modalities:
- npz
A Skull-Adaptive Framework for AI-Based 3D Transcranial Focused Ultrasound Simulation
Vinkle Srivastav, Juliette Puel, Jonathan Vappou, Elijah Van Houten, Paolo Cabras*, Nicolas Padoy*
*co-last authors
Access the paper
Introduction
Transcranial focused ultrasound (tFUS) is an emerging modality for non-invasive… See the full description on the dataset page: https://huggingface.co/datasets/vinkle-srivastav/TFUScapes.MapPool
MapPool - Bubbling up an extremely large corpus of maps for AI
MapPool is a dataset of 75 million potential maps and textual captions. It has been derived from CommonPool, a dataset consisting of 12 billion text-image pairs from the Internet. The images have been encoded by a vision transformer and classified into maps and non-maps by a support vector machine. This approach outperforms previous models and yields a validation accuracy of 98.5%. The MapPool dataset may help to train… See the full description on the dataset page: https://huggingface.co/datasets/sraimund/MapPool.srtm30m-mergedlatent-sr-embeddings
Latent-SR Embeddings: Precomputed VAE Latents for Medical Image Super-Resolution
Precomputed VAE latent embeddings from the paper:
"Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution"Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi.arXiv:2604.12152 (2026)… See the full description on the dataset page: https://huggingface.co/datasets/sebasmos/latent-sr-embeddings.srtm-global-void-filledThis dataset mirrors the
Shuttle Radar Topography Mission (SRTM) Void Filled digital elevation data from USGS.
It consists of the 15,417 GeoTIFFs available on USGS EarthExplorer in the "SRTM Void Filled" (srtm_v2) dataset.
Each GeoTIFF covers 1x1 degrees.
The data is in WGS84, with a resolution of 1 arc-second/pixel in the United States and 3 arc-seconds/pixel elsewhere.
Coverage is limited to "80% of the Earth's land surface between 60° north and 56° south latitude".
The data is attributed to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/srtm-global-void-filled.srankmonsternobehemothdakedonekotomachigawareteelfmusumenopettoshitekurashitemasu
Bangumi Image Base of S-rank Monster No "behemoth" Dakedo, Neko To Machigawarete Elf Musume No Pet Toshite Kurashitemasu
This is the image base of bangumi S-Rank Monster no "Behemoth" dakedo, Neko to Machigawarete Elf Musume no Pet toshite Kurashitemasu, we detected 56 characters, 4649 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/srankmonsternobehemothdakedonekotomachigawareteelfmusumenopettoshitekurashitemasu.ipfs_srilanka_laws
Laws of Sri Lanka
Research snapshot of official legislation collected from documents.gov.lk / parliament.lk.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-26
Coverage
catalog-backed incomplete
Source
documents.gov.lk / parliament.lk
Collector
scrapers/collect_lk.py
Laws / instruments
6567
Articles
5290
Language
en
Jurisdiction
Sri Lanka
License… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_srilanka_laws.srp-staging-controlled
srp-staging-controlled
Flat asset repo of Simready Asset Packages.
Branches
main — the published lane. Holds
.packages/simready.hf.nvidia.srp-staging-controlled.<asset>/<YYYY.MM.DD_NN>/
and nothing else.
from_ovstorage — the corpus assets are selected from, mirrored from rc.storage
under main/simready.ov/simready_content/assets_staging.
ticket_<user>_<timestamp> — one per submission, created at submit and deleted
when the ticket closes.
Versions
A… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/srp-staging-controlled.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.Ocean
OCEAN Big Five Personality Dataset
Dataset Summary
OCEAN Big Five Personality Dataset contains structured text samples and questionnaires annotated with OCEAN (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) Big Five personality trait scores.
Dataset Structure
Language: English (en)
Fields: Text prompts, personality trait ratings.
DORI-ONCWarning: This is a pre-release version. Labels in this dataset are subject to change without notice.
This is a marine mammal vocalisation dataset primarily focused on the southern resident killer whale habitat.
The portion of this data comes from Ocean Networks Canada, and is shared under a CC-BY License.
Alternatively, the data can be downloaded from Ocean Networks Canada Directly:
https://data.oceannetworks.ca/SearchHydrophoneData
Marine mammal detections are conducted by an amature labeller… See the full description on the dataset page: https://huggingface.co/datasets/DORI-SRKW/DORI-ONC.MMLU-SR
MMLU-SR Dataset
This is the dataset for the paper "MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models".
Dataset Structure
This dataset contains three different variants:
Question Only: Key terms in questions are replaced with dummy words and their definitions, while answer choices remain unchanged.
Answer Only: Key terms in answer choices are replaced with dummy words and their definitions, while questions remain unchanged.
Question… See the full description on the dataset page: https://huggingface.co/datasets/NiniCat/MMLU-SR.CountQA
Dataset Summary
CountQA is the new benchmark designed to stress-test the Achilles' heel of even the most advanced Multimodal Large Language Models (MLLMs): object counting. While modern AI demonstrates stunning visual fluency, it often fails at this fundamental cognitive skill, a critical blind spot limiting its real-world reliability.
This dataset directly confronts that weakness with over 1,500 challenging question-answer pairs built on real-world images, hand-captured to feature… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/CountQA.skillevo-trajectoriesScicode-test-data-h5llm-srbench
LLM-SRBench: Benchmark for Scientific Equation Discovery with LLMs
We introduce LLM-SRBench, a comprehensive benchmark with 239 challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization.
Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorization… See the full description on the dataset page: https://huggingface.co/datasets/nnheui/llm-srbench.SR-Ground
SR-Ground Dataset and Image Quality Grounding Models
This repository accompanies the paper SR-Ground: Image Quality Grounding for Super-Resolved Content and provides the SR-Ground dataset, proposed image quality grounding models inference code and weights.
Repository Structure
datasets/Contains all images. Each sample is located in a folder named according to the pattern: datasets/<sr>_<preset>/
<sr> – name of the Super‑Resolution method used for upscaling.… See the full description on the dataset page: https://huggingface.co/datasets/Divotion/SR-Ground.Mixkit-Srcsr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09-lossless-3
SR-HOME-MEA_001-UN-BATHROOM_C004_L1_F1
High-quality synthetic render data-pack featuring a detailed residential bathroom scene, captured from Camera_04 with 100 rendered outputs and 1,200 files across 12 production-ready render passes.
This dataset is designed for computer vision, 3D perception, material analysis, segmentation, depth estimation, and synthetic data workflows. It includes beauty renders plus rich auxiliary passes such as albedo, depth, normals, UVW, material index… See the full description on the dataset page: https://huggingface.co/datasets/DaiPatrick/sr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09-lossless-3.srt-coco-thumbsICDAR2019-SROIE
ICDAR2019's Scanned Receipts OCR and Information Extraction (SROIE)
The ICDAR2019 SROIE dataset was originally published by Huang et al. for the
15th International Conference on Document Analysis and Recognition (ICDAR2019)
Robust Reading Challenge on Scanned Receipts OCR and Information Extraction
(SROIE).
This work presents an extension of the original ICDAR2019 SROIE dataset, including 14
receipt annotations missing from the original Task 3 test dataset, in a format
integrated… See the full description on the dataset page: https://huggingface.co/datasets/jsdnrs/ICDAR2019-SROIE.Sewer-pipe-defectsAndroid-Malware-Datasetsrsd-feynman_hard
Dataset Card for SRSD-Feynman (Hard set)
Dataset Summary
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
We carefully reviewed the properties of each formula and its variables in the Feynman Symbolic Regression Database to design reasonably realistic sampling range of values so that our SRSD datasets can be used for evaluating the potential of SRSD such as whether or not an SR method con (re)discover… See the full description on the dataset page: https://huggingface.co/datasets/yoshitomo-matsubara/srsd-feynman_hard.sr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09
SR-HOME-MEA_001-UN-BATHROOM_C004_L1_F1
High-quality synthetic render data-pack featuring a detailed residential bathroom scene, captured from Camera_04 with 100 rendered outputs and 1,200 files across 12 production-ready render passes.
This dataset is designed for computer vision, 3D perception, material analysis, segmentation, depth estimation, and synthetic data workflows. It includes beauty renders plus rich auxiliary passes such as albedo, depth, normals, UVW, material index… See the full description on the dataset page: https://huggingface.co/datasets/DaiPatrick/sr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09.srsd-feynman_medium
Dataset Card for SRSD-Feynman (Medium set)
Dataset Summary
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
We carefully reviewed the properties of each formula and its variables in the Feynman Symbolic Regression Database to design reasonably realistic sampling range of values so that our SRSD datasets can be used for evaluating the potential of SRSD such as whether or not an SR method con (re)discover… See the full description on the dataset page: https://huggingface.co/datasets/yoshitomo-matsubara/srsd-feynman_medium.LSDIR_SR_imagessrsd-feynman_easy
Dataset Card for SRSD-Feynman (Easy set)
Dataset Summary
Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery.
We carefully reviewed the properties of each formula and its variables in the Feynman Symbolic Regression Database to design reasonably realistic sampling range of values so that our SRSD datasets can be used for evaluating the potential of SRSD such as whether or not an SR method con (re)discover… See the full description on the dataset page: https://huggingface.co/datasets/yoshitomo-matsubara/srsd-feynman_easy.SROIE_2019_with_labelsThis is a fork of Urban Knupleš's "SROIE datasetv2", with ground truth labels for invoice numbers. train/labels.json and test/test_labels.json contains the ground truth invoice numbers for the associated image.
Below is a copy + paste from his repo. Here is my dataset documentation and notes. See there as well for heuristics used to label the dataset. The latest version tag is data-v2.1.
Scanned receipts OCR and information extraction (SROIE) + LayoutLM (base)
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ryanznie/SROIE_2019_with_labels.
