datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MapPool
MapPool - Bubbling up an extremely large corpus of maps for AI
MapPool is a dataset of 75 million potential maps and textual captions. It has been derived from CommonPool, a dataset consisting of 12 billion text-image pairs from the Internet. The images have been encoded by a vision transformer and classified into maps and non-maps by a support vector machine. This approach outperforms previous models and yields a validation accuracy of 98.5%. The MapPool dataset may help to train… See the full description on the dataset page: https://huggingface.co/datasets/sraimund/MapPool.srankmonsternobehemothdakedonekotomachigawareteelfmusumenopettoshitekurashitemasu
Bangumi Image Base of S-rank Monster No "behemoth" Dakedo, Neko To Machigawarete Elf Musume No Pet Toshite Kurashitemasu
This is the image base of bangumi S-Rank Monster no "Behemoth" dakedo, Neko to Machigawarete Elf Musume no Pet toshite Kurashitemasu, we detected 56 characters, 4649 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/srankmonsternobehemothdakedonekotomachigawareteelfmusumenopettoshitekurashitemasu.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.CountQA
Dataset Summary
CountQA is the new benchmark designed to stress-test the Achilles' heel of even the most advanced Multimodal Large Language Models (MLLMs): object counting. While modern AI demonstrates stunning visual fluency, it often fails at this fundamental cognitive skill, a critical blind spot limiting its real-world reliability.
This dataset directly confronts that weakness with over 1,500 challenging question-answer pairs built on real-world images, hand-captured to feature… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/CountQA.imagenet_resized_64x64This is an upload of imagenet_resized/64x64 from tensorflow datasets, (but shuffled before uploading).
The homepage of imagenet_resized is: https://patrykchrabaszcz.github.io/Imagenet32/
imagenet_resized is a derivative of imagenet (and also available to download from there): https://image-net.org/index.php
Warning: The integer labels used are defined by the authors and do not match those from the other ImageNet datasets provided by Tensorflow datasets. See the original label list, and the… See the full description on the dataset page: https://huggingface.co/datasets/sradc/imagenet_resized_64x64.eagle-iter267-sra-phase1-enclocale-benchmark-sra500
Locale Embedding Benchmark : SRA 500
What
This benchmark contains embeddings of raw genomic read sequences produced by the LOCALE DNA transformer model to test the use of vector search over large sequence repositories like the NIH Sequence Read Archive. The benchmark contains the embeddings of 163,578,486 sequence embeddings coming from 500 SRA Accessions. All vectors are Float32, D=768, roughly 500GB of total data.
Given a read, we want to find accessions… See the full description on the dataset page: https://huggingface.co/datasets/rsynk/locale-benchmark-sra500.HueManity
HueManity: A Benchmark for Testing Human-Like Visual Perception in MLLMs
Paper | Code
Dataset Description
HueManity is a benchmark dataset featuring 83,850 images designed to test the fine-grained visual perception of Multimodal Large Language Models (MLLMs). Each image presents a two-character alphanumeric string embedded within Ishihara-style dot patterns, challenging models to perform precise pattern recognition in visually cluttered environments.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/HueManity.chunked-wikipedia20220301en-bookcorpusopen
Dataset Card for "chunked-wikipedia20220301en-bookcorpusopen"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The logic only splits on spaces, so the chunks are likely to be slightly larger than 820 chars.
The dataset has been normalized into lower case… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-wikipedia20220301en-bookcorpusopen.chunked-shuffled-wikipedia20220301en-bookcorpusopen
Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The order of the items in this dataset has been shuffled,
meaning you don't have to use dataset.shuffle,
which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.SRA-Bench
SRA-Bench: Skill Retrieval Augmentation for Agentic AI
SRA-Bench is a benchmark for studying Skill Retrieval Augmentation (SRA): how LLM agents retrieve, select, and use external skills to solve capability-intensive tasks.
RAG retrieves knowledge. SRA retrieves capabilities.
Modern LLM agents increasingly rely on external resources such as tools, APIs, memory, workflows, executable code, and reusable skills.SRA-Bench provides a large-scale benchmark for evaluating whether… See the full description on the dataset page: https://huggingface.co/datasets/WeihangSu/SRA-Bench.sr_atad5_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/sr_atad5_tox21.sr-artifact-detection-demo-datasetfinewebedu-climate-v2runaround_ep1
runaround_ep1
LeRobot v2.1 format dataset for robot manipulation.
Dataset Structure
Episodes: 1 episodes of robot manipulation
Total Frames: 125 frames
Cameras: 3 camera views per episode
observation.images.base_camera_sensor_image_raw
observation.images.arm1_camera_sensor_image_raw
observation.images.arm2_camera_sensor_image_raw
Robot: Bimanual manipulator with 34 joints
Format: LeRobot v2.1
Usage with LeRobot
from… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/runaround_ep1.finewebedu-climatesubset-0-official-lerobot
subset-0-official
Robot dataset converted from ROS2 bag file.
Frames: 310
FPS: 10.0
Joints: 6 DOF
Duration: 31.00s
SRA
SRA-Bench
MTEB v2 text retrieval dataset layout for SRA-Bench. The brief subset uses the compact corpus, and the full subset uses the expanded corpus-full source text.
sravaani-indic-diarbench-oracle-v1
SraVaani Indic DiarBench Oracle ASR
This is a portable evaluation-only oracle-turn view derived from
sarvamai/indic-diarbench at the
immutable revision 92877bad8aab6e598167d91c6ee02aa8ca6ede09. It contains Hindi and Telugu only.
Do not use these turns for fine-tuning if Indic DiarBench will remain an
external benchmark. Training on this export contaminates the test set.
Configurations
Config
Test rows
Audio hours
Purpose
primary
2,417
3.8233
Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhirl/sravaani-indic-diarbench-oracle-v1.locale-benchmark-sra50
Locale Embedding Benchmark : SRA 50
What
This benchmark contains embeddings of raw genomic read sequences produced by the LOCALE DNA transformer model to test the use of vector search over large sequence repositories like the NIH Sequence Read Archive. The benchmark contains the embeddings of 9,688,220 sequence embeddings coming from 50 SRA Accessions. All vectors are Float32, D=768, roughly 30GB of total data.
Given a read, we want to find accessions containing… See the full description on the dataset page: https://huggingface.co/datasets/rsynk/locale-benchmark-sra50.test_outout_shrugmaster
Shrugmaster_Bimanual_Manipulation
Robot manipulation dataset with 4 episodes.
Dataset Structure
Episodes: 4 episodes of robot manipulation
Cameras: 3 camera views (base, arm1, arm2)
Joints: 34 joint states
Format: LeRobot v2.1
Usage
from lerobot.common.datasets import LeRobotDataset
dataset = LeRobotDataset("Sraghvi/shrugmaster-robot-dataset")
Files
data/chunk-000/: Parquet files with episode data
videos/chunk-000/: MP4 video files for each… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/test_outout_shrugmaster.BioFormBench
BioFormBench: Two Open Benchmarks for Biologics Formulation Research
This release contains two datasets, kept separate because they answer different
questions and neither should be read as a substitute for the other.
Data integrity note (please read before using either dataset)
An earlier internal draft of BioFormBench-Real contained 67 rows. Auditing
source_title/source_id provenance found that 18 of those rows had no
citation and used placeholder identifiers… See the full description on the dataset page: https://huggingface.co/datasets/Sravankumarbonthada/BioFormBench.dual-arm-robot-dataset
Dual-Arm Robot Dataset
A LeRobot-compatible dataset containing joint position data from a dual-arm robotic system with 34 degrees of freedom.
Dataset Description
This dataset contains robot manipulation data recorded from a dual-arm robotic system. The data includes joint positions for both arms and their associated grippers, captured at 100Hz for precise motion analysis and learning.
Dataset Summary
Total Episodes: 1
Total Frames: 3,105
Duration: 31.05… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/dual-arm-robot-dataset.train-policy
buster_episode_0001
LeRobot v2.1 format dataset for robot manipulation.
Dataset Structure
Episodes: 14 episodes of robot manipulation
Total Frames: 21923 frames
Cameras: 3 camera views per episode
observation.images.base_camera_sensor_image_raw
observation.images.arm1_camera_sensor_image_raw
observation.images.arm2_camera_sensor_image_raw
Robot: Bimanual manipulator with 34 joints
Format: LeRobot v2.1
Usage with LeRobot
from… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/train-policy.ros2-robot-dataset-v21so101_sra_hri_tasks_2camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 48,
"total_frames": 16577,
"total_tasks": 1,
"total_videos": 96,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:48"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cchenalds17/so101_sra_hri_tasks_2cam.buster-smooth-sync
BUSTER_Episode_2
LeRobot v2.1 format dataset for robot manipulation.
Dataset Structure
Episodes: 2 episodes of robot manipulation
Total Frames: 2303 frames
Cameras: 3 camera views per episode
observation.images.base_camera_sensor_image_raw
observation.images.arm1_camera_sensor_image_raw
observation.images.arm2_camera_sensor_image_raw
Robot: Bimanual manipulator with 34 joints
Format: LeRobot v2.1
Usage with LeRobot
from… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/buster-smooth-sync.rocessed_so101_combinedshrugmaster-test
Shrugmaster Test
LeRobot v2.1 format dataset for robot manipulation.
Dataset Structure
Episodes: 4 episodes of robot manipulation
Total Frames: 1,440 frames
Cameras: 3 camera views per episode
observation.images.base_camera_sensor_image_raw
observation.images.arm1_camera_sensor_image_raw
observation.images.arm2_camera_sensor_image_raw
Robot: Bimanual manipulator with 34 joints
Format: LeRobot v2.1
Usage with LeRobot
from… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/shrugmaster-test.uploadmerge
Bone to Pick - Combined Dataset (10_08 + 10_09)
This dataset is a merged combination of two Bone to Pick datasets from Hugging Face:
oordonez/Bone_to_Pick_10_08
oordonez/Bone_to_Pick_10_09
Dataset Description
This dataset was created using LeRobot and contains bi-manual robot manipulation data for a bone picking task. The dataset combines 20 episodes (10 from each original dataset) with updated episode numbering to avoid conflicts.
Key Features:… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/uploadmerge.
