datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hub_weekly_snapshots
Sample code
To query the dataset to see which snapshots are observable, use e.g.:
import json
from datasets import load_dataset
from huggingface_hub import HfApi
REPO_ID = "hfmlsoc/hub_weekly_snapshots"
hf_api = HfApi()
all_files = hf_api.list_repo_files(repo_id=REPO_ID, repo_type="dataset")
repo_type_to_snapshots = {}
for repo_fpath in all_files:
if ".parquet" in repo_fpath:
repo_type = repo_fpath.split("/")[0]
repo_type_to_snapshots[repo_type] =… See the full description on the dataset page: https://huggingface.co/datasets/hfmlsoc/hub_weekly_snapshots.esa-hubble
Dataset Card for ESA Hubble Deep Space Images & Captions
Dataset Summary
The ESA Hubble Deep Space Images & Captions dataset is composed primarily of Hubble deep space scans as high-resolution images,
along with textual descriptions written by ESA/Hubble. Metadata is also included, which enables more detailed filtering and understanding of massive space scans.
The purpose of this dataset is to enable text-to-image generation methods for generating high-quality deep space… See the full description on the dataset page: https://huggingface.co/datasets/Supermaxman/esa-hubble.hubble_trailsKMMMU
KMMMU (Korean MMMU)
technical report https://arxiv.org/abs/2604.13058
link to evaluation tutorial! https://github.com/HAE-RAE/KMMMU
KMMMU is a Korean version of MMMU: a multimodal benchmark designed to evaluate college-/exam-level reasoning that requires combining images + Korean text.
This dataset contains 3,466 questions collected from Korean exam sources including:
Civil service recruitment exams
National Technical Qualifications
National Competency Standard (NCS) exams… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMMU.hubble-8b-unlearning-resultsHUBT_from_a_drones_perspective
HUTB From a Drone's Perspective
Dataset Summary
HUTB From a Drone's Perspective is a synthetic, multi-map UAV dataset generated
with OpenHUTB/CARLA. It provides synchronized visible RGB, metric depth,
surface normals, target semantic segmentation, LiDAR, and vehicle/pedestrian
detection annotations from elevated oblique viewpoints.
The current release contains:
4,081 synchronized frames at 1920 x 1080 pixels.
8 simulated maps.
6 weather and illumination… See the full description on the dataset page: https://huggingface.co/datasets/yutiangu/HUBT_from_a_drones_perspective.Batch_2_09img-hubgz_hubble
GZ Campaign Datasets
Dataset Summary
Galaxy Zoo volunteers label telescope images of galaxies according to their visible features: spiral arms, galaxy-galaxy collisions, and so on.
These datasets share the galaxy images and volunteer labels in a machine-learning-friendly format. We use these datasets to train our foundation models. We hope they'll help you too.
Curated by: Mike Walmsley
License: cc-by-nc-sa-4.0. We specifically require all models trained on these… See the full description on the dataset page: https://huggingface.co/datasets/mwalmsley/gz_hubble.nonastreda
Nonastreda: Multimodal Dataset for Tool Wear State Monitoring
Nonastreda is a multimodal dataset for efficient tool wear state monitoring in milling. It contains 512 sample-level records combining visual, time-frequency, and force-signal-derived representations of tool wear.
This Hugging Face repository is a mirror and machine-learning-friendly access point for the dataset. The canonical scholarly description is the associated Data in Brief article, and the canonical archived… See the full description on the dataset page: https://huggingface.co/datasets/hubtru/nonastreda.chahuadev-hub
Chahuadev Hub
@chahuadev/chahuadev-hub-app
Chahuadev Hub — The official desktop hub by Chahuadev. Browse repositories, download apps, explore npm packages, and chat with the community — all in one dark-mode Electron app.
⚙️ Installation
Requires Node.js 18+. Use --foreground-scripts to see install progress.
🪟 Windows — Global Install
npm install -g @chahuadev/chahuadev-hub-app --foreground-scripts --force
The installer… See the full description on the dataset page: https://huggingface.co/datasets/chahuadev/chahuadev-hub.hubble-anomaly
Hubble Anomaly
This dataset contains 20,000 HST ACS/WFC F814W cutouts assembled from a
223,195-source parent catalogue. The parent catalogue was built from a
10-million-row Hubble Source Catalog v3 extended-source pull, associated
with Hubble Advanced Products, and deduplicated to a minimum separation of
10 arcsec.
Anomaly labels and types come from O'Ryan, D. & Gómez, P. (2025),
Identifying astrophysical anomalies in 99.6 million source cutouts from
the Hubble Legacy Archive… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/hubble-anomaly.HAERAE-VISION
HAERAE-VISION
A Korean visual QA benchmark featuring real-world, under-specified questions.
Dataset Description
This dataset includes two question types:
original: Under-specified, authentic user queries
explicit: Clarified queries with full context
Both share the same images and reference answers, allowing controlled evaluation of query under-specification.
Evaluation Code
See our GitHub repository for evaluation scripts.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAERAE-VISION.hle_futurehouse_goldLactylationChartVerse-RL-10K-auditedmjswan-hub-data
mjswan Hub Demo Registry
Database for mjswan Hub.
Structure
demos.json — List of all registered demos
thumbnails/ — Thumbnail images for each demo
hle_futurehouse_bronzeDeepseek-zebra-puzzleOvarianUltrasoundFeatureExtraction
OvarianUltrasoundFeatureExtraction
tags: Regression, Feature Learning, Gynecological Imaging
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'OvarianUltrasoundFeatureExtraction' dataset comprises high-resolution ultrasound images of the ovaries from various gynecological imaging sources. Each image is annotated with key features and corresponding labels, such as follicular cysts, corpus luteum, and endometriomas. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/OvarianUltrasoundFeatureExtraction.UCF101-I2Vname-of-repo-on-the-hub7ABO-I2V
ABO-I2V
Image-to-video product retrieval derived from Amazon Berkeley Objects (ABO).
Built for MTEB as ABOI2VRetrieval.
Each product contributes one video (its 360-degree turntable "spin") and one query image
(an independently shot catalog photograph of the same product, in a room or scene). The two assets
were captured at different times, in different places, under different lighting, so a query is
never a frame of its own positive video and frame leakage is structurally… See the full description on the dataset page: https://huggingface.co/datasets/hubxrt/ABO-I2V.svlm-preprocessed-datasets-v2gz_hubble_wdsdemo_ds_hf_hub_segformer_fine_tuned_ADE20k_format_rgb_crudo_oi_IS_v1
Dataset Card for primer_demo_ejemplo_ds_semantic
Dataset categories
Id
Name
Description
1
objeto_interes
[128, 0, 0]
2
agua
[0, 128, 0]
hf-hub-walkthrough-assetsname-of-repo-on-the-hub1links-hub-demobutterflies_and_moths_vqa
Butterflies and Moths VQA
Dataset Summary
butterflies_and_moths_vqa is a visual question answering (VQA) dataset focused on butterflies and moths. It features tasks such as fine-grained species classification and ecological reasoning. The dataset is designed to benchmark Vision-Language Models (VLMs) for both image-based and text-only training approaches.
Key Features
Fine-Grained Classification (Type1): Questions requiring detailed species identification.… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/butterflies_and_moths_vqa.
