datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
robopoint-data
RoboPoint Dataset Card
Dataset details
This dataset contains 1432K image-QA instances used to fine-tune RoboPoint, a VLM for spatial affordance prediction. It consists of the following parts:
347K object reference instances from a synthetic data pipeline;
320K free space reference instances from a synthetic data pipeline;
100K object detection instaces from LVIS;
150K GPT-generated instruction-following instances from liuhaotian/LLaVA-Instruct-150K;
515K general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/wentao-yuan/robopoint-data.medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.tvqa-framesdanbooru2023-webp-4Mpixel-224The data set is just resized to 224*224
https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel
Pseudo code for processing
def resize_image(file_path):
with Image.open(file_path) as img:
resized_img = img.resize((224, 224))
resized_img.save(file_path)
cc12m-webdataset
CC12M WebDataset
这是CC12M数据集的WebDataset格式版本。
数据集信息
文件数量: 1098
总大小: 888796.33 MB
上传时间: 2025-03-18 14:45:49
使用方法
import webdataset as wds
dataset = wds.WebDataset("https://huggingface.co/yangyang857658468/cc12m-webdataset/resolve/main/cc12m_*.tar")
Glint360k
Dataset Card for Glint360K
Citiation by InsightFace Repository
We clean, merge, and release the largest and cleanest face recognition dataset Glint360K, which contains 17091657 images of 360232 individuals. By employing the Patial FC training strategy, baseline models trained on Glint360K can easily achieve state-of-the-art performance. Detailed evaluation results on the large-scale test set (e.g. IFRT, IJB-C and Megaface) are as follows:
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yayoimizuha/Glint360k.ppv2_imagetarUniSAR-7M
UniSAR-7M
A large-scale, multi-source synthetic aperture radar image corpus for self-supervised representation learning.
UniSAR-7M contains 7,047,666 single-channel SAR image samples assembled from public SAR datasets and openly available imagery from commercial satellite constellations. It provides the pretraining corpus for DINOSAR, a self-supervised learning framework that uses Content-Aware Multi-Crop (CAMC) to construct informative views of SAR imagery.
Associated… See the full description on the dataset page: https://huggingface.co/datasets/YTang/UniSAR-7M.GlobalGeoTree
GlobalGeoTree Dataset
GlobalGeoTree is a comprehensive global dataset for tree species classification, comprising 6.3 million geolocated tree occurrences spanning 275 families, 2,734 genera, and 21,001 species across hierarchical taxonomic levels. Each sample is paired with Sentinel-2 image time series and 27 auxiliary environmental variables.
Dataset Structure
This repository contains three main components:
1. GlobalGeoTree-6M
Training dataset with around 6M… See the full description on the dataset page: https://huggingface.co/datasets/yann111/GlobalGeoTree.Echo-4o-Image
Echo-4o-Image Dataset
Paper | Project Page | Code
Introduction
Echo-4o-Image is a 180K-scale synthetic dataset generated by GPT-4o, designed to advance open-source models in image generation. While real-world image datasets are valuable, synthetic images offer crucial advantages, especially in addressing blind spots in real-world coverage:
Complementing Rare Scenarios: Synthetic data can generate examples for scenarios less represented in real-world datasets, such as… See the full description on the dataset page: https://huggingface.co/datasets/Yejy53/Echo-4o-Image.dense-sc-wdsyuxuan_good_dataset_dtsc-wdsDynamicvlmyuxuan_good_dataset_sttcc3m-subset-100kyfcc15mVideoSSR-30kVoT-video-latent-archivesd-yffSEE-600Kgrounding-YT-dataset
Grounding YouTube Dataset
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
arxiv
This dataset is packed in WebDataset format.
The dataset is present in three styles:
Untrimmed videos + annotations within the entire video
Action clips extracted from the videos + annotations in each clip
Action frames extracted from the videos + annotation of the frame
Example usage for clips:… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/grounding-YT-dataset.yuxuan_dataset_dtTraceGenLibero
TraceGen – LIBERO (Derived Subset)
Overview
This folder contains a derived subset generated from the LIBERO dataset using the
TraceForge pipeline as part of the TraceGen project.
TraceGen Project Website: https://tracegen.github.io/
Evaluation Protocol
This dataset defines the official evaluation protocol for the TraceGen benchmark.
Models are evaluated on five environments with the following metrics:
Mean Squared Error (MSE)
Mean Absolute Error (MAE)… See the full description on the dataset page: https://huggingface.co/datasets/yoonkyojung/TraceGenLibero.YTVIS2021_Splityuxuan_dataset_sttIllumiCraft
IllumiCraft Dataset
This repository contains the dataset released with:
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark, Ming-Hsuan Yang
🔗 Links
📄 Paper: https://arxiv.org/abs/2506.03150
🌐 Project Page: https://yuanze-lin.me/IllumiCraft_page/
💻 GitHub: https://github.com/yuanze-lin/IllumiCraft
🎥 YouTube: https://youtu.be/qAV58sADEzo
🤗 Checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/YuanzeLin/IllumiCraft.ElysiumTrack-1M
Dataset Card
ElysiumTrack-1M dataset is a million-scale object perception video dataset. It supports the following tasks:
Single Object Tracking (SOT): Predicting the location of a specific object in consecutive frames by referencing its initial position in the first frame.
Referring Single Object Tracking (RSOT): Identifying and locating a specific object within an entire video based on the given language expression. This task provides a more flexible tracking format and… See the full description on the dataset page: https://huggingface.co/datasets/sty-yyj/ElysiumTrack-1M.lora_model
