datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sekaiarxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot.
hpa10m
HPA10M Dataset
A large-scale immunohistochemistry (IHC) image dataset derived from the Human Protein Atlas (HPA, https://www.proteinatlas.org/), containing approximately 10.5 million pathology and tissue images with detailed annotations.
Dataset Overview
Statistic
Value
Total Images
10,495,672
Training Set
10,493,672 images (10,497 tar files)
Validation Set
2,000 images (1 tar file)
Image Types
Pathology (7,970,595) / Tissue (2,525,077)
Format
JPEG… See the full description on the dataset page: https://huggingface.co/datasets/nirschl-lab/hpa10m.VISTA-400K
VISTA-400K
This repo contains all subsets for VISTA-400K. VISTA is a video spatiotemporal augmentation method that generates long-duration and high-resolution video instruction-following data to enhance the video understanding capabilities of video LMMs.
This repo is under construction. Please stay tuned.
🌐 Homepage | 📖 arXiv | 💻 GitHub | 🤗 VISTA-400K | 🤗 Models | 🤗 HRVideoBench
Video Instruction Data Synthesis Pipeline
VISTA leverages insights from… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VISTA-400K.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.Optimus-2-MGOAThis repository contains the data presented in Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy.
Code: https://github.com/lizaijing/Optimus-2
NatHEARPlease see LICENSE.txt for each dataset licenses
WSYue-ASR-eval
WSYue-ASR-eval: Cantonese ASR Benchmark
To address the unique linguistic characteristics of Cantonese in speech recognition, we propose WSYue-ASR-eval, a benchmark specifically designed for evaluating Cantonese ASR systems. It is tailored to assess model performance across diverse lengths, domains, and linguistic phenomena of Cantonese speech.
The test set annotations are provided by Beijing AISHELL Technology Co., Ltd.
Key features:
Annotated through multiple rounds of manual… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/WSYue-ASR-eval.Thalia
Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring
Paper | GitHub | Interactive Demo (Colab)
Thalia is a global, multi-modal dataset for volcanic activity monitoring through Satellite-based Interferometric Synthetic Aperture Radar (InSAR) imagery. Building upon the Hephaestus dataset, Thalia provides higher-resolution, multi-source, and multi-temporal data in a machine-learning-ready format.
Dataset Overview
Thalia consists of 38 spatiotemporal… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Thalia.cfc26
CFC26 Dataset Card
Dataset Summary
CFC26 is a large-scale benchmark dataset for fish detection, tracking, and counting in underwater ARIS sonar video.
It is designed to evaluate generalization under distribution shift and deployment-relevant performance, particularly in ecologically diverse and acoustically challenging environments.
The dataset spans multiple river systems with substantial variation in fish size, density, sonar range, and background structure, and… See the full description on the dataset page: https://huggingface.co/datasets/perona-lab/cfc26.imagenetpp-laion-t2iDataset Card for ImageNet++'s LAION Text-to-Image Split
LLaVA-OneVision-Mid-Data
Dataset Card for LLaVA-OneVision
Due to unknow reasons, we are unable to process dataset with large amount into required HF format. So we directly upload the json files and image folders (compressed into tar.gz files).
You can use the following link to directly download and decompress them.
https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data/tree/main/evol_instruct
We provide the whole details of LLaVA-OneVision Dataset. In this dataset, we include the data splits… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data.JP-LLM-Corpus-PII-Filtered-10B
CommonCrawl Japanese (Filtered PPI) Dataset
本データセットは、CommonCrawlより抽出した約100億(10B)トークン規模の日本語テキストデータから、特に配慮が必要な「要配慮個人情報」をフィルタリング処理したものです。
データセットの概要
元データソース: CommonCrawl(https://commoncrawl.org/)
トークン数: 約10Bトークン
言語: 日本語
処理内容: 要配慮個人情報をルールベースおよび機械学習分類器を用いてフィルタリング
フィルタリングには以下のコードを使用しております。https://github.com/matsuolab/jp-llm-corpus-pii-filter/
注意事項
本データセットは、非常に大規模なテキストから自動的に要配慮個人情報を除去したものであり、完全な排除を保証するものではありません。そのため、二次的な活用に際しては、目的に応じた適切な管理・配慮が必要です。… See the full description on the dataset page: https://huggingface.co/datasets/matsuo-lab/JP-LLM-Corpus-PII-Filtered-10B.wds_vtab-smallnorb_label_elevationAISHELL6-Whisper
🗣️ AISHELL6-Whisper
AISHELL6-Whisper is a large-scale open-source Chinese Mandarin audio-visual whisper speech dataset,containing 30 hours each of whisper and parallel normal speech, with synchronized frontal RGB facial videos.
📘 Dataset Summary
Property
Description
Language
Chinese (Mandarin, ZH)
License
CC BY-NC-SA 4.0
Duration
~60 hours total (30 h whisper + 30 h normal)
Speakers
167 total (121 with RGB-D, 46 audio-only)
Environment
Controlled… See the full description on the dataset page: https://huggingface.co/datasets/SMIIP-lab/AISHELL6-Whisper.imagenetpp-laion-i2iwds_vtab-dsprites_label_x_positionNat-HEAR-Ambisonicswds_vtab-dsprites_label_orientationwds_vtab-smallnorb_label_azimuthhuanhuan_sadtalkerinfini-thor-niehwds_vtab-dsprites_label_y_positionalexandria-aeternum-1k
Alexandria Aeternum — Genesis
Your Entry Point to Cognitive Nutrition
1,000 curated paintings · Masters only · 4,000+ tokens each · Free sample
Not scraped. Not auto-captioned. Translated from human knowledge.
Monet, Van Gogh, Rembrandt, Degas, Hokusai, Cezanne, and 100+ master artists.
Full 10K Dataset · MCP Access (2M+ Artworks) · Research Paper · Explore Full Archive · Scale With Us
MCP Access — AI Agent Marketplace
The complete high-resolution… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/alexandria-aeternum-1k.BinauralRIRsTestOnlyProteomeLM-datasetinfini-thor-trainugspeech-akan-clean-100hrsGRAM-Naturalistic-Scenes-Binaurallabelgs_datasets
