CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mvp-lab /Sekaitext1M<n<10M0 likes24k downloads10mo agoHugging Face02TIGER-Lab /arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot. image10M<n<100M5 likes5.8k downloads1y agoHugging Face03nirschl-lab /hpa10m HPA10M Dataset A large-scale immunohistochemistry (IHC) image dataset derived from the Human Protein Atlas (HPA, https://www.proteinatlas.org/), containing approximately 10.5 million pathology and tissue images with detailed annotations. Dataset Overview Statistic Value Total Images 10,495,672 Training Set 10,493,672 images (10,497 tar files) Validation Set 2,000 images (1 tar file) Image Types Pathology (7,970,595) / Tissue (2,525,077) Format JPEG… See the full description on the dataset page: https://huggingface.co/datasets/nirschl-lab/hpa10m.image10M<n<100M7 likes3k downloads4mo agoHugging Face04TIGER-Lab /VISTA-400K VISTA-400K This repo contains all subsets for VISTA-400K. VISTA is a video spatiotemporal augmentation method that generates long-duration and high-resolution video instruction-following data to enhance the video understanding capabilities of video LMMs. This repo is under construction. Please stay tuned. 🌐 Homepage | 📖 arXiv | 💻 GitHub | 🤗 VISTA-400K | 🤗 Models | 🤗 HRVideoBench Video Instruction Data Synthesis Pipeline VISTA leverages insights from… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VISTA-400K.textquestion-answering100K<n<1M6 likes1.6k downloads2y agoHugging Face05ASLP-lab /Easy-Turn-Trainset Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems Guojian Li1, Chengyou Wang1, Hongfei Xue1, Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2, Yuke Lin2, Wenjie Li2, Longshuai Xiao2, Zhonghua Fu1,╀, Lei Xie1,╀ 1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University 2 Huawei Technologies, China 🎤 Demo Page 🤖 Easy Turn Model 📑 Paper 🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.imageautomatic-speech-recognition1K<n<10K12 likes1.2k downloads11mo agoHugging Face06iLearn-Lab /Optimus-2-MGOAThis repository contains the data presented in Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy. Code: https://github.com/lizaijing/Optimus-2 text10M<n<100M2 likes988 downloads1y agoHugging Face07labhamlet /NatHEARPlease see LICENSE.txt for each dataset licenses audio10K<n<100K0 likes495 downloads8mo agoHugging Face08ASLP-lab /WSYue-ASR-eval WSYue-ASR-eval: Cantonese ASR Benchmark To address the unique linguistic characteristics of Cantonese in speech recognition, we propose WSYue-ASR-eval, a benchmark specifically designed for evaluating Cantonese ASR systems. It is tailored to assess model performance across diverse lengths, domains, and linguistic phenomena of Cantonese speech. The test set annotations are provided by Beijing AISHELL Technology Co., Ltd. Key features: Annotated through multiple rounds of manual… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/WSYue-ASR-eval.text1K<n<10K4 likes434 downloads1y agoHugging Face09orion-ai-lab /Thalia Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring Paper | GitHub | Interactive Demo (Colab) Thalia is a global, multi-modal dataset for volcanic activity monitoring through Satellite-based Interferometric Synthetic Aperture Radar (InSAR) imagery. Building upon the Hephaestus dataset, Thalia provides higher-resolution, multi-source, and multi-temporal data in a machine-learning-ready format. Dataset Overview Thalia consists of 38 spatiotemporal… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Thalia.textimage-classification10K<n<100K3 likes331 downloads5mo agoHugging Face10perona-lab /cfc26 CFC26 Dataset Card Dataset Summary CFC26 is a large-scale benchmark dataset for fish detection, tracking, and counting in underwater ARIS sonar video. It is designed to evaluate generalization under distribution shift and deployment-relevant performance, particularly in ecologically diverse and acoustically challenging environments. The dataset spans multiple river systems with substantial variation in fish size, density, sonar range, and background structure, and… See the full description on the dataset page: https://huggingface.co/datasets/perona-lab/cfc26.image1M<n<10M1 likes279 downloads5mo agoHugging Face11nyu-dice-lab /imagenetpp-laion-t2iDataset Card for ImageNet++'s LAION Text-to-Image Split image100K<n<1M0 likes228 downloads2y agoHugging Face12lmms-lab /LLaVA-OneVision-Mid-Data Dataset Card for LLaVA-OneVision Due to unknow reasons, we are unable to process dataset with large amount into required HF format. So we directly upload the json files and image folders (compressed into tar.gz files). You can use the following link to directly download and decompress them. https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data/tree/main/evol_instruct We provide the whole details of LLaVA-OneVision Dataset. In this dataset, we include the data splits… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data.imagetext-generation100K<n<1M21 likes142 downloads2y agoHugging Face13matsuo-lab /JP-LLM-Corpus-PII-Filtered-10B CommonCrawl Japanese (Filtered PPI) Dataset 本データセットは、CommonCrawlより抽出した約100億(10B)トークン規模の日本語テキストデータから、特に配慮が必要な「要配慮個人情報」をフィルタリング処理したものです。 データセットの概要 元データソース: CommonCrawl(https://commoncrawl.org/) トークン数: 約10Bトークン 言語: 日本語 処理内容: 要配慮個人情報をルールベースおよび機械学習分類器を用いてフィルタリング フィルタリングには以下のコードを使用しております。https://github.com/matsuolab/jp-llm-corpus-pii-filter/ 注意事項 本データセットは、非常に大規模なテキストから自動的に要配慮個人情報を除去したものであり、完全な排除を保証するものではありません。そのため、二次的な活用に際しては、目的に応じた適切な管理・配慮が必要です。… See the full description on the dataset page: https://huggingface.co/datasets/matsuo-lab/JP-LLM-Corpus-PII-Filtered-10B.text10K<n<100K1 likes135 downloads1y agoHugging Face14clip-benchmark /wds_vtab-smallnorb_label_elevationimage10K<n<100K0 likes111 downloads4y agoHugging Face15SMIIP-lab /AISHELL6-Whispergated 🗣️ AISHELL6-Whisper AISHELL6-Whisper is a large-scale open-source Chinese Mandarin audio-visual whisper speech dataset,containing 30 hours each of whisper and parallel normal speech, with synchronized frontal RGB facial videos. 📘 Dataset Summary Property Description Language Chinese (Mandarin, ZH) License CC BY-NC-SA 4.0 Duration ~60 hours total (30 h whisper + 30 h normal) Speakers 167 total (121 with RGB-D, 46 audio-only) Environment Controlled… See the full description on the dataset page: https://huggingface.co/datasets/SMIIP-lab/AISHELL6-Whisper.audio10K<n<100K8 likes111 downloads8mo agoHugging Face16nyu-dice-lab /imagenetpp-laion-i2iimage100K<n<1M0 likes99 downloads2y agoHugging Face17clip-benchmark /wds_vtab-dsprites_label_x_positionimage100K<n<1M0 likes66 downloads4y agoHugging Face18labhamlet /Nat-HEAR-Ambisonicsaudio10K<n<100K0 likes65 downloads9mo agoHugging Face19clip-benchmark /wds_vtab-dsprites_label_orientationimage100K<n<1M0 likes56 downloads4y agoHugging Face20clip-benchmark /wds_vtab-smallnorb_label_azimuthimage10K<n<100K0 likes55 downloads4y agoHugging Face21LLLM-Lab /huanhuan_sadtalkertextn<1K0 likes54 downloads3y agoHugging Face22PEARLS-Lab /infini-thor-niehimage100K<n<1M0 likes43 downloads7mo agoHugging Face23clip-benchmark /wds_vtab-dsprites_label_y_positionimage100K<n<1M0 likes42 downloads4y agoHugging Face24Metavolve-Labs /alexandria-aeternum-1k Alexandria Aeternum — Genesis Your Entry Point to Cognitive Nutrition 1,000 curated paintings · Masters only · 4,000+ tokens each · Free sample Not scraped. Not auto-captioned. Translated from human knowledge. Monet, Van Gogh, Rembrandt, Degas, Hokusai, Cezanne, and 100+ master artists. Full 10K Dataset · MCP Access (2M+ Artworks) · Research Paper · Explore Full Archive · Scale With Us MCP Access — AI Agent Marketplace The complete high-resolution… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/alexandria-aeternum-1k.imagetext-to-image1K<n<10K1 likes42 downloads7mo agoHugging Face25labhamlet /BinauralRIRsTestOnlytext10K<n<100K0 likes38 downloads9mo agoHugging Face26Bitbol-Lab /ProteomeLM-datasettext10K<n<100K1 likes34 downloads1y agoHugging Face27PEARLS-Lab /infini-thor-traintext1K<n<10K0 likes34 downloads7mo agoHugging Face28hci-lab-dcug /ugspeech-akan-clean-100hrsaudio10K<n<100K0 likes32 downloads1y agoHugging Face29labhamlet /GRAM-Naturalistic-Scenes-Binauraltext10K<n<100K0 likes30 downloads8mo agoHugging Face30zhangyupeng90 /labelgs_datasetsimage1K<n<10K1 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.