CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IPEC-COMMUNITY /FastUMI_100k_lerobot FastUMI-100K: Advancing Data-Driven Robotic Manipulation with a Large-Scale UMI-Style Dataset [paper] [dataset] ## Overview FastUMI-100K is a large-scale, high-quality UMI-style dataset designed for data-driven robotic manipulation learning. Featuring over **100K+ demonstration trajectories** across **54 diverse tasks** and hundreds of object types, the dataset provides multi-view wrist-mounted fisheye images and high-frequency end-effector states. To… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/FastUMI_100k_lerobot.8 likes278k downloads5mo agoHugging Face02builddotai /Egocentric-100Kgated Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here. Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation. Dataset Statistics Attribute Value Total Hours 100,405 Total Frames 10.8 billion Video Clips 2,010,759 Median Clip Length 180.0 seconds Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.text1M<n<10M142 likes140k downloads7mo agoHugging Face03mlabonne /FineTome-100k FineTome-100k The FineTome dataset is a subset of arcee-ai/The-Tome (without arcee-ai/qwen2-72b-magpie-en), re-filtered using HuggingFaceFW/fineweb-edu-classifier. It was made for my article "Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth". text100K<n<1M279 likes26k downloads2y agoHugging Face04TencentARC /TimeLens-100K TimeLens-100K 📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data ✨ Dataset Description TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro. 📊 Dataset Statistics Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.textvideo-text-to-text10K<n<100K7 likes24k downloads9mo agoHugging Face05RidheshBhati /Complete_Data_Source_100K_HOURS Multi-Language Audio Collection (100K Hours) This repository is physically reorganized for Absolute 100% Data Visibility. 🏗️ Global Consolidator Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here. audio1M<n<10M4 likes16k downloads5mo agoHugging Face06obswork /arxiv-ai-ml-100k-papers license: other tags: - arxiv - ocr - machine-learning --- # obswork/arxiv-ai-ml-100k A 99,999-paper stratified subset of [`Rendra8631/arxiv-papers`](https://huggingface.co/datasets/Rendra8631/arxiv-papers) at revision `a2c6afb51332d2744b46308df6917697582f8cd4`, filtered to the primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. Only papers submitted in `2023`-`2025` are included. This dataset is a build artifact of the OCR… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-papers.1 likes14k downloads5mo agoHugging Face07lavita /ChatDoctor-HealthCareMagic-100k Dataset Card for "ChatDoctor-HealthCareMagic-100k" More Information needed text100K<n<1M117 likes13k downloads3y agoHugging Face08obswork /arxiv-ai-ml-100k-pages license: other tags: - arxiv - ocr - machine-learning --- # obswork/arxiv-ai-ml-100k-pages A **page-bounded** stratified subset of the raw pool dataset [`obswork/arxiv-ai-ml-100k`](https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k), filtered to primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. The raw pool is itself a 100k-paper stratified sample from… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-pages.0 likes5.6k downloads5mo agoHugging Face09meituan-longcat /Q-Eval-100K Q-Eval-100K Dataset (CVPR 2025 Oral) 📝 Introduction The Q-Eval-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). We utilize multiple popular text-to- image and text-to-video models to ensure diversity, which include FLUX, Lumina-T2X, PixArt, Stable Diffusion 3, Stable Diffusion XL, DALL·E 3, Wanx, Midjourney, Hunyuan-DiT… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/Q-Eval-100K.11 likes5.6k downloads1y agoHugging Face10ADSKAILab /Zero-To-CAD-100k Zero-to-CAD 100K A curated subset of 100,000 geometrically diverse CAD construction sequences selected from Zero-to-CAD 1M. Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data Mohammadmehdi Ataei, Farzaneh Askari, Kamal Rahimi Malekshan, Pradeep Kumar Jayaraman Autodesk Research Related Resources Resource Link 📄 Paper Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without… See the full description on the dataset page: https://huggingface.co/datasets/ADSKAILab/Zero-To-CAD-100k.tabulartext-to-3d100K<n<1M19 likes5.4k downloads5mo agoHugging Face11zhihefang /UltraHR-100K UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset       UltraHR-100K Ultra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle the first challenge, we introduce… See the full description on the dataset page: https://huggingface.co/datasets/zhihefang/UltraHR-100K.imagetext-to-imagen<1K37 likes5.2k downloads10mo agoHugging Face12zhaojiao /ai-images-100kimagen<1K0 likes3.8k downloads2mo agoHugging Face13AMLGentex /sweden_100K_difficulttabular10M<n<100M0 likes3.2k downloads1y agoHugging Face14Whiteboat /MLIC-Train-100K Abstract The latent representation in learned image compression encompasses channel-wise, local spatial, and global spatial correlations, which are essential for the entropy model to capture for conditional entropy minimization. Efficiently capturing these contexts within a single entropy model, especially in high-resolution image coding, presents a challenge due to the computational complexity of existing global context modules. To address this challenge, we propose the Linear… See the full description on the dataset page: https://huggingface.co/datasets/Whiteboat/MLIC-Train-100K.100K<n<1M3 likes3.1k downloads1y agoHugging Face15ManiTwin /ManiTwin-100K ManiTwin-100K: Manipulation-Ready Digital Object Twins Project Page | Paper ManiTwin-100K is a large-scale dataset of manipulation-ready digital object twins designed for robotic manipulation research. Each object includes simulation-ready 3D meshes, physical properties, functional point annotations, grasp configurations, and rich language descriptions—all validated through physics-based simulation. Note: We are currently releasing approximately 1K sample objects with a… See the full description on the dataset page: https://huggingface.co/datasets/ManiTwin/ManiTwin-100K.robotics100K<n<1M13 likes2.8k downloads6mo agoHugging Face16qihoo360 /RevealLayer-100K RevealLayer Open Dataset RevealLayer Open is the open-source dataset accompanying RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition. Paper: https://arxiv.org/html/2605.11818v1 Accepted by ICML 2026 RevealLayer studies box-guided layered image decomposition for natural images. Given an RGB image and instance bounding boxes, the task is to decompose the scene into a clean background and object-level foreground layers, where each… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/RevealLayer-100K.imageimage-to-image1M<n<10M7 likes2.4k downloads4mo agoHugging Face17Dhanush944 /ManiTwin-100K ManiTwin-100K: Manipulation-Ready Digital Object Twins Project Page | Paper ManiTwin-100K is a large-scale dataset of manipulation-ready digital object twins designed for robotic manipulation research. Each object includes simulation-ready 3D meshes, physical properties, functional point annotations, grasp configurations, and rich language descriptions—all validated through physics-based simulation. Note: We are currently releasing approximately 1K sample objects with a… See the full description on the dataset page: https://huggingface.co/datasets/Dhanush944/ManiTwin-100K.robotics100K<n<1M0 likes2k downloads6mo agoHugging Face18Coraxor /JAVEdit-100k Yinan Chen 1★ · Chuming Lin 2★ · Zhennan Chen 3 · Yuxiang Zeng 4 · Junwei Zhu 2 · Yali Bi 1 · Xijie Huang 5 · Chengming Xu 2 · Donghao Luo 2 · Zhucun Xue 1 · Xiaobin Hu 6 · Chengjie Wang 2 · Yong Liu 1 · Jiangning Zhang 1,2 📧 · Shuicheng Yan 6 1 Zhejiang University 2 YouTu Lab, Tencent 3 Nanjing University 4 University of Auckland 5 Fudan University 6 National University of Singapore 😊 Dataset Introduction JAVEdit-100k is the official dataset… See the full description on the dataset page: https://huggingface.co/datasets/Coraxor/JAVEdit-100k.videotext-to-videon<1K14 likes1.9k downloads4mo agoHugging Face19Xkev /LLaVA-CoT-100k Dataset Card for LLaVA-CoT The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.textvisual-question-answering10K<n<100K106 likes1.9k downloads9mo agoHugging Face20chatelet /political-leaning-tweets-100k 🗳️ political-leaning-tweets-100k Châtelet AI presents a 100,000+ dataset of tweets labelled for political leaning: neutral, liberal, conservative.Labels are machine-generated using a SOTA thinking-enabled LLM. The dataset is intended for research on political language modelling, ideology detection, robustness, and safety evaluation. 📦 Dataset Card Name: chatelet/political-leaning-tweets-100k Publisher: Châtelet AI Licence: MIT with additional restrctions against… See the full description on the dataset page: https://huggingface.co/datasets/chatelet/political-leaning-tweets-100k.texttext-classification100K<n<1M2 likes1.8k downloads1y agoHugging Face21HuggingFaceTB /cosmopedia-100k Dataset description This is a 100k subset of Cosmopedia dataset. A synthetic dataset of textbooks, blogposts, stories, posts and WikiHow articles generated by Mixtral-8x7B-Instruct-v0.1. Here's how you can load the dataset from datasets import load_dataset ds = load_dataset("HuggingFaceTB/cosmopedia-100k", split="train") text100K<n<1M49 likes1.8k downloads3y agoHugging Face22atom-in-the-universe /libgen-10k-100k0 likes1.7k downloads3y agoHugging Face23BoyangZ /VisualGenome_VG_100K_1_and_2 image0 likes1.6k downloads2y agoHugging Face24Elriggs /openwebtext-100k Dataset Card for "openwebtext-100k" More Information needed text100K<n<1M8 likes1.4k downloads3y agoHugging Face25kucklily /Q-Eval-100K Q-Eval-100K Dataset (CVPR 2025 Oral) 📝 Introduction The Q-Eval-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). We utilize multiple popular text-to- image and text-to-video models to ensure diversity, which include FLUX, Lumina-T2X, PixArt, Stable Diffusion 3, Stable Diffusion XL, DALL·E 3, Wanx, Midjourney, Hunyuan-DiT… See the full description on the dataset page: https://huggingface.co/datasets/kucklily/Q-Eval-100K.0 likes1.4k downloads7mo agoHugging Face26nvidia /vipe-dynpose-100kpp ViPE Dataset Release This dataset contains the camera pose, depth, and intrinsics estimated using ViPE. For more details of the dataset, please refer to the Github link. Please consider citing the following paper if you found this dataset helpful: @article{huang2025vipe, title={Vipe: Video pose engine for 3d geometric perception}, author={Huang, Jiahui and Zhou, Qunjie and Rabeti, Hesam and Korovko, Aleksandr and Ling, Huan and Ren, Xuanchi and Shen, Tianchang and Gao, Jun and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/vipe-dynpose-100kpp.text10K<n<100K9 likes1.4k downloads1y agoHugging Face27Andrew613 /PICA-100K PICABench: How Far Are We from Physically Realistic Image Editing? Paper | Project Page | Code Overview PICABench probes how far current editing models are from physically realistic image manipulation. It ties together: PICABench benchmark – physics-aware editing cases spanning eight laws across Optics, Mechanics, and State Transition, each labeled with superficial/intermediate/explicit difficulty tiers. PICAEval metric – region-grounded, QA-based verification with… See the full description on the dataset page: https://huggingface.co/datasets/Andrew613/PICA-100K.imageimage-to-image100K<n<1M7 likes1.3k downloads11mo agoHugging Face28Yuanhao-Harry-Wang /fitvto-100k FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On The official preview dataset from the paper "FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On". This dataset supports garment-centric virtual try-on and try-off research, containing 100,000 training and 5,000 evaluation triplets. Each sample pairs a person image with a layflat garment image and body/garment measurements. Dataset Structure Each split (train / eval) contains four aligned modalities — all… See the full description on the dataset page: https://huggingface.co/datasets/Yuanhao-Harry-Wang/fitvto-100k.imageimage-to-image100K<n<1M10 likes1.3k downloads3mo agoHugging Face29ouasdg /vids-100k0 likes1.3k downloads2y agoHugging Face30lmarena-ai /arena-human-preference-100k Overview This dataset contains leaderboard conversation data collected between June 2024 and August 2024. It includes English human preference evaluations used to develop Arena Explorer. Additionally, we provide an embedding file, which contains precomputed embeddings for the English conversations. These embeddings are used in the topic modeling pipeline to categorize and analyze these conversations. For a detailed exploration of the dataset and analysis methods, refer to the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-100k.tabular100K<n<1M49 likes1.3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.