CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simple-world-lab /HiFi-UMI-2K HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data 2,000 hours released · 6 synchronized camera views · 480+ scenes · 3 mm pose accuracy · <40 µs synchronization 🌐 Project Website | 📦 Dataset | 📄 Paper: arXiv:2607.25895 Examples from the HiFi-UMI corpus. Click the image to play the video. 📚 Introduction HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.… See the full description on the dataset page: https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K.tabularrobotics100M<n<1B55 likes113k downloads2mo agoHugging Face02mhenrichsen /alpaca_2k_testtext1K<n<10K27 likes34k downloads3y agoHugging Face03DL3DV /DL3DV-ALL-2Kgated DL3DV-Dataset This repo has all the 2K frames with camera poses of DL3DV-10K Dataset. We are working hard to review all the dataset to avoid sensitive information. Thank you for your patience. Download If you have enough space, you can use git to download a dataset from huggingface. See this link. 480P/960P versions should satisfies most needs. If you do not have enough space, we further provide a download script here to download a subset. The usage: usage: download.py… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-ALL-2K.n>1T6 likes18k downloads2y agoHugging Face04ngailapdi /MEBench-2K1Uimage10K<n<100K0 likes11k downloads1y agoHugging Face05RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes4.7k downloads7mo agoHugging Face06Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.40546875 Action score: 0.475 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face07Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face08Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.39921875 Action score: 0.44375 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face09Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38359375 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face10LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K33 likes2.8k downloads6mo agoHugging Face11filipwx /the-secrets-of-ceos-book-2k The-Secrets-Of-Ceos-Book-2k Made with ❤️ using 🦥 Unsloth Studio Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 2,000 📋 Columns: 4 📋 Schema & Statistics Column Type Column Type Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.text1K<n<10K0 likes2.7k downloads5mo agoHugging Face12fozziethebeat /alpaca_messages_2k_dpo_testtext1K<n<10K2 likes2.2k downloads2y agoHugging Face13HuggingFaceTB /everyday-conversations-llama3.1-2k Everyday conversations for Smol LLMs finetunings This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic. The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include: 20 everyday topics with 100 subtopics each 43 elementary science topics with 10… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k.text1K<n<10K138 likes2k downloads2y agoHugging Face14maxspeer /assessment2_spheres_and_cube_2k Dataset Card for cilp_assessment_all This is a FiftyOne dataset with 2000 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("maxspeer/assessment2_spheres_and_cube_2k_2") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/maxspeer/assessment2_spheres_and_cube_2k.imageimage-classification1K<n<10K0 likes1.4k downloads9mo agoHugging Face15eddyfox8812 /ai-vs-real-2k-imagesimage1K<n<10K0 likes719 downloads7mo agoHugging Face16BitRobot /FrodoBots-2K FrodoBots 2K Dataset The FrodoBots 2K Dataset is a diverse collection of camera footage, GPS, IMU, audio recordings & human control data collected from ~2,000 hours of tele-operated sidewalk robots driving in 10+ cities. This dataset is collected from Earth Rovers, a global scavenger hunt "Drive to Earn" game developed by FrodoBots Lab. Please join our Discord for discussions with fellow researchers/makers! If you're interested in contributing driving data, you can buy your own… See the full description on the dataset page: https://huggingface.co/datasets/BitRobot/FrodoBots-2K.reinforcement-learning16 likes666 downloads1y agoHugging Face17dmarsili /RSVQA-LR-2kA 2k subset of the validation split of the RSVQA LR dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation. For more information and attribution please refer to the original dataset: https://rsvqa.sylvainlobry.com/#dataset imagevisual-question-answering1K<n<10K0 likes594 downloads3mo agoHugging Face18Yuheng02 /dl3dv_2k0 likes592 downloads25d agoHugging Face19mlfoundations-dev /sci_question_exp__scp_116k__training_2k_for_GPQAtext100K<n<1M1 likes572 downloads2y agoHugging Face20sameedkhan /medconceptsqa-sample_medarc_2k Dataset Card for MedConceptsQA 2K Sample for MedARC This is a hierarchically stratified sample of the ICD10-CM coding system from MedConceptsQA. This dataset was sampled such that maximum coverage across all levels of the ICD-10 hierarchy to be as representative as possible while constraining the dataset to a reasonable size for evaluation and potential reinforcement learning. Users can generate their own sampling of MedConceptsQA via this generator script. Original MedConceptsQA… See the full description on the dataset page: https://huggingface.co/datasets/sameedkhan/medconceptsqa-sample_medarc_2k.texttext-classification10K<n<100K0 likes565 downloads11mo agoHugging Face21nachid /CHERRY-2K0 likes530 downloads10mo agoHugging Face22Eedi /Question-Anchored-Tutoring-Dialogues-2k Question-Anchored-Tutoring-Dialogues-2k This dataset contains dialogues from math tutoring interventions recorded on Eedi. Dataset Details Dataset Description Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data: DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.tabulartext-generation10K<n<100K10 likes488 downloads7mo agoHugging Face23Brench /R1_Annotated_AIME_True_2Ktext1K<n<10K0 likes479 downloads2y agoHugging Face24depth-anything /DA-2K DA-2K Evaluation Benchmark Introduction DA-2K is proposed in Depth Anything V2 to evaluate the relative depth estimation capability. It encompasses eight representative scenarios of indoor, outdoor, non_real, transparent_reflective, adverse_style, aerial, underwater, and object. It consists of 1K diverse high-quality images and 2K precise pair-wise relative depth annotations. Please refer to our paper for details in constructing this benchmark. Usage Please… See the full description on the dataset page: https://huggingface.co/datasets/depth-anything/DA-2K.image1K<n<10K17 likes476 downloads2y agoHugging Face25TigerResearch /tigerbot-kaggle-leetcodesolutions-en-2kTigerbot 基于leetcode-solutions数据集,加工生成的代码类sft数据集 原始来源:https://www.kaggle.com/datasets/erichartford/leetcode-solutions Usage import datasets ds_sft = datasets.load_dataset('TigerResearch/tigerbot-kaggle-leetcodesolutions-en-2k') text1K<n<10K18 likes470 downloads3y agoHugging Face26Voxel51 /deeplesion-balanced-2k DeepLesion Benchmark Subset (Balanced 2K) This dataset is a curated subset of the DeepLesion dataset, prepared for demonstration and benchmarking purposes. It consists of 2,000 CT lesion samples, balanced across 8 coarse lesion types, and filtered to include lesions with a short diameter > 10mm. Dataset Details Source: DeepLesion Institution: National Institutes of Health (NIH) Clinical Center Subset size: 2,000 images Lesion types: lung, abdomen, mediastinum, liver… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/deeplesion-balanced-2k.imageobject-detection1K<n<10K0 likes438 downloads1y agoHugging Face27LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes437 downloads8mo agoHugging Face28dmarsili /RSVQA-HR-2kA 2k subset of the validation split of the RSVQA HR dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation. For more information and attribution please refer to the original dataset: https://rsvqa.sylvainlobry.com/#dataset imagevisual-question-answering1K<n<10K0 likes410 downloads3mo agoHugging Face29hfwang01 /DL3DV-VIPED-final-2K3d1K<n<10K0 likes340 downloads5mo agoHugging Face30Henryoung /WRIT-2K WRIT-2K WRIT-2K is a 2,000-trajectory supervised fine-tuning dataset for multi-turn, tool-using customer-service agents on tau2-bench style airline and retail tasks. This dataset accompanies the paper WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents. Project homepage: https://hengrui-gu.github.io/WRIT/ Dataset Summary WRIT-2K contains complete multi-turn trajectories with user messages, assistant natural-language responses… See the full description on the dataset page: https://huggingface.co/datasets/Henryoung/WRIT-2K.texttext-generation1K<n<10K5 likes322 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.