CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hoshipu /roboreal_data8 likes861k downloads5mo agoHugging Face02google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K258 likes488k downloads3y agoHugging Face03jat-project /jat-dataset-tokenized Dataset Card for "jat-dataset-tokenized" More Information needed timeseries10M<n<100M32 likes472k downloads3y agoHugging Face04jat-project /jat-dataset JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.imagereinforcement-learning100M<n<1B78 likes410k downloads3y agoHugging Face05applied-ai-018 /peacock-data-public-datasets-idc0 likes376k downloads2y agoHugging Face06jhu-clsp /ettin-pretraining-data Ettin Pre-training Data Phase 1 of 3: Diverse pre-training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre-training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description DCLM 837.2 49.1% High-quality web crawl data CC Head 356.6… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/ettin-pretraining-data.text-generation10 likes276k downloads1y agoHugging Face07challenge-2026 /challenge_data PrimeBot Household Bimanual Manipulation Challenge Dataset 中文 | English 中文 目录 关于我们 更新日志 真机遥操作数据 训练集说明 验证集说明 数据集字段说明 URDF 图像 语言指令 本体感知与动作 机器人推理接口 UMI数据 数据概览 目录结构 数据集字段说明 图像 本体感知与动作 索引字段 标注与 IMU 关于我们 我们来自上纬新材-启元研究院,我们的使命是加速个人机器人时代到来,加速家用机器人时代到来。我们开源高质量面向家庭操作的双臂操作数据集,同时开放机器人硬件描述以供可视化、可复现研究。 如果本数据集对您的工作有帮助,感谢引用: @misc{xu2026scalingbimanualhouseholdmanipulation, title={Scaling Bimanual Household Manipulation from 1,500… See the full description on the dataset page: https://huggingface.co/datasets/challenge-2026/challenge_data.12 likes268k downloads19h agoHugging Face08ACERobotics /ACE-Data-0 ACE-Data-0 Human-Centric Ambient Capture as Embodied Data Engine S-Lab, Nanyang Technological University, Singapore &nbsp;·&nbsp; ACE Robotics ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI. ▶ Demo video &nbsp;·&nbsp; Full story, figures, and interactive examples on the blog What this is Learning to act in the physical… See the full description on the dataset page: https://huggingface.co/datasets/ACERobotics/ACE-Data-0.videorobotics10K<n<100K47 likes248k downloads11d agoHugging Face09google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M40 likes220k downloads3y agoHugging Face10huggingface /DEH-image-scan-data21 likes206k downloads4m agoHugging Face11GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K45 likes183k downloads25d agoHugging Face12pjpjq /blofin-oi-data6 likes175k downloads5mo agoHugging Face13IPEC-COMMUNITY /fractal20220817_data_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "google_robot", "total_episodes": 87212, "total_frames": 3786400, "total_tasks": 599, "total_videos": 87212, "total_chunks": 88, "chunks_size": 1000, "fps": 3, "splits": { "train": "0:87212" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.videorobotics13 likes159k downloads2y agoHugging Face14boltzgen /inference-data0 likes156k downloads1y agoHugging Face15community-datasets /quarel Dataset Card for "quarel" Dataset Summary QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 0.63 MB Size of the generated dataset: 1.53 MB Total amount of disk used: 2.17 MB An example of 'train'… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/quarel.text1K<n<10K2 likes140k downloads2y agoHugging Face16mvp-lab /LLaVA-OneVision-2-Data LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. At a Glance The dataset is split across two Hugging Face repositories because of its size: Repository What it contains Part 1 (this repository) ~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.imagevideo-text-to-textn<1K39 likes131k downloads23d agoHugging Face17jzr99 /mesh4d_dataset3d100K<n<1M0 likes124k downloads11mo agoHugging Face18vincewin /CREST_data CREST forcing (parquet) EF5/CREST hourly forcing for CONUS, 2016-present, packed as one tar per variable/year. dir variable source cadence mrms/ precipitation MRMS QPE (corrected) hourly temp/ 2 m temperature NLDAS-2 FORA hourly pet/ potential ET FEWS NET daily PET daily Each *.tar expands to individual .pqf (Apache Arrow parquet) grids readable by the EF5 v4.5 native parquet reader. Used by the Space vincewin/CREST_AI. Download + extract one year, e.g.: from… See the full description on the dataset page: https://huggingface.co/datasets/vincewin/CREST_data.5 likes122k downloads3m agoHugging Face19hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes119k downloads2y agoHugging Face20llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes117k downloads2y agoHugging Face21defeatbeta /yahoo-finance-data The Financial data from Yahoo! *** Key Points to Note *** All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes. I will update the data regularly, and you are welcome to follow this project and use the data. Each time the data is updated, I will record the update time in spec.json. Data Usage Instructions Use DuckDB… See the full description on the dataset page: https://huggingface.co/datasets/defeatbeta/yahoo-finance-data.100M<n<1B126 likes111k downloads18h agoHugging Face22adams-story /datacomp200m Datacomp200m This is a smaller version of the datacomp_1b dataset. Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows. The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling. Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.image100M<n<1B3 likes108k downloads3y agoHugging Face23wegrthj /e94fjt-v654-data9 likes108k downloads4mo agoHugging Face24Salesforce /lotsa_data LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for time series forecasting. It was collected for the purpose of pre-training Large Time Series Models. See the paper and codebase for more information. Citation If you're using LOTSA data in your research or applications, please cite it using this BibTeX: BibTeX: @article{woo2024unified, title={Unified Training of Universal Time Series Forecasting Transformers}… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lotsa_data.text1M<n<10M97 likes96k downloads2y agoHugging Face25datania /boedocument10K<n<100K1 likes94k downloads20h agoHugging Face26YWZBrandon /webshop-data4 likes94k downloads1y agoHugging Face27mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes87k downloads3y agoHugging Face28tencent /Hy-Embodied-0.5-VLA-Data Hy-Embodied-0.5-VLA From Vision-Language-Action Models to a Real-World Robot Learning Stack Tencent Robotics X × Tencent Hy Team 📖 Abstract We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.tabularroboticsn<1K23 likes81k downloads3mo agoHugging Face29eddmpython /dartlab-data DartLab 데이터 종목코드 하나로 읽는 한국 DART + 미국 SEC EDGAR 공시 데이터 Structured Korean (DART) and US (SEC EDGAR) disclosure data, ready as Parquet. 무엇인가요? DartLab이 한국 DART 전자공시와 미국 SEC EDGAR 공시를 종목코드 하나로 비교 가능한 표로 가공해 Parquet으로 올려둔 데이터셋입니다. 한국 전 상장사(약 2,700사)와 미국 주요 상장사(약 1,000사)의 재무제표, 사업보고서 본문, 정형 공시, 주가, 거시지표가 들어 있습니다. 이 데이터셋은 DartLab의 데이터 층입니다. dartlab.Company("005930")을 호출하면 라이브러리가 필요한 parquet을 여기서 자동으로 내려받습니다. 숫자는 원문 그대로 보존합니다(반올림·추정·보간 없음). 코드 없이도 바로 씁니다… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/dartlab-data.table-question-answering1M<n<10M14 likes81k downloads32m agoHugging Face30google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K127 likes80k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.