CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IPEC-COMMUNITY /bridge_orig_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "widowx", "total_episodes": 53192, "total_frames": 1893026, "total_tasks": 19974, "total_videos": 212768, "total_chunks": 54, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:53192" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot.videorobotics27 likes165k downloads2y agoHugging Face02orionweller /generic_data_v20 likes57k downloads2y agoHugging Face03orionweller /mmBERT-pretrain-p2-fineweb2-remaining mmBERT Pre-training Data P2 Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite. NOTE: this is only P2 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p2-fineweb2-remaining.fill-mask0 likes46k downloads1y agoHugging Face04Xaira-Therapeutics /X-Atlas-Orion X-Atlas/Orion X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.tabular1M<n<10M28 likes45k downloads1y agoHugging Face05orionweller /ProLong-TextFull0 likes29k downloads2y agoHugging Face06orionweller /reddit_mds_incremental0 likes26k downloads2y agoHugging Face07orionweller /cc_en_middle_mds_incremental0 likes18k downloads2y agoHugging Face08slaf-project /X-Atlas-Orion X-Atlas Orion Dataset (SLAF Format) Attribution This is a re-release of data originally generated by Xaira Therapeutics. Original Dataset: Xaira-Therapeutics/X-Atlas-Orion Original Format: Parquet files This Release: Same data in SLAF (Sparse Lazy Array Format) License: CC-BY-NC-SA-4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0) Original Citation: @article{huang2025xatlasorion, title={X-Atlas/Orion: Genome-wide Perturb-seq Datasets via a Scalable… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/X-Atlas-Orion.tabular10B<n<100B0 likes15k downloads8mo agoHugging Face09orionweller /ProLongText0 likes15k downloads2y agoHugging Face10Mander0608 /scenetok_originvideo10K<n<100K0 likes7.6k downloads3mo agoHugging Face11orionweller /chunk-ext0 likes7.3k downloads2y agoHugging Face12orionweller /mmBERT-pretrain-p3-others mmBERT Pre-training Data P3 Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite. NOTE: this is only P3 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p3-others.fill-mask0 likes6.7k downloads1y agoHugging Face13ryanqian1994 /bridge_orig_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "widowx", "total_episodes": 53192, "total_frames": 1893026, "total_tasks": 19974, "total_videos": 212768, "total_chunks": 54, "chunks_size": 1000, "fps": 5, "splits": { "train": "0:53192" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ryanqian1994/bridge_orig_lerobot.videorobotics0 likes6.1k downloads9mo agoHugging Face14orionweller /old-data-store-real0 likes6.1k downloads2y agoHugging Face15ori-drs /oxford_spires_datasetWe present the Oxford Spires Dataset, captured in and around well-known landmarks in Oxford using a custom-built multi-sensor perception unit as well as a millimetre-accurate map from a terrestrial LiDAR scanner (TLS). The perception unit includes three global shutter colour cameras, an automotive 3D LiDAR scanner, and an inertial sensor — all precisely calibrated. Project page Paper Arxiv Video Code Sample Usage Download the Dataset You can download the dataset from… See the full description on the dataset page: https://huggingface.co/datasets/ori-drs/oxford_spires_dataset.robotics10 likes5.4k downloads4mo agoHugging Face16orionweller /mmBERT-pretrain-p1-fineweb2-langs mmBERT Pre-training Data P1 Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite. NOTE: this is only P1 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p1-fineweb2-langs.fill-mask7 likes5.4k downloads1y agoHugging Face17SharpaIT /Robotic_Origami_Challengegated Robotic Origami Challenge: Fold Plane Demonstrations Real-world LeRobot demonstrations for dexterous paper-airplane folding. Overview Robotic Origami Challenge: Fold Plane Demonstrations is a real-world teleoperation dataset for folding a paper airplane with a bimanual dexterous robot system. It is released by Sharpa in a LeRobot-compatible format for the Robotic Origami Challenge community. Origami is a demanding benchmark for embodied AI:… See the full description on the dataset page: https://huggingface.co/datasets/SharpaIT/Robotic_Origami_Challenge.video10K<n<100K7 likes5.1k downloads2mo agoHugging Face18originlab /brand-assetsimagen<1K2 likes5k downloads8d agoHugging Face19orionweller /mmBERT-pretraining-data-chunk1 mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretraining-data-chunk1.fill-mask0 likes4.7k downloads1y agoHugging Face20orionweller /tulu_flan_mds_incremental-tokens0 likes4.1k downloads2y agoHugging Face21fangqi /bridge_orig0 likes3.7k downloads1y agoHugging Face22oripress /AlgoTune Website  |   Paper   |   Code How good are language models at coming up with new algorithms? To try to answer this, we built a benchmark, AlgoTune, comprised of 154 widely used math, physics, and computer science functions. For each function, the goal is to write code that produces the same outputs as the original function, while being faster. In addition to the benchmark, we also provide an agent, AlgoTuner, which allows language models to easily optimize code.… See the full description on the dataset page: https://huggingface.co/datasets/oripress/AlgoTune.tabularn<1K1 likes3.7k downloads8mo agoHugging Face23ShareLab-SII /thinking_bridge_orig_lerobot_output_qwen3vl0 likes3.5k downloads6mo agoHugging Face24open-world-agents /D2E-Original D2E-Original Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 273.4 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents. What's included: Video + Audio: H.264 encoded at FHD/QHD 60fps with game audio. Input events: Keyboard… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/D2E-Original.videoroboticsn<1K4 likes3.4k downloads5mo agoHugging Face25orionweller /tulu_flan_mds_incremental0 likes3.4k downloads2y agoHugging Face26orionweller /refinedweb_mds_incremental0 likes3.1k downloads2y agoHugging Face27Emulated-Inc /ogb-full-original OGB full original archives Public, byte-for-byte mirror of 17 official Open Graph Benchmark (OGB) and OGB Large-Scale Challenge archive downloads used by the Emulated-Inc graph benchmark environment. Original ZIP archives are stored under archives//. Each dataset directory includes metadata.json with the authoritative source URL, exact byte size, SHA-256 digest, and repository archive path. Archives were transferred directly from the official SNAP/DGL hosts through ephemeral… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/ogb-full-original.0 likes2.9k downloads26d agoHugging Face28orionweller /cc_news_mds_incremental-tokens0 likes2.8k downloads2y agoHugging Face29orionweller /cc_en_head_mds_incremental0 likes2.7k downloads2y agoHugging Face30orion-ai-lab /S4ASen4AgriNet is a Sentinel-2 based time series multi country benchmark dataset, tailored for agricultural monitoring applications with Machine and Deep Learning. It is annotated from farmer declarations collected via the Land Parcel Identification System (LPIS) for harmonizing country wide labels. These declarations have only recently been made available as open data, allowing for the first time the labelling of satellite imagery from ground truth data. We proceed to propose and standardise a new crop type taxonomy across Europe that address Common Agriculture Policy (CAP) needs, based on the Food and Agriculture Organization (FAO) Indicative Crop Classification scheme. Sen4AgriNet is the only multi-country, multi-year dataset that includes all spectral information. It is constructed to cover the period 2016-2020 for Catalonia and France, while it can be extended to include additional countries.9 likes2.7k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.