CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Xaira-Therapeutics /X-Atlas-Orion X-Atlas/Orion X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.tabular1M<n<10M28 likes47k downloads1y agoHugging Face02mlfoundations /datacomp_xlarge DataComp XLarge Pool This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.image10B<n<100B21 likes40k downloads3y agoHugging Face03Xnhyacinth /LongBenchtabular1K<n<10K7 likes31k downloads1y agoHugging Face04wangyz1999 /X-EGO-CS X-Ego-CS Ten players. One match. Ten simultaneous first-person recordings, each paired with a 64 Hz stream of that player's exact keyboard, mouse and view-angle inputs — all on a common, measured clock. Paper · Paper code · Collection pipeline Cross-Ego Demo (Pistol Round) Your browser cannot play this video — download it instead. All ten players' points of view, from the same pistol round, on one clock. Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.tabularvideo-classification10K<n<100K2 likes30k downloads5d agoHugging Face05asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes20k downloads2y agoHugging Face06slaf-project /X-Atlas-Orion X-Atlas Orion Dataset (SLAF Format) Attribution This is a re-release of data originally generated by Xaira Therapeutics. Original Dataset: Xaira-Therapeutics/X-Atlas-Orion Original Format: Parquet files This Release: Same data in SLAF (Sparse Lazy Array Format) License: CC-BY-NC-SA-4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0) Original Citation: @article{huang2025xatlasorion, title={X-Atlas/Orion: Genome-wide Perturb-seq Datasets via a Scalable… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/X-Atlas-Orion.tabular10B<n<100B0 likes16k downloads8mo agoHugging Face07asahi417 /seamless-align-enA-frA.speaker-embedding.hubert-xltabular1M<n<10M0 likes16k downloads2y agoHugging Face08cambridgeltl /xcopa Dataset Card for "xcopa" Dataset Summary XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning The Cross-lingual Choice of Plausible Alternatives dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across languages. The dataset is the translation and reannotation of the English COPA (Roemmele et al. 2011) and covers 11 languages from 11 families and several areas around the globe. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/xcopa.tabularquestion-answering10K<n<100K22 likes15k downloads3y agoHugging Face09asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face10asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes10k downloads2y agoHugging Face11asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.8k downloads2y agoHugging Face12asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes9.6k downloads2y agoHugging Face13zackyabd /ptb-xl-processedtabular10K<n<100K0 likes9.2k downloads1y agoHugging Face14asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.5k downloads2y agoHugging Face15bigcode /the-stack-smol-xl Dataset Description A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.tabulartext-generation100K<n<1M11 likes8.4k downloads4y agoHugging Face16dwjoeikr /xehe_rezako_fokjhimage1B<n<10B1 likes8.1k downloads2mo agoHugging Face17xxxspatialencoderwds4 /data_4 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds4/data_4.tabularobject-detectionn<1K1 likes7.8k downloads6d agoHugging Face18asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes7.1k downloads2y agoHugging Face19geodesic-research /pa-warm-start-sft-xl-50b-mix geodesic-research/pa-warm-start-sft-xl-50b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.tabular10M<n<100M0 likes6.9k downloads13d agoHugging Face20xxxspatialencoderwds3 /data_3 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds3/data_3.tabularobject-detectionn<1K0 likes6.8k downloads6d agoHugging Face21cy0307 /ropedia-xperience-10m-task-suite-artifacts Ropedia Xperience-10M Task Suite Artifacts This dataset repository stores small derived artifacts for the Ropedia Xperience-10M task-suite project: metrics, predictions, manifests, reports, figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the Cosmos3-Nano plus Cosmos3-Super diagnostic packages. Project Identity The Project identity mark is shared across the GitHub README, GitHub Pages dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.tabular10K<n<100K1 likes6.2k downloads3mo agoHugging Face22xxxspatialencoderwds2 /data_2 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds2/data_2.tabularobject-detectionn<1K5 likes6.1k downloads6d agoHugging Face23xxxspatialencoderwds1 /data_1 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds1/data_1.tabularobject-detectionn<1K0 likes5.9k downloads6d agoHugging Face24asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.8k downloads2y agoHugging Face25asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes5.8k downloads2y agoHugging Face26lerobot /xarm_lift_mediumThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 800, "total_frames": 20000, "total_tasks": 1, "total_videos": 800, "total_chunks": 1, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:800" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_lift_medium.tabularrobotics10K<n<100K2 likes5.8k downloads1y agoHugging Face27asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M0 likes5.7k downloads2y agoHugging Face28kshitijd /mmu-xmatch-20m MMU 20M positional crossmatch spine This release contains 20,415,199 unique DESI DR1 TARGETID anchors and recomputed positional associations to Multimodal Universe HATS catalogs. It is a reproducible positional crossmatch product, not a claim that every association is a certain physical identity. Recommended views spine/ — every anchor with positional and conservative secure modality flags. spine_positional_multimodal/ — 4,184,585 anchors with at least one… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-xmatch-20m.tabular1K<n<10K0 likes5.6k downloads2mo agoHugging Face29bigcode /the-stack-smol-xs\tabulartext-generation1K<n<10K11 likes5.3k downloads4y agoHugging Face30xuejun72 /HR-VILAGE-3K3M HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.tabularzero-shot-classification1K<n<10K3 likes5.1k downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.