datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Atlas-Orion
X-Atlas/Orion
X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human
protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular
identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.datacomp_xlarge
DataComp XLarge Pool
This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.LongBenchX-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.seamless-align-enA-viA.speaker-embedding.xlsr-2bX-Atlas-Orion
X-Atlas Orion Dataset (SLAF Format)
Attribution
This is a re-release of data originally generated by Xaira Therapeutics.
Original Dataset: Xaira-Therapeutics/X-Atlas-Orion
Original Format: Parquet files
This Release: Same data in SLAF (Sparse Lazy Array Format)
License: CC-BY-NC-SA-4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0)
Original Citation:
@article{huang2025xatlasorion,
title={X-Atlas/Orion: Genome-wide Perturb-seq Datasets via a Scalable… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/X-Atlas-Orion.seamless-align-enA-frA.speaker-embedding.hubert-xlxcopa
Dataset Card for "xcopa"
Dataset Summary
XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning
The Cross-lingual Choice of Plausible Alternatives dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across
languages. The dataset is the translation and reannotation of the English COPA (Roemmele et al. 2011) and covers 11 languages from 11 families and several areas around
the globe. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/xcopa.seamless-align-deA-enA.speaker-embedding.xlsr-2bseamless-align-enA-hiA.speaker-embedding.hubert-xlseamless-align-enA-frA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.xlsr-2bptb-xl-processedseamless-align-enA-zhA.speaker-embedding.xlsr-2bthe-stack-smol-xl
Dataset Description
A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset.
Languages
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c',
'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.xehe_rezako_fokjhdata_4
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds4/data_4.seamless-align-enA-zhA.speaker-embedding.hubert-xlpa-warm-start-sft-xl-50b-mix
geodesic-research/pa-warm-start-sft-xl-50b-mix
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.data_3
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds3/data_3.ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.data_2
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds2/data_2.data_1
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds1/data_1.seamless-align-enA-jaA.speaker-embedding.xlsr-2bseamless-align-enA-hiA.speaker-embedding.xlsr-2bxarm_lift_mediumThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 800,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/xarm_lift_medium.seamless-align-enA-jaA.speaker-embedding.hubert-xlmmu-xmatch-20m
MMU 20M positional crossmatch spine
This release contains 20,415,199 unique DESI DR1 TARGETID anchors and
recomputed positional associations to Multimodal Universe HATS catalogs. It is
a reproducible positional crossmatch product, not a claim that every association
is a certain physical identity.
Recommended views
spine/ — every anchor with positional and conservative secure modality flags.
spine_positional_multimodal/ — 4,184,585 anchors with
at least one… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-xmatch-20m.the-stack-smol-xs\HR-VILAGE-3K3M
HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression
This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.
