datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CortexJEPAData
CortexJEPAData
Minimal cortex spatial transcriptomics H5AD files prepared for CortexJEPA.
Data Contents
Each .h5ad file keeps:
X: expression matrix
obs_names: spot/cell identifiers
var_names: gene identifiers
obsm["spatial"]: spatial coordinates
fine-tuning labels only for selected training splits:
c.macaque/pretrain: obs["layer"]
d.marmoset/pretrain: obs["layer"] and obs["PrAl"]
developmental-stage grouping for d.marmoset/test/development: obs["segment"] only… See the full description on the dataset page: https://huggingface.co/datasets/BGI-Hangzhou-OmicsAI/CortexJEPAData.CORTEX
CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph
English | 简体中文
Abstract
The continuous evolution of large language models drives escalating demands on data scale and quality, and as different training stages impose increasingly tailored data requirements, systematic organization of high-quality corpora becomes indispensable. Existing corpus construction pipelines confine the resulting corpora… See the full description on the dataset page: https://huggingface.co/datasets/zjukg/CORTEX.tb2-k5-cd1-cortex-dsh-flashswe-forge
SWE-Forge Dataset
20 validated tasks for evaluating software engineering agents.
Each task contains:
workspace.yaml - Task configuration (repo, commits, install commands, test commands)
patch.diff - The ground-truth patch
tests/ - Generated test files (fail before patch, pass after)
evaluate.sh - Binary evaluator (score 0 or 1)
Docker Images
Pre-built images on Docker Hub: platformnetwork/swe-forge:<task_id>
Each image has the repo cloned at base_commit with… See the full description on the dataset page: https://huggingface.co/datasets/CortexLM/swe-forge.14062026-yambox1-fold-a-cloth-napkin-1MCAP-Housing
MCAP-Housing: Egocentric RGB-D Household Manipulation Dataset
MCAP-Housing is an egocentric RGB + Depth + IMU dataset of human household manipulation activities, packaged in robotics-native .mcap (ROS2) format. Designed for robotics research, policy learning, and embodied AI.
This is a sample release. We can scale to custom episode counts, new activities, and specific environments on request. Contact us to discuss your requirements.
Quick Facts
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/cortexdatalabs/MCAP-Housing.open-cortex-fx-v3
Open Cortex FX v3
A curated dataset of videos depicting human manual labor and physical work, organized by task categories.
Dataset Description
Standout's Cortex FX v3 is a video dataset focusing on human manual labor and physical work activities. Each video has been carefully annotated to identify work-related content and categorized into specific labor types.
Dataset Statistics
Total Videos: 708 videos showing human manual labor
Categories: 28 distinct labor… See the full description on the dataset page: https://huggingface.co/datasets/Standout/open-cortex-fx-v3.07022026-afternoonThis dataset was created using LeRobot.
Dataset Description
Trash Items: CRUSHED_TISSUE_PAPER, TOILET_PAPER_ROLL, CRUSHED_CAN, WHOLE_CAN, PLASTIC_CUP.
Non-trash items : 1 DETTOL, 1 AIR_FRESHNER,
Variation Set_B009 : PAGE 10 to 39
Reject episode: 2-29 (left arm joint data not recorded)
Sink #1 , rectangle sink
3 cameras (top, right, left)
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_piper",
"total_episodes": 9… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs-cortexai/07022026-afternoon.Cortexucsc-tracks-cortexeval_pi05_multi_packing_200k_03122026_snacks_02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_yam_follower",
"total_episodes": 25,
"total_frames": 92816,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cortexairobot/eval_pi05_multi_packing_200k_03122026_snacks_02.cortex-cafe-capsula-copoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ilustraviz/cortex-cafe-capsula-copo.cortex-data26052026-yambox-dispose-food-wasteThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_yam_follower",
"total_episodes": 5,
"total_frames": 15516,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cortexairobot/26052026-yambox-dispose-food-waste.eval_pi05_multi_packing_200k_03122026_snacks_01This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_yam_follower",
"total_episodes": 25,
"total_frames": 96326,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cortexairobot/eval_pi05_multi_packing_200k_03122026_snacks_01.cortex-codellamacortex_zdkidney-cortex-medulla-maskrcnn-predictionscortex-cafe-capsula-copo_20260917_121853This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ilustraviz/cortex-cafe-capsula-copo_20260917_121853.cortex-20b
cortex-20b
A 20 B-token (100 B-character), cap-balanced English/Japanese
pretraining mixture assembled for general-purpose language model
training. Release tag v0.9 (mixture build balanced-pretraining-v8,
finalized 2026-09-15).
The dataset is partitioned into three splits (train, validation,
test) and stores 22.4 M text records on disk in zstd-compressed
Parquet. The source stream spans web text, encyclopedias, news,
public-domain books, academic articles, patents, court… See the full description on the dataset page: https://huggingface.co/datasets/m8than/cortex-20b.cortex
CORTEX
A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
CORTEX (Clinically Organized Reasoning and sTructured EXplanation) is a structured reasoning benchmark for multimodal large language models working with 3D chest CT. It converts chest CT question-answering and report-generation tasks into auditable, radiologist-inspired reasoning traces.
This Hugging Face repository is the data release for the paper:
Hashmat Shadab Malik, Anees Ur Rehman… See the full description on the dataset page: https://huggingface.co/datasets/aneesurhashmi/cortex.cortex07032026-afternoon-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_piper",
"total_episodes": 10,
"total_frames": 17612,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs-cortexai/07032026-afternoon-1.cortex_z1cortex_zcortex-ibm-tabformer-embeddings
cortex embeddings — IBM TabFormer
Per-transaction hidden states (128-d, float32) produced by Neospace's cortex transaction
foundation model over every transaction in the public IBM TabFormer credit-card dataset
(card_transaction.v1.csv, 24,386,900 transactions, 2,000 cardholders, 1991–2020).
These embeddings let you reproduce the NeoLDM benchmark
without running cortex: feed them to the gradient-boosted-tree fraud classifier in that repo.
Contents
dir… See the full description on the dataset page: https://huggingface.co/datasets/luizcoroo/cortex-ibm-tabformer-embeddings.cortex-vire-datakidney-cortex-medulla-maskrcnn-predictions-2bimanual-yam-samplescortex-1-market-analysis
NEAR Cortex-1 Market Analysis Dataset
Dataset Summary
This dataset contains blockchain market analyses combining historical and real-time data with chain-of-thought reasoning. The dataset includes examples from Ethereum, Bitcoin, and NEAR chains, demonstrating high-quality market analysis with explicit calculations, numerical citations, and actionable insights.
The dataset has been enhanced with examples generated by GPT-4o and Claude 3.7 Sonnet, providing diverse… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/cortex-1-market-analysis.
