datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stellar-blade-gameplay-data
剑星
This public dataset repository contains local gameplay data uploaded from F:\剑星.
Contents
Files: 225
Total local size: 167.48 GB
Generated: 2026-06-07 11:56:57 UTC
File Types
.jsonl: 67
.png: 57
.json: 51
.parquet: 17
.mkv: 17
.txt: 16
Notes
This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs.
The license is marked as other; review game footage, audio, and asset… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/stellar-blade-gameplay-data.H3-IR
H3-IR
H3-IR contains privacy-reviewed prompt/Context-IR pairs for training H3 prompt
enhancers. The public export is fail-closed: a row is included only when its
text, annotation, and every referenced media asset pass both privacy and
redistribution-rights gates.
Splits
Split
Rows
train
1110
validation
81
total
1191
Privacy Review
All source rows and unique visual assets were reviewed with gpt-5.6-sol at
reasoning_effort=xhigh… See the full description on the dataset page: https://huggingface.co/datasets/StellarVoyager/H3-IR.details_sequelbox__StellarBright
Dataset Card for Evaluation run of sequelbox/StellarBright
Dataset Summary
Dataset automatically created during the evaluation run of model sequelbox/StellarBright on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_sequelbox__StellarBright.meridian-stellar-cacheSTELAR-topo_vision_reasoning_SFT_50k
Stellar-Neuron/STELAR-topo_vision_reasoning_SFT_50k
[Paper] [HF Collection] [Project Page]
The dataset was released as part of STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision. STELAR is a more accurate, faster and greener intelligent system for Vision Language Reasoning.
Contact: chenli4@andrew.cmu.edu
Dataset Summary
This dataset was created by STELAR TopoAug from two base datasets: Math-V and VLM_S2H. Each question includes… See the full description on the dataset page: https://huggingface.co/datasets/Stellar-Neuron/STELAR-topo_vision_reasoning_SFT_50k.STELAR-topo_vision_reasoning_100k
Stellar-Neuron/STELAR-topo_vision_reasoning_100k
[Paper] [HF Collection] [Project Page]
The dataset was released as part of STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision. STELAR is a more accurate, faster and greener intelligent system for Vision Language Reasoning.
Contact: chenli4@andrew.cmu.edu
Dataset Summary
This dataset was created by STELAR TopoAug from two base datasets: Math-V and VLM_S2H. Each question includes responses… See the full description on the dataset page: https://huggingface.co/datasets/Stellar-Neuron/STELAR-topo_vision_reasoning_100k.details_Weyaxi__Stellaris-internlm2-20b-r512
Dataset Card for Evaluation run of Weyaxi/Stellaris-internlm2-20b-r512
Dataset automatically created during the evaluation run of model Weyaxi/Stellaris-internlm2-20b-r512 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Stellaris-internlm2-20b-r512.STELAR-topo_vision_reasoning_preference_123k
Stellar-Neuron/STELAR-topo_vision_reasoning_preference_123k
[Paper] [HF Collection] [Project Page]
The dataset was released as part of STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision. STELAR is a more accurate, faster and greener intelligent system for Vision Language Reasoning.
Contact: chenli4@andrew.cmu.edu
Dataset Summary
This dataset was created by STELAR TopoAug from two base datasets: Math-V and VLM_S2H.
Each question includes… See the full description on the dataset page: https://huggingface.co/datasets/Stellar-Neuron/STELAR-topo_vision_reasoning_preference_123k.stellarStellariumGaiagalah-dr4-stellar-abundances
GALAH DR4 — Stellar Abundances for 917k Stars
Credit: NASA/ESA/Hubble
Part of a dataset collection on Hugging Face.
Dataset description
The fourth data release of the GALactic Archaeology with HERMES (GALAH) survey, providing radial velocities, stellar parameters, and up to 31 elemental abundances for 917,588 stars observed with the HERMES spectrograph on the Anglo-Australian Telescope.
GALAH DR4 is one of the largest stellar spectroscopic surveys… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/galah-dr4-stellar-abundances.deep-multimodal-representation-learning-for-stellar-spectraDataset used in the paper "Deep Multimodal Representation Learning for Stellar Spectra".
Dataset of Milky Way stars based on selection from https://ui.adsabs.harvard.edu/abs/2024A&A...682A...9G,
which is based on ESA/Gaia/DPAC and APOGEE surveys.
This work has made use of data from the European Space Agency (ESA) mission Gaia (https://www.cosmos.esa.int/gaia),
processed by the Gaia Data Processing and Analysis Consortium (DPAC, https://www.cosmos.esa.int/web/gaia/dpac/consortium).
Funding… See the full description on the dataset page: https://huggingface.co/datasets/christianschwarz/deep-multimodal-representation-learning-for-stellar-spectra.Stellarx-4b-swe1
Dataset Card for "Stellarx-4b-swe1"
More Information needed
FoldPlanet-500
FoldPlanet-500折叠星球
衣物折叠In-the-wild Human数据集
Version: 1.0
Author: 上海星际硅途技术有限公司
Date: 2025-10-24
公司介绍(Company Introduction)
上海星际硅途技术有限公司,成立于2025年4月,2025年9月入驻上海人形机器人孵化器。
我们是一家具身智能数据解决方案服务商,致力于通过“动作捕捉+视觉感知+语义标注”的多模态技术,进行“in-the-wild”场景下的“Human Data”采集,建立通专融合、覆盖千行百业的数据生态,推动具身智能数据行业宽度和深度的发展,促进具身智能大模型的快速迭代。
数据集简介(Dataset Overview)
专为具身智能人形机器人训练而设计的,高质量、结构化、可学习的真实泛化场景叠衣动作数据集。
它旨在帮助模型学习人类的行为逻辑、操作方式、物体交互特征以及任务理解能力。… See the full description on the dataset page: https://huggingface.co/datasets/stellarnexrobotics/FoldPlanet-500.Stellar-Motion-WA2-Track1
Stellar-Motion
WorldArena 2.0 Track 1 submission package from Nanyang Technological
University, Singapore.
Model name: Stellar-Motion
Version: v2-rank08
Inference seed: 23
Contact: ziying.song@ntu.edu.sg
Number of videos: 1000
Resolution: 640 x 480
Frames per video: 121
Frame rate: 24 FPS
Files
Stellar-Motion_WA2_Track1_submission.tar.gz: submission archive
Stellar-Motion_WA2_Track1_submission.tar.gz.sha256: archive checksum… See the full description on the dataset page: https://huggingface.co/datasets/ZI-YING/Stellar-Motion-WA2-Track1.sspp-dr9-stellar-spectra
SDSS SSPP DR9 Stellar Spectra Dataset
Dataset Summary
This dataset contains 233,013 stellar spectra from the
Sloan Digital Sky Survey (SDSS) Stellar Parameter Pipeline (SSPP) DR9,
cross-matched with SEGUE1/SEGUE2 spectroscopic observations.
Each sample contains a flux-calibrated 1D spectrum (3,878 wavelength bins,
~3,820–9,200 Å) plus stellar atmospheric parameters measured by the SSPP.
Data Source
Field
Value
Survey
SDSS SEGUE (DR9)… See the full description on the dataset page: https://huggingface.co/datasets/OneAstronomy/sspp-dr9-stellar-spectra.stellar_blade_recordings_01
剑星 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_f2424870b371b67d324df2e6eaa56246
Collection: general (泛数据)
Recordings: 19
Layout: recordings/<recording_id>/<raw component>
demo
Stellaromics Pyxa demo dataset
Output from the Stellaromics Pyxa platform, used as a demo dataset for the
spatialdata-io pyxa reader
(experimental). No patient-identifiable information.
Datasets
Each folder is a self-contained dataset with the same layout.
Folder
Extent (x × y × z)
Cells
Transcripts
Size
Purpose
small/
~1310 × 443 × 130 µm
4,372
395,667
~290 MB
full demo region
xsmall/
100 × 100 × 100 µm crop
187
23,495
~8 MB
CI tests and quick visual… See the full description on the dataset page: https://huggingface.co/datasets/Stellaromics/demo.test_stellarStellarX-4b-SWE2
Dataset Card for "StellarX-4b-SWE2"
More Information needed
SII-Stellarator-Configuration-Dataset
SII Stellarator Configuration Dataset
Wenyang Li
Shanghai Innovation Institute | PhD student
AI-Driven Controlled Fusion Simulation, Control & Design Lab
Phone: +86-156-2003-5216
Email: lwydsg@mail.nankai.edu.cn
Overview
This directory hosts a stellarator configuration dataset grouped by number
of field periods (nfp). Each sample keeps only three things:
Fourier boundary coefficients, nfp, and 9 VMEC evaluation metrics —
every field is guaranteed to be non-None /… See the full description on the dataset page: https://huggingface.co/datasets/SII-AI4Fusion/SII-Stellarator-Configuration-Dataset.details_Dampish__StellarX-4B-V0
Dataset Card for Evaluation run of Dampish/StellarX-4B-V0
Dataset Summary
Dataset automatically created during the evaluation run of model Dampish/StellarX-4B-V0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Dampish__StellarX-4B-V0.Stellar-chat-full
Dataset Card for "StellarX-FULL"
More Information needed
gaia-dr3-young-stellar-objects
Gaia DR3 Young Stellar Objects
Part of the Astronomy Datasets collection on Hugging Face.
The Gaia DR3 young stellar object (YSO) catalog, containing 79,375 YSO candidates
identified by the ESA Gaia mission's variability classification pipeline. Each source includes
a YSO classification confidence score, variability statistics (amplitudes, standard deviations,
skewness, kurtosis), astrometry (positions, parallax, proper motions), and multi-band
photometry (G, BP, RP).… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/gaia-dr3-young-stellar-objects.stellar_blade_bc_01
剑星 BC parquet archives
Collection mode: general
Subset: default
Archives: 2
Encrypted bytes: 54281761060
Generated by the game data platform BC repository consolidator.
details_Weyaxi__Stellaris-internlm2-20b-r128
Dataset Card for Evaluation run of Weyaxi/Stellaris-internlm2-20b-r128
Dataset automatically created during the evaluation run of model Weyaxi/Stellaris-internlm2-20b-r128 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Stellaris-internlm2-20b-r128.near-axis-stellarators
NearAxisStellarators
The NearAxisStellarators dataset contains stellarator configurations generated with the pyQSC near-axis expansion code.It provides input design parameters (magnetic axis Fourier coefficients, field strength coefficients, number of field periods, pressure) and the resulting plasma properties (rotational transform, elongation, Mercier stability, quasisymmetry, etc.).
This dataset supports both forward modeling (parameters → properties) and inverse design (desired… See the full description on the dataset page: https://huggingface.co/datasets/pedrocurvo/near-axis-stellarators.stellar-classification-eda
Stellar Classification: Can we tell what's in space from telescope data?
Dataset Overview
I chose the Stellar Classification Dataset (SDSS17) from Kaggle, based on real data from the Sloan Digital Sky Survey. It contains 100,000 observations of celestial objects with 18 columns, mostly numeric measurements like light filters, sky coordinates, and redshift.
Main question: Can we classify whether a celestial object is a Star, Galaxy, or Quasar just from the numbers the… See the full description on the dataset page: https://huggingface.co/datasets/idoyaaran/stellar-classification-eda.repro-stellar-testing-framework-traces
Agent traces
Agent sessions published from a Trackio Logbook.
details_Dampish__StellarX-4B-V0.2
Dataset Card for Evaluation run of Dampish/StellarX-4B-V0.2
Dataset Summary
Dataset automatically created during the evaluation run of model Dampish/StellarX-4B-V0.2 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Dampish__StellarX-4B-V0.2.
