datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stl10wds_stl10xlerobot-head-v13-stl
XLeRobot 头部 V13 · 可打印 STL 全套
所有数字都是从导出的网格实测的(对每个三角形用散度定理算体积),
不是从 .scad 源码读的标称值。这两者在本项目中已被证实会不一致,以网格为准。
最后更新 2026-09-09。
一、打印什么
文件
体积
材质
说明
quarter0.stl
55.66 cm³
白色 PLA
先打这块 — 带眼睛,最复杂
quarter1.stl
44.38 cm³
白色 PLA
左耳罩安装台
quarter2.stl
52.28 cm³
白色 PLA
后梁板
quarter3.stl
44.51 cm³
白色 PLA
右耳罩安装台
cam_frame.stl
58.47 cm³
白色 PLA
相机框(领结),含 USB 出线口
core_link.stl
22.13 cm³
白色 PLA
内部承力链 — 整个头挂在它上面
skin0.stl
30.19 cm³
金属质感
外皮肤,含眼睛开口
skin1.stl
31.81… See the full description on the dataset page: https://huggingface.co/datasets/Suyang99/xlerobot-head-v13-stl.stl10
Dataset Card for "stl10"
More Information needed
doc-stl
[!NOTE]
Dataset origin: https://www.ortolang.fr/market/corpora/doc-stl
Description
Le projet DOC (Didactique, Oral, Corpus) de l’UMR STL CNRS 8163 - Université de Lille, est porté par Juliette DELAHAIE et Emmanuelle CANUT (Professeures des Université en Sciences du langage), en collaboration avec Antonio BALVET (Maitre de conférences en Sciences du langage), Driss SADOUN (Président de PostLab, Chercheur EMSE (IMT) et chercheur associé ERTIM/INALCO), Liliane SANTOS (Maitre de… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/doc-stl.Zephyrus
ZephyrusBench
ZephyrusBench is a weather-science benchmark released with the paper Zephyrus: An Agentic Framework for Weather Science. It contains 2,230 question-answer pairs across 49 tasks spanning geospatial reasoning, temporal reasoning, forecasting, simulation, climatology, and scientific question answering.Accepted at the International Conference on Learning Representations, 2026.
Paper and Resources
Paper: arXiv
Poster: ICLR 2026 Poster
Code: Rose-STL-Lab/Zephyrus… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/Zephyrus.stl10_binaryClimaQA
ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models (ICLR 2025)
Check the paper's webpage and GitHub for more info!
The ClimaQA benchmark is designed to evaluate Large Language Models (LLMs) on climate science question-answering tasks by ensuring scientific rigor and complexity. It is built from graduate-level climate science textbooks, which provide a reliable foundation for generating questions with precise terminology and complex scientific theories.… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/ClimaQA.drone_stl_env0000_v1_nogroundThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "uav",
"total_episodes": 500,
"total_frames": 100000,
"total_tasks": 5,
"total_videos": 1000,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Celina717/drone_stl_env0000_v1_noground.stocks-STLTECH-1D-candlesstl10sigma01_stl_gated_lerobotSimulCost-Bench
SimulCost-Bench
📖 Paper | 🛠️ Code | 🌐 Website | 💾 Cache (Baseline) | 💾 Cache (Full)
SimulCost is a cost-aware benchmark and toolkit for evaluating how well LLM agents tune simulation parameters under realistic computational budgets. Unlike prior evaluations that focus on correctness while implicitly treating tool usage as “free,” SimulCost explicitly measures both: (1) whether a proposed configuration meets an accuracy target and (2) how much simulation compute it consumes.The… See the full description on the dataset page: https://huggingface.co/datasets/Rose-STL-Lab/SimulCost-Bench.stl_formulae_variantsSTL-MCQA-results
STL Prompting for Zero-Shot MCQA
Results of zero-shot Multiple-Choice Question Answering (MCQA) experiments for the paper:
Dang, Q. P., Tran-Truong, P. T., Vu, D. L., Nguyen, L. S. T., Vo, Q. T. N., & Quan, T. (2026).
Enhancing large language model performance for automatic zero-shot multiple-choice question answering via single-token logit prompting.
Computers and Education: Artificial Intelligence. DOI: 10.1016/j.caeai.2026.100578
Source code:… See the full description on the dataset page: https://huggingface.co/datasets/p-storm/STL-MCQA-results.stl_high_complexitystl10
Dataset Card for "stl10"
More Information needed
eval_act_stl_4cam_viewdrop_v060_test_20260923_162208This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/KoukiHagiwara/eval_act_stl_4cam_viewdrop_v060_test_20260923_162208.stl_prompt_sft_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "uav",
"total_episodes": 63,
"total_frames": 6300,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:63"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Celina717/stl_prompt_sft_lerobot.stl10
Dataset Card for "stl10"
More Information needed
SWIR-Fruit_and_Vegetable_DatasetThe RGB-SWIR Fruits and Vegetables Dataset is published in this paper.
Citation
If this work has been helpful to you, please cite our paper.
@article{songswir2024,
AUTHOR = {Song, Hanbin and Yeo, Sanghyeop and Jin, Youngwan and Park, Incheol and Ju, Hyeongjin and Nalcakan, Yagiz and Kim, Shiho},
TITLE = {Short-Wave Infrared (SWIR) Imaging for Robust Material Classification: Overcoming Limitations of Visible Spectrum Data},
JOURNAL = {Applied Sciences},
VOLUME =… See the full description on the dataset page: https://huggingface.co/datasets/STL-Yonsei/SWIR-Fruit_and_Vegetable_Dataset.stl_formulaeThis dataset contains the data used in the paper "Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers" (Candussio et al. 2025).
More in detail:
train.csv is the training set for the random models. It contains formulae spanning from depth 2 to depth 23;
TODO: complete with test set.
cube_pick_testThis dataset was created using LeRobot.
STL_dataset_5.12STL_dataset_5.12_maxminZxcJvmRASMD
RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions
Current autonomous driving algorithms heavily rely on the visible spectrum, which is prone to performance degradation in adverse conditions like fog, rain, snow, glare, and high contrast. Although other spectral bands like near-infrared (NIR) and long-wave infrared (LWIR) can enhance vision perception in such situations, they have limitations and lack large-scale datasets and benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/STL-Yonsei/RASMD.so101_cube_pick6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 10,
"total_frames": 1000,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/stlee601/so101_cube_pick6.stl_dataset_5_12_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "uav",
"total_episodes": 160,
"total_frames": 16000,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:160"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Celina717/stl_dataset_5_12_lerobot.
