datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Curr-ReFT-data
Curr-ReFT-data
[📂 GitHub][📝 Paper]
[🤗 HF Dataset] [🤗 HF-Model: Curr-ReFT-3B]
[🤗 HF-Model: Curr-ReFT-7B]
Dataset Overview
Curr-ReFT-data contains training data for both stages of the Curr-ReFT methodology. The proposed Curr-ReFT post-training paradigm consists of two consecutive training stages: 1. Curriculum Reinforcement Learning: Gradually increasing task difficulty through reward mechanisms that match task complexity. 2. Rejected Sample based… See the full description on the dataset page: https://huggingface.co/datasets/ZTE-AIM/Curr-ReFT-data.realmirror-extra-datasets
Statement
This dataset is only used for data augmentation and not for training!
If you need to obtain benchmark train data for RealMirror, please refer to:
RealMirror train data
Citation
If you find this work helpful in your research, please consider giving this repo a star ⭐ and citing our paper:
@article{tai2025realmirror,
title={RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI},
author={Tai, Cong and Zheng, Zhaoyu and Long… See the full description on the dataset page: https://huggingface.co/datasets/zte-terminators/realmirror-extra-datasets.seed2scale-example-data
Seed2Scale Multi-GPU Data Generation Results
(Disclaimer: Due to capacity limitations, this warehouse only provides partial trajectory examples. For complete examples or business cooperation, please contact the corresponding author: shen.tao5@zte.com.cn)
This folder presents a real multi-GPU Seed2Scale data generation experiment conducted on a single workstation. The results summarize both data generation throughput and trajectory quality under different GPU configurations, with all… See the full description on the dataset page: https://huggingface.co/datasets/zte-terminators/seed2scale-example-data.realmirror-datasets
Statement
This dataset contains complete train data for the RealMirror benchmark.
Citation
If you find this work helpful in your research, please consider giving this repo a star ⭐ and citing our paper:
@article{tai2025realmirror,
title={RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI},
author={Tai, Cong and Zheng, Zhaoyu and Long, Haixu and Wu, Hansheng and Xiang, Haodong and Long, Zhengbin and Xiong, Jun and Shi, Rong and… See the full description on the dataset page: https://huggingface.co/datasets/zte-terminators/realmirror-datasets.realmirror-asset
Acknowledgments
Thank you to AGIBOT for granting open source authorization to the A2 robot model!
Citation
If you find this work helpful in your research, please consider giving this repo a star ⭐ and citing our paper:
@article{tai2025realmirror,
title={RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI},
author={Tai, Cong and Zheng, Zhaoyu and Long, Haixu and Wu, Hansheng and Xiang, Haodong and Long, Zhengbin and Xiong… See the full description on the dataset page: https://huggingface.co/datasets/zte-terminators/realmirror-asset.SpherIQ-datarealmirror-hdf5-raw-data32B_LLM_AdaptiveCode_data
English |
中文
datasets:
ZTE-AIM/32B_LLM_AdaptiveMath_data
ZTE-AIM/32B_LLM_AdaptiveCode_data
base_model:
DeepSeek-R1-Distill-Qwen-32B
32B_LLM_AdaptiveMath_data
[🤗 HF Dataset]
LLM-Adaptive-CoT-Code-data
[🤗 HF Dataset]
LLM-Adaptive-ZMath-model-32B
[🤗 LLM-Adaptive-ZMath-model-32B]
LLM-Adaptive-ZCode-model-32B
[🤗 LLM-Adaptive-ZCode-model-32B]
Model Overview
This work presents a fine-tuned reasoning model built on the… See the full description on the dataset page: https://huggingface.co/datasets/ZTE-AIM/32B_LLM_AdaptiveCode_data.32B_LLM_AdaptiveMath_data
English |
中文
datasets:
ZTE-AIM/32B_LLM_AdaptiveMath_data
ZTE-AIM/32B_LLM_AdaptiveCode_data
base_model:
DeepSeek-R1-Distill-Qwen-32B
32B_LLM_AdaptiveMath_data
[🤗 HF Dataset]
LLM-Adaptive-CoT-Code-data
[🤗 HF Dataset]
LLM-Adaptive-ZMath-model-32B
[🤗 LLM-Adaptive-ZMath-model-32B]
LLM-Adaptive-ZCode-model-32B
[🤗 LLM-Adaptive-ZCode-model-32B]
Model Overview
This work presents a fine-tuned reasoning model built on the… See the full description on the dataset page: https://huggingface.co/datasets/ZTE-AIM/32B_LLM_AdaptiveMath_data.zte_embodied_2026
ZTE Embodied 2026 Dataset (LeRobot v3.0)
Overview
This dataset is converted from the raw data of the 2026 16th ZTE Cup Global Elite Challenge - Algorithm Elite Challenge - Embodied Intelligence (Preliminary).
Competition Link: ZTE Challenge - Embodied Intelligence
Format: LeRobot v3.0
Conversion Script: scripts/datasets/convert_zte_to_lerobot.py
Dataset Info
Property
Value
Codebase Version
v3.0
Robot Type
Dual-arm Dexterous (dual 7-DOF arms +… See the full description on the dataset page: https://huggingface.co/datasets/hello3x3/zte_embodied_2026.seed2scale-example-assetsContains relevant assets used to reproduce the trajectory generated by Seed2Scale:
(1) Scenario asset: Collected_kitchen
(2) Robot asset: AgiBotA2
Telecom-Function-Calling-Evaluation
TFCE(Telecom Function-Calling Evaluation)数据集
数据集摘要
TFCE是一个评估通信领域函数调用能力的数据集,由1800余个函数构成917道Python题目,并应用于通信领域的Simple(简单函数)、Multiple(多函数)、Parallel(并行函数)、Parallel-Multiple(并行多函数)等场景,涉及4G LTE、5G技术与6G探索、无线通信与网络优化、物联网(IoT)与M2M通信、移动通信系统与实施、网络安全与协议等方面的内容。
语言
数据集中question的文本是中文;其他部分的文本是英文。
数据集结构
TECE数据集中的数据按照“question-function-required”的结构;其中,“function”由“name”、“description”、“parameters”组成;“parameters”由“type”和“properties”组成。… See the full description on the dataset page: https://huggingface.co/datasets/ZTE-AIM/Telecom-Function-Calling-Evaluation.NTele-R1-Datahf5934_ztest_candidate_01ztextstoryZtEiQmLEZte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22Zte1loiWIfsxLjmJ22
