datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Belle_1.4M-SLAM-Omni
Belle_1.4M
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round Chinese spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/Belle_1.4M-SLAM-Omni.UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.VoiceAssistant-400K-SLAM-Omni
VoiceAssistant-400K (Modified)
This dataset is prepared for the reproduction of SLAM-Omni.
This is a single-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/VoiceAssistant-400K-SLAM-Omni.Belle_1.4M-SLAM-Omni
Belle_1.4M
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round Chinese spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/Belle_1.4M-SLAM-Omni.slam-asrUltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/UltraChat-300K-SLAM-Omni.VoiceAssistant-400K-SLAM-Omni
VoiceAssistant-400K (Modified)
This dataset is prepared for the reproduction of SLAM-Omni.
This is a single-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These… See the full description on the dataset page: https://huggingface.co/datasets/mwei/VoiceAssistant-400K-SLAM-Omni.fold_new_slamThis dataset was created using LeRobot.
Dataset Description
Data Distribution Overview
This figure summarizes the data distribution of the ywxia/fold_new_slam dataset, auto-generated after each conversion via analysis/postprocess_with_overview.py. It shows episode-length distribution, the 3-D EEF workspace, per-dimension state histograms, per-arm action magnitudes, and a sample of frames from each camera.
Task: fold the box on the desk
Episodes: 59 | Frames: 21264 |… See the full description on the dataset page: https://huggingface.co/datasets/ywxia/fold_new_slam.fold_box_dual_arm_60_slamThis dataset was created using LeRobot.
Dataset Description
Data Distribution Overview
This figure summarizes the data distribution of the ywxia/fold_box_dual_arm_60_slam dataset, auto-generated after each conversion via analysis/postprocess_with_overview.py. It shows episode-length distribution, the 3-D EEF workspace, per-dimension state histograms, per-arm action magnitudes, and a sample of frames from each camera.
Task: fold the box on the desk
Episodes: 59 |… See the full description on the dataset page: https://huggingface.co/datasets/ywxia/fold_box_dual_arm_60_slam.test-slam-axis-fixThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 3,
"total_frames": 1058,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/simheo/test-slam-axis-fix.SLAM_project
Omni Instrument SLAM Project Dataset
The Omni Instrument SLAM Project Dataset is a compact robotics dataset designed for evaluating stereo, visual-inertial, and visual-inertial odometry (VIO) pipelines.
It provides:
Stereo Image Pairs
Inertial measurements (IMU)
Ground-truth 6 DoF pose (for VIO)
Raw ROS 1 and ROS 2 recordings
Overview
The dataset is structured into three splits:
Split
Description
stereo
Stereo-only (IMU stationary)
stereoinertial… See the full description on the dataset page: https://huggingface.co/datasets/OmniInstrument/SLAM_project.slam-es-en-knowledge-tracing
SLAM Spanish-English Knowledge Tracing
This repository contains a Hugging Face-compatible conversion of the Spanish-English track from the 2018 Second Language Acquisition Modeling (SLAM) shared task.
Each row represents one learner exercise attempt. The original token-level mistake labels are retained in token_labels: 0 means OK and 1 means a mistake. The correct field is strict whole-attempt correctness: it is 1 only when every token label is 0.
This dataset is suitable for… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/slam-es-en-knowledge-tracing.slam_pickandplaceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 26651,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jeongsu-moon/slam_pickandplace.slam-en-es-knowledge-tracing
SLAM English-Spanish Knowledge Tracing
This repository contains a Hugging Face-compatible conversion of the English-Spanish track from the 2018 Second Language Acquisition Modeling (SLAM) shared task. Each row represents one learner exercise attempt. Source reference-token, linguistic, and official mistake-label annotations are preserved.
The dataset supports longitudinal knowledge tracing: the same anonymous learner can have many interactions over multiple relative days, with… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/slam-en-es-knowledge-tracing.it-slama-4MULTI_VALUE_qqp_perfect_slam
Dataset Card for "MULTI_VALUE_qqp_perfect_slam"
More Information needed
it-slama-10slam-en-es-vietnamese-prompts
SLAM English-Spanish Knowledge Tracing with Vietnamese Prompts
This dataset is a derived version of bihungba1101/slam-en-es-knowledge-tracing. It preserves every source row and field and adds prompt_vi, a machine-generated Vietnamese translation of the Spanish prompt.
The source track contains English learners who already speak Spanish. Consequently:
prompt is the original Spanish text shown to the learner when text is available.
prompt_vi is a Vietnamese translation of prompt.… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/slam-en-es-vietnamese-prompts.MULTI_VALUE_mnli_perfect_slam
Dataset Card for "MULTI_VALUE_mnli_perfect_slam"
More Information needed
MULTI_VALUE_mrpc_perfect_slam
Dataset Card for "MULTI_VALUE_mrpc_perfect_slam"
More Information needed
MULTI_VALUE_cola_perfect_slam
Dataset Card for "MULTI_VALUE_cola_perfect_slam"
More Information needed
MULTI_VALUE_sst2_perfect_slam
Dataset Card for "MULTI_VALUE_sst2_perfect_slam"
More Information needed
MULTI_VALUE_wnli_perfect_slam
Dataset Card for "MULTI_VALUE_wnli_perfect_slam"
More Information needed
MULTI_VALUE_rte_perfect_slam
Dataset Card for "MULTI_VALUE_rte_perfect_slam"
More Information needed
g1_grasp_bottleMULTI_VALUE_stsb_perfect_slam
Dataset Card for "MULTI_VALUE_stsb_perfect_slam"
More Information needed
it-slama-cs
