datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clean_data2gutenberg-en-v1-clean
gutenberg - clean
dataset_info:
- config_name: default
features:
- name: text
dtype: string
- name: label
dtype: string
- name: score
dtype: float64
- name: sha256dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 3384868097
num_examples: 9978
- name: validation
num_bytes: 195405579
num_examples: 574
- name: test
num_bytes: 189439446
num_examples: 565
download_size: 2317462261
dataset_size:… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/gutenberg-en-v1-clean.pytorch-issues-dataset-cleanpi-c0-c8-technique-dataset-clean_harmful_only_with_benign_tech2_self_reminder_v5_split70_30lma_clean_raw_datasetlma_tamil_clean_datasetclean_dataturkish-wikipedia-dataset-clean
Turkish Wikipedia Dataset
A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments.
The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata.
Dataset Summary
Language: Turkish (tr)
Source: Turkish Wikipedia
Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.clean_solar_panel_clean_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/dhirajdg/clean_solar_panel_clean_dataset.clean_solar_panel_no_variations_clean_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/dhirajdg/clean_solar_panel_no_variations_clean_dataset.new_fold_towel_2_og_dataset_20260919_183952_cleanpick_and_clean_end_to_end_wrist_cam_blocked_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/dhirajdg/pick_and_clean_end_to_end_wrist_cam_blocked_dataset.clean_sorting_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 32,
"total_frames": 13037,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/trossenai/clean_sorting_data.clean_surface_dataset_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 751,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankits0052/clean_surface_dataset_1.new_fold_towel_1_og_dataset_20260919_175141_cleanmg_fr_v1_dataset_cleanopenarm_vr_dataset_v7_clean_fixed
OpenArm VR Dataset (v7 clean fixed)
This dataset was recorded using the OpenArm VR teleoperation system.
Contains isolated actuator movements for testing.
Total episodes: 28
Total frames: 16749
dataset-push-cleanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/vasco281204/dataset-push-clean.drug_dataset_cleanrobomind_benchmark1_0_release_tienkung_gello_1rgb_clean_table_2_241211
benchmark1_0_release_tienkung_gello_1rgb_clean_table_2_241211
This dataset converts the Robomain format uniformly into LeRobot V3.0.
Dataset Statistics
本体: tienkung_gello_1rgb
末端执行器: 夹爪
任务平台显示版: 清洁餐桌part_1
total_episodes: 38
total_tasks: 1
size: 1.0G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── images
│ └── observation.images.camera_top
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/robomind_benchmark1_0_release_tienkung_gello_1rgb_clean_table_2_241211.clean_surface_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 12549,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankits0052/clean_surface_dataset.Ultrafeedback-llama3-8b-instruct-v0.2-on-policy-clean-4-binned-datamg_fr_v2_dataset_cleanlerobot_dataset_dc_a_cleanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 43,
"total_frames": 15488,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:43"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DCC1015/lerobot_dataset_dc_a_clean.ai-hdlcoder-dataset-cleanCleaned dataset for experiments and fast pretokenizing. The SQL query to create the dataset is the following:
SELECT
f.repo_name,
f.path,
c.copies,
c.size,
c.content,
l.license
FROM
(select f.*, row_number() over (partition by id order by path desc) as seqnum
from `bigquery-public-data.github_repos.files` AS f) f
JOIN
`bigquery-public-data.github_repos.contents` AS c
ON
f.id = c.id AND seqnum=1
JOIN
`bigquery-public-data.github_repos.licenses` AS l
ON
f.repo_name =… See the full description on the dataset page: https://huggingface.co/datasets/AWfaw/ai-hdlcoder-dataset-clean.NED2_Test_Dataset_CLEAN_15This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ned2_ros2_follower",
"total_episodes": 3,
"total_frames": 583,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/5AMBASH/NED2_Test_Dataset_CLEAN_15.Ultrafeedback-llama3-8b-instruct-v0.2-on-policy-clean-2-binned-dataUltrafeedback-llama3-8b-instruct-v0.2-on-policy-clean-8-binned-datalexi-coding-web-clean-datasets
lexi-coding-web-clean-datasets
lexi-coding-web-clean-datasets by Reallexi LLC AI Model Builder — llm.reallexi.io
Copyright (c) 2026 Reallexi LLC. All rights reserved.
A retrieval index built by Reallexi AI Model Builder: source text chunked, embedded, and stored for nearest-neighbor retrieval. This is not a causal-language-model checkpoint and cannot be loaded with AutoModelForCausalLM.
Contents
Source data… See the full description on the dataset page: https://huggingface.co/datasets/reallexi/lexi-coding-web-clean-datasets.clean_data2
