datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cucumber-place-classifier-eval071526-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "vibeboard_follower_tilt",
"total_episodes": 74,
"total_frames": 2908,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-eval071526-v1-trim.cucumber-place-classifier-filtered071126This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-filtered071126.classifier_source
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.arc-agi-impabs-dpolr1e-7-beta0.01-classifiersft5e-7ham10000-skin-lesion-classifier-datagaia-dr3-vari-classifier-definition
Gaia DR3 variability classifier definition
This reference table describes the classifier used for Gaia DR3 variability classification. ESA's DR3 release contains one row, for the nTransits:5+ classifier. The related vari_classifier_class_definition table describes its class vocabulary, while vari_classifier_result contains its per-source results.
Use
python -m venv .venv && .venv/bin/pip install datasets pyarrow
from datasets import load_dataset
classifier =… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-vari-classifier-definition.eval_Aibot2_control_poses_classifier_20260817-094847This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aibot2",
"total_episodes": 1,
"total_frames": 221,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/eval_Aibot2_control_poses_classifier_20260817-094847.toxicity-multi-label-classifier
Part of a course titled "Generative AI application design & development"
https://genai.acloudfan.com/
Created from a dataset available on Kaggle.
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data
gaia-dr3-vari-classifier-class-definition
Gaia DR3 variability classifier class definitions
This reference table describes the published variability classes used by the Gaia DR3 classifier recorded in vari_classifier_definition and applied in vari_classifier_result. ESA's DR3 table contains the classes for the nTransits:5+ classifier.
Use
python -m venv .venv && .venv/bin/pip install datasets pyarrow
from datasets import load_dataset
classes =… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-vari-classifier-class-definition.ML-Music-Classifier-dataset-and-model-name-Models
🎧 Spotify Music Preference Analysis
🧠 Project Overview
This project analyzes Spotify music data to predict song preferences using machine learning models. The analysis is based on a dataset of 195 songs (100 liked, 95 disliked) with various audio features extracted from Spotify's API.
📂 Dataset Description
📥 Data Collection Process
Liked Songs (100 tracks):
🎵 Primarily French Rap
🎸 Some American Rap, Rock, and Electronic music
✅… See the full description on the dataset page: https://huggingface.co/datasets/Jack1808/ML-Music-Classifier-dataset-and-model-name-Models.classifierThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 20,
"total_frames": 14656,
"total_tasks": 1,
"total_videos": 80,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jannick-st/classifier.push-cube-classifier_cropped_resizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 17,
"total_frames": 2807,
"total_tasks": 1,
"total_videos": 51,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jannick-st/push-cube-classifier_cropped_resized.pick_up_classifier_33This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 15,
"total_frames": 381,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aiwhisperer/pick_up_classifier_33.hil-serl-push-circle-classifierThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 20,
"total_frames": 6747,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/helper2424/hil-serl-push-circle-classifier.alignment-classifier-training-chunked-unlabeledfm_queries_classifier
Dataset Card for "fm_queries_classifier"
More Information needed
pick_up_classifier_32This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 15,
"total_frames": 315,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aiwhisperer/pick_up_classifier_32.intent-classifier-v4tulu-v2-sft-mixture-first-stage-classifier-outputsgithub_fetch_huggingface_terminal_9134_x9v2m6_derived_emotion_classifier
Emotion Classifier Data
A derived dataset used to train an emotion-classification model.
Overview
Dataset ID: DRV-EMOTION
Catalog: CUSTOMER-FEEDBACK-ANALYTICS
Origin: Derived from SRC-ALPHA and SRC-BETA with manual annotation
Records: 8,700
Product Line: Customer Feedback Analytics
Source
This dataset originates from: Derived from SRC-ALPHA and SRC-BETA with manual annotation.
Contents
Cleaned and re-labeled samples for emotion… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/github_fetch_huggingface_terminal_9134_x9v2m6_derived_emotion_classifier.book-text-classifier
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shhossain/book-text-classifier.VV-classifier-2.0-payment-runs-v1.2ignorance-classifier-testing-datafm_classifier_mutable-1-n
Dataset Card for "fm_classifier_mutable-1-n"
More Information needed
hil-serl-pusht-reward-classifier-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 2,
"total_frames": 529,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/helper2424/hil-serl-pusht-reward-classifier-test.mutability_classifier-1-n
Dataset Card for "mutability_classifier-1-n"
More Information needed
samyx-classifier-train-v3VV-classifier-2.0-payment-runs_v1.1tulu-v2-sft-mixture-first-stage-classifier-outputs-v1classifier_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 4,
"total_frames": 2742,
"total_tasks": 1,
"total_videos": 12,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jannick-st/classifier_2.
