datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
4c_command_change_related_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch_follower",
"total_episodes": 61,
"total_frames": 17051,
"total_tasks": 1,
"total_videos": 122,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:61"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ethanCSL/4c_command_change_related_test.code-commande-publique
Code de la commande publique, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-commande-publique.speech_commands_enrichment_only
Dataset Card for SpeechCommands
Dataset Summary
📊 Data-centric AI principles have become increasingly important for real-world use cases.At Renumics we believe that classical benchmark datasets and competitions should be extended to reflect this development.
🔍 This is why we are publishing benchmark datasets with application-specific enrichments (e.g. embeddings, baseline results, uncertainties, label error scores). We hope this helps the ML community in the… See the full description on the dataset page: https://huggingface.co/datasets/renumics/speech_commands_enrichment_only.1s-crypto-data
1-Second Crypto OHLCV Data (Binance)
Historical 1-second kline (OHLCV) data for 6 major cryptocurrencies
downloaded from Binance Vision.
Updated daily — never more than 24h behind.
Assets
Symbol
Start
Updated
BTCUSDT
2019-01
daily
ETHUSDT
2019-01
daily
BNBUSDT
2019-01
daily
XRPUSDT
2019-01
daily
DOGEUSDT
2019-07
daily
SOLUSDT
2020-10
daily
File Structure
data/{SYMBOL}_1s.parquet ← full history, one file per asset… See the full description on the dataset page: https://huggingface.co/datasets/commanderzee/1s-crypto-data.details_CohereForAI__c4ai-command-r-v01
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r-v01
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r-v01.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r-v01.details_CohereForAI__c4ai-command-r-plus
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r-plus
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r-plus.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r-plus.details_CohereLabs__c4ai-command-r-plus-08-2024_private
Dataset Card for Evaluation run of CohereLabs/c4ai-command-r-plus-08-2024
Dataset automatically created during the evaluation run of model CohereLabs/c4ai-command-r-plus-08-2024.
The dataset is composed of 3 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/sasha/details_CohereLabs__c4ai-command-r-plus-08-2024_private.cobe-firas-commanded-instrument-gains
COBE/FIRAS Actual Values of Commanded Instrument Gains
This table contains the actual commanded gains for the four COBE/FIRAS
bolometer channels. LAMBDA states that the values are used to normalize
interferograms and publishes the columns in the order RH, RL, LH, and
LL: right-high, right-low, left-high, and left-low.
Instrument
COBE Far Infrared Absolute Spectrophotometer (FIRAS)
Product
Actual Values of Commanded Instrument Gains
Structure
8 settings × 4… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cobe-firas-commanded-instrument-gains.speech_commands_mel_preprocessedfr3-leader-command-pilot
FR3 Leader-Command Pilot
Four FR3 + GELLO teleoperation episodes recorded with the corrected action convention:
action is the GELLO leader's commanded joint position, following ACT/ALOHA. This is a
pilot run verifying the fix before full-scale collection — not a training set.
Contents
Robot / teleop
Franka Research 3 + GELLO leader arm
Task
pick up the skyblue cup and place it on the yellow bowl
Episodes / frames
4 / 1092 (319, 208, 347, 218)… See the full description on the dataset page: https://huggingface.co/datasets/knu-physical-ai/fr3-leader-command-pilot.eval_act_wiping_100k_command_stiffThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "custom_manipulator",
"total_episodes": 10,
"total_frames": 5319,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucazanett/eval_act_wiping_100k_command_stiff.eval_act_wiping_100k_no_command_stiffThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "custom_manipulator",
"total_episodes": 10,
"total_frames": 7040,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lucazanett/eval_act_wiping_100k_no_command_stiff.fineweb-principle_A_c_command-100M-99.99-0.01details_CohereForAI__c4ai-command-r-08-2024
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r-08-2024
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r-08-2024.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r-08-2024.gesture-commandsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/satyadevineni/gesture-commands.fineweb-principle_A_c_command-100M-99.9-0.1stackoverflow-commandline-inst
Dataset Card for "stackoverflow-commandline-inst"
More Information needed
fineweb-principle_A_c_command-100M-98-2speech_commands_m5_16k_3000msCohereForAI__c4ai-command-r-plus-08-2024fineweb-principle_A_c_command-100M-99.8-0.2speech_commands-ast-finetuned-results
Dataset Card for "speech_commands-ast-finetuned-results"
More Information needed
speech_commands_ru
Набор данных русских речевых команд
Описание
Датасет содержит аудиозаписи русских речевых команд для задач классификации аудио.
Команды
стоп, пауза, включи
Структура данных
speech_commands_dataset/
├── train/
│ ├── включи/
│ ├── стоп/
│ └── пауза/
└── test/
├── включи/
└── стоп/
MIT License
command_r_plus_classifications_multi_humanspeech-commands_thinnable-bifsmndetails_CohereForAI__c4ai-command-r-plus-08-2024
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r-plus-08-2024
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r-plus-08-2024.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r-plus-08-2024.hue_commands_synth_5k_v7mnesh-shell-commandsfluent_speech_commands_femalespeech_commands_ru
