datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
record-screw-urThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "cta_ur_follower",
"total_episodes": 10,
"total_frames": 7519,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-screw-ur.d2l4asr-wiki-jaikema_youtube_asr_full_with_longSLR35_javanesekillkan
Killkan: Speech Recognition dataset for Kichwa
Killkan (Kichwa uyachkata payllatak killkak anta) is the first automatic speech recognition (ASR) dataset for the Kichwa language.
See also our paper (https://arxiv.org/abs/2404.15501).
d2l4asr-wiki-en_audiod2l4asr-wiki-enHarmfulVsEthical_redteaming_eval_v3ikema_dict_asrrecord-test-labThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "cta_ur_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-test-lab.record-test-urThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "cta_ur_follower",
"total_episodes": 1,
"total_frames": 426,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-test-ur.record-screw-ur-jointsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "cta_ur_follower",
"total_episodes": 2,
"total_frames": 918,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-screw-ur-joints.deon_train_llama2_v3record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 2473,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-test.record-unity-urThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "cta_ur_follower",
"total_episodes": 1,
"total_frames": 450,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-unity-ur.CTA-synthetic-dataset
German Instagram Political Communication 2021 - Synthetic Call to Action Dataset
Dataset Overview
This dataset consists of synthetic training data used to detect Calls to Action (CTAs) in German political Instagram content from the 2021 Federal Election. The synthetic data was generated using OpenAI's GPT-4o to augment original human-annotated examples. The dataset aims to address class imbalance issues for improved model performance in political communication studies.… See the full description on the dataset page: https://huggingface.co/datasets/chaichy/CTA-synthetic-dataset.record-screw-joints-anywhereThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "cta_ur_follower",
"total_episodes": 1,
"total_frames": 459,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-screw-joints-anywhere.util_rewardtrainerd2l4asr-wiki-en_contextrecord-test-ur-jointsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "cta_ur_follower",
"total_episodes": 1,
"total_frames": 267,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alex-cta/record-test-ur-joints.conlang_eval_dataset
Evaluation data for IASC
This dataset contains the evaluation data used in the research paper "Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs".
License
This dataset is available under the CC BY-SA 4.0 license. The dataset is gated to minimize evaluation data contamination.
Citation
@misc{taguchi2026creatingconlangsprobemetalinguistic,
title={Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs}… See the full description on the dataset page: https://huggingface.co/datasets/ctaguchi/conlang_eval_dataset.papi_asr_testikema_dictionary_examples_datasetCTAButil_train_llama2_v3deon_eval_llama2_v3virt_train_llama2_v3papi_asrboc
Dataset Card for "boc"
More Information needed
HHH_redteaming_eval_v3
