datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vtg-featurestwin_test_new_features5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 2,
"total_frames": 406,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/twin_test_new_features5.twin_test_new_features_degreesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 3,
"total_frames": 646,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/twin_test_new_features_degrees.twin_test_new_features_radsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 2,
"total_frames": 465,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/twin_test_new_features_rads.new_features_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "stringman",
"total_episodes": 1,
"total_frames": 155,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 60,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/naavox/new_features_test.new_features_test_30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "stringman",
"total_episodes": 3,
"total_frames": 924,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/naavox/new_features_test_30.microvent-features
microvent-features
Derived signals for the microvent core release: per-keyframe OCR text,
per-chunk ASR transcripts, and an embedding zoo (keyframe-level vision,
keyframe-OCR text, audio-level, video-level, omni-modal).
This card covers only the features. For the source videos, audio,
keyframes, and the public eval annotations, see the microvent dataset
card. All artifacts here key on the same chunk_id and follow the same
WebDataset shard layout, so joining feature shards back… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/microvent-features.pjepa_features
P-JEPA feature archives
Precomputed input features and aligned annotations for the
P-JEPA encoders. These are backbone/pooler
outputs, before P-JEPA temporal encoding. No raw videos or P-JEPA output caches
are included. All feature tensors are float32.
The files use the directory layout expected by the
P-JEPA code. That GitHub repository is currently
private pending the code release. The plain PyTorch example below works without
it; the full training/evaluation recipes require… See the full description on the dataset page: https://huggingface.co/datasets/FelTris/pjepa_features.
