datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
berry-eval30-detectionso101_eval3_broad_grounding
eval3_objectvla_vl_pairs
ObjectVLA-style vision-language co-training data for the Eval 3 face/celebrity
identification task. Companion dataset to
HBOrtiz/so101_eval3_aug_v3_200celebs.
Contents
manifest.parquet — 176 670 VL pairs (one row per (frame, portrait, caption_type)).
data.tar.zst — 1.7 GB compressed archive containing:
images/chunk-{000..003}/*.jpg — 29 445 wrist-cam frames (480 × 640, 3 sampled per episode at 20% / 50% / 80%).
references/<episode>__ref.jpg — 9… See the full description on the dataset page: https://huggingface.co/datasets/HBOrtiz/so101_eval3_broad_grounding.train_reflection_eval3eval3-text-benchmark
Asset from the SCALEMED Framework
This model/dataset is an asset released as part of the SCALEMED framework, a project focused on developing scalable and resource-efficient medical AI assistants.
Project Overview
The models, known as DermatoLlama, were trained on versions of the DermaSynth dataset, which was also generated using the SCALEMED pipeline.
For a complete overview of the project, including all related models, datasets, and the source code, please visit our main… See the full description on the dataset page: https://huggingface.co/datasets/DermaVLM/eval3-text-benchmark.so101_eval3_cotrain_grounding
eval3_track3_vl_pairs (v3 — image↔label mispairing fixed)
ObjectVLA-style vision-language co-training data for the Track 3 TOY 3-celeb Eval 3 task. v3 fixes a critical bug where ~98% of rows had image_path pointing to a frame extracted from a different episode than the row's labels described.
Companion to HBOrtiz/so101_eval3_track3_v3_baseline.
What changed vs prior pushes
Bug fix (v3): Earlier builders keyed video lookups by enumeration index in their own filtered… See the full description on the dataset page: https://huggingface.co/datasets/HBOrtiz/so101_eval3_cotrain_grounding.eval-3-epoch-resultsfinepdfs-eval3eval-3-epocheval_3B_on_traineval_3B_on_train_armoEval3-ManualSplit-2025_02_06_error_supervisor_rl_training_data.1eval3_track3_vl_pairs
eval3_track3_vl_pairs (v3 — image↔label mispairing fixed)
ObjectVLA-style vision-language co-training data for the Track 3 TOY 3-celeb Eval 3 task. v3 fixes a critical bug where ~98% of rows had image_path pointing to a frame extracted from a different episode than the row's labels described.
Companion to HBOrtiz/so101_eval3_track3_v3_baseline.
What changed vs prior pushes
Bug fix (v3): Earlier builders keyed video lookups by enumeration index in their own filtered… See the full description on the dataset page: https://huggingface.co/datasets/userdarius/eval3_track3_vl_pairs.eval3_TOY_video_images
Eval3 TOY Video Images
Image-question-answer grounding dataset extracted from the Eval 3 TOY celebrity permutation videos.
Each episode contributes three frames from the first 6 seconds:
start frame
middle frame
end frame
The labels use the user-provided episode block ground truth:
episodes 0-29: Taylor Swift
episodes 30-59: Barack Obama
episodes 60-89: Yann LeCun
within each 30-episode block: first 10 left, next 10 middle, last 10 right
Files:
all.jsonl: all 270 examples… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval3_TOY_video_images.test_reflection_eval3eval3-phase2-sceneseval-3-epoch-results-newtrain_reflection_eval3_with_rewardseval_3B_base_on_train_armoeval3-celeb-faceseval3-dataset
My Undersampled Dataset
This dataset contains balanced labels by undersampling overrepresented classes.
