unlabeled
so101-expanded-v3-success-plus-unlabeled
SO101 Expanded v3 — Success + Unlabeled with Filtered Training View
This public dataset combines a canonical LeRobot v3.0 SO100/SO101 single-arm release with the
fully audited joint-delta filtering sidecars used to select training observations and targets.
It is the success-plus-unlabeled view only. It does not contain the companion release's
explicitly failed or suboptimal trajectories.
Release summary
Field
Value
Canonical episodes
17,745
Canonical… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/so101-expanded-v3-success-plus-unlabeled.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.VietMed_unlabeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) unlabeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the unlabeled set: 966h - 230k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-unlabeled.py
need to do: check misspelling, restore foreign words phonetised to… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_unlabeled.magvit_1m_unlabeledchexpert_unlabeledunlabeled_data
