scotus-sim/scotus-elizabeth_b_prelogar-audio
SCOTUS-sim audio: elizabeth_b_prelogar Per-utterance audio clips from Oyez oral-argument mp3s, sliced at the start_time / stop_time timestamps stored in the companion scotus-sim/scotus-elizabeth_b_prelogar-training dataset. Alignment clip_NNNNN.wav in the tarball corresponds exactly to audio_segments.jsonl[NNNNN] in the training companion dataset. In metadata.jsonl each row carries the same 0-padded index in idx. This supersedes the v1 tarball, which had… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-elizabeth_b_prelogar-audio.
v2: rebuilt from Oyez mp3 source — 962 clips (1.67h, 29 cases), indexed by segments.jsonl ordinal so clip_NNN.wav == segments[NNN]. Replaces v1 which had systematic audio↔segment index misalignment. (README)
v2: rebuilt from Oyez mp3 source — 962 clips (1.67h, 29 cases), indexed by segments.jsonl ordinal so clip_NNN.wav == segments[NNN]. Replaces v1 which had systematic audio↔segment index misalignment. (metadata.jsonl)
v2: rebuilt from Oyez mp3 source — 962 clips (1.67h, 29 cases), indexed by segments.jsonl ordinal so clip_NNN.wav == segments[NNN]. Replaces v1 which had systematic audio↔segment index misalignment. (tarball)
Upload metadata.csv with huggingface_hub
Upload elizabeth_b_prelogar.tar.gz with huggingface_hub
initial commit
