scotus-sim/scotus-ketanji_brown_jackson-audio
SCOTUS-sim audio: ketanji_brown_jackson Per-utterance audio clips from Oyez oral-argument mp3s, sliced at the start_time / stop_time timestamps stored in the companion scotus-sim/scotus-ketanji_brown_jackson-training dataset. Alignment clip_NNNNN.wav in the tarball corresponds exactly to audio_segments.jsonl[NNNNN] in the training companion dataset. In metadata.jsonl each row carries the same 0-padded index in idx. This supersedes the v1 tarball, which had… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-ketanji_brown_jackson-audio.
SCOTUS-sim audio: ketanjibrownjackson
Per-utterance audio clips from Oyez oral-argument mp3s, sliced at the start_time / stop_time timestamps stored in the companion scotus-sim/scotus-ketanji_brown_jackson-training dataset.
Alignment
clip_NNNNN.wav in the tarball corresponds exactly to audio_segments.jsonl[NNNNN] in the training companion dataset. In metadata.jsonl each row carries the same 0-padded index in idx.
This supersedes the v1 tarball, which had systematic audio↔segment index misalignment (Apr 2026).
Stats
- clips: 4024
- total duration: 7.04 hours
- unique cases: 213
Files
ketanji_brown_jackson.tar.gz— all clips; members namedketanji_brown_jackson_NNNNN.wav, 24 kHz mono PCM 16-bit.metadata.jsonl— one row per clip withtext,speaker,duration,case_id,audio_url,idx.
License
Audio is derived from Oyez.org (CC-BY-NC 4.0). Derivative TTS training artifacts inherit the NC term.
