CoolFace
Datasetpublic

scotus-sim/scotus-jeffrey_l_fisher-audio

SCOTUS-sim audio: jeffrey_l_fisher Per-utterance audio clips from Oyez oral-argument mp3s, sliced at the start_time / stop_time timestamps stored in the companion scotus-sim/scotus-jeffrey_l_fisher-training dataset. Alignment clip_NNNNN.wav in the tarball corresponds exactly to audio_segments.jsonl[NNNNN] in the training companion dataset. In metadata.jsonl each row carries the same 0-padded index in idx. This supersedes the v1 tarball, which had systematic… See the full description on the dataset page: https://huggingface.co/datasets/scotus-sim/scotus-jeffrey_l_fisher-audio.

sourceHugging Facecc-by-nc-4.0updated 5mo agoView on Hugging Face
0likes9downloads
Dataset Card

SCOTUS-sim audio: jeffreylfisher

Per-utterance audio clips from Oyez oral-argument mp3s, sliced at the start_time / stop_time timestamps stored in the companion scotus-sim/scotus-jeffrey_l_fisher-training dataset.

Alignment

clip_NNNNN.wav in the tarball corresponds exactly to audio_segments.jsonl[NNNNN] in the training companion dataset. In metadata.jsonl each row carries the same 0-padded index in idx.

This supersedes the v1 tarball, which had systematic audio↔segment index misalignment (Apr 2026).

Stats

  • —clips: 1325
  • —total duration: 2.25 hours
  • —unique cases: 47

Files

  • —jeffrey_l_fisher.tar.gz — all clips; members named jeffrey_l_fisher_NNNNN.wav, 24 kHz mono PCM 16-bit.
  • —metadata.jsonl — one row per clip with text, speaker, duration, case_id, audio_url, idx.

License

Audio is derived from Oyez.org (CC-BY-NC 4.0). Derivative TTS training artifacts inherit the NC term.