datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oral-arguments-us
US Court Oral Arguments -- metadata and transcripts
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every recording catalogued here was collected and catalogued by CourtListener, and this dataset is
sliced… See the full description on the dataset page: https://huggingface.co/datasets/docketx/oral-arguments-us.nepal-oral-demo
Nepal Oral Demo
Public, always-safe fixtures for the Nepal oral-language backbone (AkAiNp).
Fictional “Demo Himalayan” track only
Schema examples for CI, export-script tests, and the Expo training app offline demo pack
No real community speakers, ever
Monorepo: nepal-multilingual-llm (local project). Source: data/packs/_demo/ + packages/schema/examples/.
Intended uses
OK
Not OK
Unit tests, pipeline dry-runs
Training production ASR/TTS as if it were… See the full description on the dataset page: https://huggingface.co/datasets/AkAiNp/nepal-oral-demo.
