Whispering-GPT/whisper-transcripts-the-verge
annotations_creators: machine-generated language: en language_creators: crowdsourced license: [] multilinguality: monolingual paperswithcode_id: wikitext-2 pretty_name: Whisper-Transcripts size_categories: 1M<n<10M source_datasets: original tags: [] task_categories: text-generation fill-mask task_ids: language-modeling masked-language-modeling
428
Final partition
Fixed concatenation of jsonl files
Fixed issue (#1)
Uploaded transcripts for around 3400 videos
Data for ~800 videos
Added Dataset Card
Uploaded data for 600 videos
initial commit
