CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Fascinat0r /taiga_corpus_subtitles Dataset Card for Taiga Corpus - TV Series Subtitles This dataset contains subtitles extracted from various TV series. The original data is sourced from the Taiga Corpus. It consists of line-level subtitle information with precise timing and additional metadata including series title and episode information. The dataset is designed for tasks such as subtitle alignment, translation, and dialogue analysis. Dataset Details Each record in the dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/Fascinat0r/taiga_corpus_subtitles.tabularquestion-answering10M<n<100M0 likes93 downloads2y agoHugging Face02diyarhamedi /HowTo100M-subtitles-small HowTo100M-subtitles-small The subtitles from a subset of the HowTo100M dataset. tabular10K<n<100K2 likes71 downloads3y agoHugging Face03IljaSamoilov /ERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR. from datasets import load_dataset, load_metric dataset = load_dataset('csv', data_files={'train': "train.tsv", \ "validation":"val.tsv", \ "test": "test.tsv"}, delimiter='\t') tabular100K<n<1M0 likes68 downloads4y agoHugging Face04PJMixers-Dev /Subtitles-rag-answers-r1 Subtitles-rag-answers-r1 You should mask everything except the last turn. The only part that matters to teach the model is the last turn, as you are teaching it to always output thinking, no matter what the user feeds it. It's setup to be trained like R1: tabular1K<n<10K0 likes43 downloads1y agoHugging Face05PJMixers-Dev /Subtitles-rag-questions-r1 Subtitles-rag-questions-r1 You should mask everything except the last turn. The only part that matters to teach the model is the last turn, as you are teaching it to always output thinking, no matter what the user feeds it. It's setup to be trained like R1: tabularn<1K0 likes42 downloads1y agoHugging Face06yhavinga /open_subtitles_en_nl Dataset Card for OpenSubtitles Dataset Summary This dataset is a subset from the en-nl open_subtitles dataset. It contains only subtitles of tv shows that have a rating of at least 8.0 with at least 1000 votes. The subtitles are also ordered and appended into buffers several lengths, with a maximum of 370 tokens as tokenized by the 'yhavinga/ul2-base-dutch' tokenizer. Supported Tasks and Leaderboards [More Information Needed] Languages The languages… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/open_subtitles_en_nl.tabulartranslation1M<n<10M2 likes41 downloads4y agoHugging Face07jungsanghyun /subtitle-summary-datasetgated Subtitle Summary & Keyword Dataset (YouTube) 국내 방송사 유튜브 자막 기반 스트리밍(누적) 요약 + 유튜브 검색 키워드 데이터셋. 방송을 5분 단위로 진행하며, 매 5분마다 (직전 누적 요약 + 새 5분 자막)으로 방송 대표 한 문장 요약을 갱신하고 검색 키워드를 생성합니다. 교사 모델 Qwen3-235B-A22B-Instruct-2507-FP8 (vLLM, 자막-only). HBKenerzai/LGUplus_summary_keyword와 동일 스키마. 규모 split 회차 레코드(5분) train 2,641 10,159 validation 294 971 합계 2,935 11,130 약 563시간 · 110개 채널 · 요약 중앙값 58자 · 도메인(유튜브 카테고리): Entertainment 2982 · Travel & Events 2081 · News &… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-dataset.tabularsummarization10K<n<100K0 likes21 downloads2mo agoHugging Face08hipml /survivor-subtitles-cleaned Survivor Subtitles Dataset (cleaned) Dataset Description A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts. This dataset is a modification of the original Survivor Subtitles dataset after cleaning up and joining subtitle fragments. This dataset is a work in progress and any contributions are welcome. Source The subtitles were… See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles-cleaned.tabular100K<n<1M0 likes14 downloads2y agoHugging Face09jungsanghyun /subtitle-summary-lgugated Broadcast Subtitle Summary & Keyword — (LGUplus source, our pipeline) HBKenerzai/LGUplus_summary_keyword의 방송 자막을 우리 형식으로 재구성한 뒤, **우리 파이프라인 (Qwen3-235B-A22B-Instruct-2507-FP8, vLLM, 자막-only)**으로 5분 단위 누적 요약 + 검색 키워드를 다시 생성한 데이터셋입니다. 원천 자막·메타 저작권은 원 출처(AI Hub / 방송사)에 있습니다. 규모 split 회차 레코드(5분) train 16,609 118,608 validation 1,846 13,177 합계 18,455 131,785 약 9,025시간 분량 · 요약 길이 중앙값 63자 · 도메인: 생활정보 9528 · 정보/토크 9458 · 퀴즈/게임 8123 · 푸드/요리 7221 · 토론/대담… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu.tabularsummarization100K<n<1M0 likes12 downloads2mo agoHugging Face10Zarakun /youtube_ua_noisy_subtitles_test The list of all subsets in the dataset Each subset is generated splitting videos from given particular ukrainiam YouTube channel All subsets are in test split "opodcast" subset is from channel "О! ПОДКАСТ" "rozdympodcast" subset is from channel "Роздум | Подкаст" "test" subset is just a small subset of samples Loading a particular subset >>> data_files = {"train": "data/<your_subset>.parquet"} >>> data = load_dataset("Zarakun/youtube_ua_subtitles_test"… See the full description on the dataset page: https://huggingface.co/datasets/Zarakun/youtube_ua_noisy_subtitles_test.tabularautomatic-speech-recognition1K<n<10K0 likes10 downloads3y agoHugging Face11Zarakun /youtube_ua_subtitles_test The list of all subsets in the dataset Each subset is generated splitting videos from given particular ukrainiam YouTube channel All subsets are in test split "opodcast" subset is from channel "О! ПОДКАСТ" "rozdympodcast" subset is from channel "Роздум | Подкаст" "test" subset is just a small subset of samples Loading a particular subset >>> data_files = {"train": "data/<your_subset>.parquet"} >>> data = load_dataset("Zarakun/youtube_ua_subtitles_test"… See the full description on the dataset page: https://huggingface.co/datasets/Zarakun/youtube_ua_subtitles_test.tabularautomatic-speech-recognition1K<n<10K0 likes9 downloads3y agoHugging Face12hipml /survivor-subtitles Survivor Subtitles Dataset Dataset Description A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts. Source The subtitles were obtained from OpenSubtitles.com. Dataset Details Coverage: Seasons: 1-47 Episodes per season: ~13-14 Total episodes: ~600 Format: Text files containing timestamped subtitle data Character… See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles.tabular100K<n<1M1 likes9 downloads2y agoHugging Face13hossam87 /el-mal-el-halal-podcast-subtitles El Mal El Halal Podcast Subtitles Dataset Summary El Mal El Halal Podcast Subtitles is a collection of manual subtitles for 18 episodes of the El Mal El Halal podcast by Eng. Mohamed Aboulnaga, covering Arabic content. This dataset is designed for research on speech processing, translation, semantic search, and Arabic NLP. Total episodes: 18 - untill the date of 03/08/2025 Total segments: 13 970 Total words: 166 505 Total duration: 20 h 50 m 56 s (75 057 s) Average… See the full description on the dataset page: https://huggingface.co/datasets/hossam87/el-mal-el-halal-podcast-subtitles.tabularautomatic-speech-recognition10K<n<100K0 likes7 downloads1y agoHugging Face14jungsanghyun /subtitle-summary-testsetgated Subtitle Summary & Keyword — Test set (YouTube) 방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다. jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다. 규모 split 회차 레코드(5분) test 114 519 약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28 스키마 program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.tabularsummarizationn<1K0 likes7 downloads2mo agoHugging Face15PJMixers-Dev /Subtitles-rag-questions-r1-splittabularn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.