datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-subtitles-bitext-miningopen-subtitles-256s-bitext-miningopen-subtitles-500-bitext-miningopen-subtitles-250-bitext-miningyoutube_subtitlesopen_subtitles_en_nl
Dataset Card for OpenSubtitles
Dataset Summary
This dataset is a subset from the en-nl open_subtitles dataset.
It contains only subtitles of tv shows that have a rating of at least 8.0 with at least 1000 votes.
The subtitles are also ordered and appended into buffers several lengths, with a maximum of 370 tokens
as tokenized by the 'yhavinga/ul2-base-dutch' tokenizer.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The languages… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/open_subtitles_en_nl.Anohana-SubtitlesHere is a structured dataset that brings together, in order of dialogue, all the content of the episodes of the Anohana series.
polish-subtitlessubtitle-summary-dataset
Subtitle Summary & Keyword Dataset (YouTube)
국내 방송사 유튜브 자막 기반 스트리밍(누적) 요약 + 유튜브 검색 키워드 데이터셋.
방송을 5분 단위로 진행하며, 매 5분마다 (직전 누적 요약 + 새 5분 자막)으로 방송 대표 한 문장 요약을 갱신하고 검색 키워드를 생성합니다.
교사 모델 Qwen3-235B-A22B-Instruct-2507-FP8 (vLLM, 자막-only). HBKenerzai/LGUplus_summary_keyword와 동일 스키마.
규모
split
회차
레코드(5분)
train
2,641
10,159
validation
294
971
합계
2,935
11,130
약 563시간 · 110개 채널 · 요약 중앙값 58자 · 도메인(유튜브 카테고리): Entertainment 2982 · Travel & Events 2081 · News &… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-dataset.subtitle-summary-sft
Subtitle Summary & Keyword — SFT (YouTube)
jungsanghyun/subtitle-summary-dataset를 SFT 학습용 messages 형식으로 가공.
형식 (system 없음, 단일 턴)
입력(user): [이전 요약] {last_summary}\n[자막] {5min_script}
출력(assistant): [요약] {한 문장}\n[키워드] {검색어} — 라벨 2줄(소형 모델 친화).
split
examples
train
10,159
validation
971
chat_template.jinja({% generation %})로 assistant만 loss. 파싱 \[요약\]\s*(.+) / \[키워드\]\s*(.+).
라이선스
cc-by-nc-4.0, 연구·비상업.
subtitle-summary-lgu
Broadcast Subtitle Summary & Keyword — (LGUplus source, our pipeline)
HBKenerzai/LGUplus_summary_keyword의 방송 자막을 우리 형식으로 재구성한 뒤, **우리 파이프라인
(Qwen3-235B-A22B-Instruct-2507-FP8, vLLM, 자막-only)**으로 5분 단위 누적 요약 + 검색 키워드를
다시 생성한 데이터셋입니다. 원천 자막·메타 저작권은 원 출처(AI Hub / 방송사)에 있습니다.
규모
split
회차
레코드(5분)
train
16,609
118,608
validation
1,846
13,177
합계
18,455
131,785
약 9,025시간 분량 · 요약 길이 중앙값 63자 · 도메인: 생활정보 9528 · 정보/토크 9458 · 퀴즈/게임 8123 · 푸드/요리 7221 · 토론/대담… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu.subtitle-summary-testset
Subtitle Summary & Keyword — Test set (YouTube)
방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다.
jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다.
규모
split
회차
레코드(5분)
test
114
519
약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28
스키마
program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.archer_subtitlessubtitle-summary-lgu-sft
Broadcast Subtitle Summary & Keyword — SFT (LGUplus source)
jungsanghyun/subtitle-summary-lgu를 SFT 학습용 messages 형식으로 가공한 버전입니다.
형식 (system 없음, 단일 턴)
입력(user): [이전 요약] {last_summary}\n[자막] {5min_script}
출력(assistant): [요약] {한 문장}\n[키워드] {검색어} — 라벨 2줄(소형 모델 친화·파싱 용이).
split
examples
train
118,608
validation
13,177
generation만 학습
chat_template.jinja({% generation %} 블록)로 assistant 응답만 loss. TRL SFTConfig(assistant_only_loss=True… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-lgu-sft.subtitle-summary-testset-sft
Subtitle Summary & Keyword — Test set SFT (YouTube)
jungsanghyun/subtitle-summary-testset를 messages 형식으로 가공한 평가용 버전.
형식 (system 없음, 단일 턴)
입력(user): [이전 요약] {last_summary}\n[자막] {5min_script}
출력(assistant): [요약] {한 문장}\n[키워드] {검색어}
split
examples
test
519
파싱 \[요약\]\s*(.+) / \[키워드\]\s*(.+). chat_template.jinja 동봉.
라이선스
cc-by-nc-4.0, 연구·비상업.
SubtitlesSubtitles-splitrp_dataset_Tr_from_subtitles
