datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Anime_subtitles_CN
Dataset Card for Dataset Name
This repo contains a csv file about anime subtitles crawl from open web.This dataset could be used for t2t,all the NLP projects expectionly of the anime domain.It's part.1,probably will have part.2.
Dataset Description
anime_subtitles.csv: Contains two features('name' and 'caption') and 4055 rows,about 400MB. Each name represent one season or movie, caption contaions all the dialogues that the characters speaks but no characters name or… See the full description on the dataset page: https://huggingface.co/datasets/cilyy/Anime_subtitles_CN.Zouhir-Amazigh-Subtitles
Dataset Card for Zouhir-Amazigh-Subtitles
This dataset provides parallel sentence-level translations and precise audio timestamps extracted from the YouTube channel of Zouhir Amazigh. It is curated to support Automatic Speech Recognition (ASR), machine translation, and text generation tasks for the Amazigh language.
Dataset Details
Dataset Summary
The dataset contains aligned speech segments, text timestamps, and sentence-level parallel data scraped… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Zouhir-Amazigh-Subtitles.
