NilanE/ParallelFiction-Ja_En-100k
Dataset details: Each entry in this dataset is a sentence-aligned Japanese web novel chapter and English fan translation. The intended use-case is for document translation tasks. Dataset format: { 'src': 'JAPANESE WEB NOVEL CHAPTER', 'trg': 'CORRESPONDING ENGLISH TRANSLATION', 'meta': { 'general': { 'series_title_eng': 'ENGLISH SERIES TITLE', 'series_title_jap': 'JAPANESE SERIES TITLE'… See the full description on the dataset page: https://huggingface.co/datasets/NilanE/ParallelFiction-Ja_En-100k.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload dataset-Ja_En-Massive-v2.jsonl
Delete dataset-Ja_En-Massive.jsonl
cleaned partial chapter titles, may leave artifacts like random commas
removed some entries where trg was a duplicate of src
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Huge dataset of (hopefully) clean Japanese web novels and English Translations
initial commit
