yhavinga/open_subtitles_en_nl
Dataset Card for OpenSubtitles Dataset Summary This dataset is a subset from the en-nl open_subtitles dataset. It contains only subtitles of tv shows that have a rating of at least 8.0 with at least 1000 votes. The subtitles are also ordered and appended into buffers several lengths, with a maximum of 370 tokens as tokenized by the 'yhavinga/ul2-base-dutch' tokenizer. Supported Tasks and Leaderboards [More Information Needed] Languages… See the full description on the dataset page: https://huggingface.co/datasets/yhavinga/open_subtitles_en_nl.
Dataset Card for OpenSubtitles
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage: http://opus.nlpl.eu/OpenSubtitles.php
- Repository: None
- Paper: http://www.lrec-conf.org/proceedings/lrec2016/pdf/62_Paper.pdf
- Leaderboard: [More Information Needed]
- Point of Contact: [More Information Needed]
Dataset Summary
This dataset is a subset from the en-nl open_subtitles dataset. It contains only subtitles of tv shows that have a rating of at least 8.0 with at least 1000 votes. The subtitles are also ordered and appended into buffers several lengths, with a maximum of 370 tokens as tokenized by the 'yhavinga/ul2-base-dutch' tokenizer.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The languages in the dataset are:
- en
- nl
Dataset Structure
Data Instances
Here are some examples of questions and facts:
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data
[More Information Needed]
Initial Data Collection and Normalization
[More Information Needed]
Who are the source language producers?
[More Information Needed]
Annotations
[More Information Needed]
Annotation process
[More Information Needed]
Who are the annotators?
[More Information Needed]
Personal and Sensitive Information
[More Information Needed]
Considerations for Using the Data
Social Impact of Dataset
[More Information Needed]
Discussion of Biases
[More Information Needed]
Other Known Limitations
[More Information Needed]
Additional Information
Dataset Curators
[More Information Needed]
Licensing Information
[More Information Needed]
Citation Information
[More Information Needed]
Contributions
Thanks to @abhishekkrthakur for adding the open_subtitles dataset.
