dzur658/opus-es-monologues
OPUS Spanish Monologues A dataset that captures monologues from the Spanish Open Subtitles Project dump and undergoes light cleaning. Monologues retained in this dataset are intances in the raw .txt dump where a single speaker is uninterrupted for more than 100 words. The dataset consists of monologues from the 2013, 2016, and 2018 OPUS Spanish monolingual datasets. Quick Dataset Facts Contains 1,481 documents Each document averages ~241.2 words The dataset… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/opus-es-monologues.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face