CoolFace
Datasetpublic

hf-test/sv_corpora_parliament_processed

Swedish text corpus created by extracting the "text" from dataset = load_dataset("europarl_bilingual", lang1="en", lang2="sv", split="train") and processing it with: import re def extract_text(batch): text = batch["translation"]["sv"] batch["text"] = re.sub(chars_to_ignore_regex, "", text.lower()) return batch

sourceHugging Faceupdated 5y agoView on Hugging Face
1likes50downloads

hf-test/sv_corpora_parliament_processed · main · files are served by the source, never re-hosted here