CoolFace
Datasetpublic

ParaPat/para_pat

ParaPat: The Multi-Million Sentences Parallel Corpus of Patents Abstracts This dataset contains the developed parallel corpus from the open access Google Patents dataset in 74 language pairs, comprising more than 68 million sentences and 800 million tokens. Sentences were automatically aligned using the Hunalign algorithm for the largest 22 language pairs, while the others were abstract (i.e. paragraph) aligned.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
16likes146downloads

ParaPat/para_pat · main · files are served by the source, never re-hosted here