CoolFace
Datasetpublic

gagan3012/ppc

Political Parliamentary Corpus (PPC) A multilingual corpus of parliamentary speech, party manifestos and (for German) historical newspapers, exposed with one config per language. Every record follows a single unified schema, so the languages are directly comparable. 44,978,179 documents (~16.5B tokens, chars/4 estimate) 5 languages: de, en, it, pl, tr Coverage 1803–2026 22 sources, unified schema Languages / configs Config Language Documents ~Tokens Years… See the full description on the dataset page: https://huggingface.co/datasets/gagan3012/ppc.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes159downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
gagan3012/ppc · CoolFace