spadeMIA/GovReport_Corpus_1024_2040
GovReport fine-tuning corpus Overview This dataset contains GovReport text sequences prepared for controlled language-model fine-tuning and evaluation. It provides three predefined splits with a binary annotation column for reproducible dataset handling. The label column is dataset metadata, not a language-model training target. Project and contributors This corpus was prepared at the SPADE Lab, Koç University. Contributor Affiliation… See the full description on the dataset page: https://huggingface.co/datasets/spadeMIA/GovReport_Corpus_1024_2040.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face