bertin-project/BOE-XSUM
BOE-XSUM Balanced Dataset - Reviewed and Cleaned Description The BOE 2025 Dataset is a collection of BOE articles with extreme summaries of them. This dataset has been carefully balanced and cleaned to ensure its quality and usefulness in natural language processing (NLP) tasks, primarily for evaluating generative models. Read more in https://arxiv.org/abs/2509.24908 Dataset Content The dataset is composed of the following subsets (splits): train:… See the full description on the dataset page: https://huggingface.co/datasets/bertin-project/BOE-XSUM.
BOE-XSUM Balanced Dataset - Reviewed and Cleaned
Description
The BOE 2025 Dataset is a collection of BOE articles with extreme summaries of them. This dataset has been carefully balanced and cleaned to ensure its quality and usefulness in natural language processing (NLP) tasks, primarily for evaluating generative models.
Read more in https://arxiv.org/abs/2509.24908
Dataset Content
The dataset is composed of the following subsets (splits):
train: Training data.validation: Validation data.test: Test data.
Dataset Columns
The dataset contains the following fields:
- id: Unique identifier of the item.
- boe_materials: A category identifier used by the BOE.
- boe_date_publication: Publication date of the BOE article.
- boe_previous: Previous BOE articles that are affected by this new BOE.
- boe_id: Identifier of the BOE.
- boe_title: Title of the BOE article.
- boe_soup_xml: Fully scraped web page.
- tweet_original: Original tweet by Eva Belmonte.
- boe_category: Category to which this item belongs.
- boe_alert: BOE classification codes in state areas.
- boe_departament: Department from which the BOE article originates.
- tweet_text_cleaned: Extreme summary generated from a meticulous review of Eva Belmonte's tweet.
- boe_subsequent: Laws that are modified by this order (only for articles referring to legislation).
