agentlans/en-document-format-classification
English Document Format Classification Dataset English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats. Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks. Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage). Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.
English Document Format Classification Dataset
English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats.
- Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks.
- Configurations: Each document type is available in its own dedicated dataset configuration (e.g.,
TutorialHow-ToGuide,PersonalAboutPage). - Splits: The
Allconfiguration contains every document type combined, featuring a 10% stratified test split.
Curation & Filtering Pipeline
This dataset is derived from a subset of the first 1 million rows of `allenai/c4`, utilizing annotations from `agentlans/en-document-classification`.
- Filtering: Samples are included only when the
weborganizer_formatfield matchesdoc_type_v2_primary. - Personally identifiable information (PII) removal: Samples were removed based on `agentlans/multilingual-e5-small-pii-detector`.
Class Distribution (All Split)
Licensing
Distributed under the Open Data Commons Attribution License (ODC-BY), matching the licensing terms of upstream sources `allenai/c4` and `agentlans/en-document-classification`.
