CoolFace
Datasetpublic

agentlans/en-document-format-classification

English Document Format Classification Dataset English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats. Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks. Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage). Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.

sourceHugging Faceodc-byupdated 23d agoView on Hugging Face
0likes227downloads
StructuredData.jsonl.zst4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:6f9b0632e048ec11adbd73d2c6fdaa119daf5b010de292954e8d0ed7483f3cce3size 6631424