TurkuNLP/WebDocumentDescriptors
Data release for the paper Task-Agnostic Web Document Annotation with LLM-Generated Descriptors (forthcoming). The descriptors are generated via a task-agnostic data annotation pipeline described in the paper (link coming soon). This Hugging Face dataset repository contains 5 distinct datasets: a descriptor-annotated version of a 10 billion token (~15 million document) sample of FineWeb. The 800k label descriptor schema The 500k document sample of FineWeb used to develop the schema along… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/WebDocumentDescriptors.
This repository belongs to TurkuNLP on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
