TurkuNLP/WebDocumentDescriptors
Data release for the paper Task-Agnostic Web Document Annotation with LLM-Generated Descriptors (forthcoming). The descriptors are generated via a task-agnostic data annotation pipeline described in the paper (link coming soon). This Hugging Face dataset repository contains 5 distinct datasets: a descriptor-annotated version of a 10 billion token (~15 million document) sample of FineWeb. The 800k label descriptor schema The 500k document sample of FineWeb used to develop the schema along… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/WebDocumentDescriptors.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face