CoolFace
Datasetpublic

TurkuNLP/WebDocumentDescriptors

Data release for the paper Task-Agnostic Web Document Annotation with LLM-Generated Descriptors (forthcoming). The descriptors are generated via a task-agnostic data annotation pipeline described in the paper (link coming soon). This Hugging Face dataset repository contains 5 distinct datasets: a descriptor-annotated version of a 10 billion token (~15 million document) sample of FineWeb. The 800k label descriptor schema The 500k document sample of FineWeb used to develop the schema along… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/WebDocumentDescriptors.

sourceHugging Faceodc-byupdated 11d agoView on Hugging Face
1likes266downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
TurkuNLP/WebDocumentDescriptors · CoolFace