JQL-AI/JQL-LLM-Edu-Annotations
π JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. π Dataset Summary Multilingual document-level quality annotations scored on a 0β5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-itβ¦ See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.
π JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
π Dataset Summary
Multilingual document-level quality annotations scored on a 0β5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it, and LLaMA-3.3-70B-it. Up to 500k documents per language from FineWeb2 are included. Annotations are aligned with human ratings and intended for quality estimation, distillation, and multilingual benchmark research.
π Languages
In total we included 35 European languages. Input documents are in their native language, but models were prompted and responded in English.
π§± Dataset Structure:
βοΈ Data Splits:
This dataset is not pre-split. Users can generate custom splits by:
- Language
- Model agreement
- Prediction validity
- Document length or other features
π― Intended Use
- Training multilingual document quality models
- Benchmarking multilingual LLM performance
- Distillation and teacher-student LLM training
- Creating filters for noisy web-scale data
β οΈ Limitations:
- LLM-generated scores, not human-authored
- Some predictions may be invalid or inconsistent
- No domain control across documents
- Educational value is a subjective, task-specific metric
π Citation
@article{ali2025judging,
title = {Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models},
author = {
Mehdi Ali,
Manuel Brack,
Max LΓΌbbering,
Elias Wendt,
Abbas Goher Khan,
Richard Rutmann,
Alex Jude,
Maurice Kraus,
Alexander Arno Weber,
Felix Stollenwerk,
David KaczΓ©r,
Florian Mai,
Lucie Flek,
Rafet Sifa,
Nicolas Flores-Herr,
Joachim KΓΆhler,
Patrick Schramowski,
Michael Fromm,
Kristian Kersting
},
year = {2025},
journal = {arXiv preprint arXiv:2505:22232}
}π Links:
- Base Dataset: FineWeb2
- Related Work: FineWeb2 LLM Judging Section
