datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu
FineWeb-Edu (Lance Format)
A Lance-formatted version of FineWeb-Edu — over 1.5 billion educational web passages with cleaned text, source metadata, language detection signals, and 384-dim text embeddings — available directly from the Hub at hf://datasets/lance-format/fineweb-edu/data/train.lance.
Key features
Cleaned passage text in the text column with the source url and title carried alongside.
Language detection signals (language, language_probability) for filtered… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fineweb-edu.EduBench
EduBench 📚
EduBench é um benchmark em português brasileiro para avaliação de Large Language Models (LLMs) em tarefas educacionais, composto por 3,149 questões discursivas extraídas de vestibulares de alta competitividade.
GitHub
Paper
Dataset Description
Fontes
USP: Universidade de São Paulo
UNICAMP: Universidade Estadual de Campinas
UNESP: Universidade Estadual Paulista
Período
2015-2025 (11 anos de provas)
Áreas do… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/EduBench.EduFeedback
EduFeedback
Alternating dataset example: a single curated multi-turn conversation yields a complete (prompt, chosen, rejected) triplet on its own — the direct early response becomes chosen and a later, less-direct response becomes rejected. Both sides come from the same real dialog, so no synthetic LLM generation is needed to fill in the rejected side.
EduFeedback is a synthetically generated, multi-turn conversational
preference dataset in an educational tutoring setting… See the full description on the dataset page: https://huggingface.co/datasets/miria0/EduFeedback.cs50-educational-rag
CS50 Pedagogical RAG Dataset
📜 Dataset Description
This repository contains the data artifacts for the undergraduate thesis, which explores the use of a pedagogical chatbot with Retrieval-Augmented Generation (RAG) for Harvard's CS50: Introduction to Computer Science course.
The project involved several stages of data processing, from raw content collection to the generation and curation of a high-quality evaluation dataset. To ensure full transparency and… See the full description on the dataset page: https://huggingface.co/datasets/dev-jonathanb/cs50-educational-rag.crypto-education-en-corpus
Crypto Education Corpus (EN)
A curated educational corpus about cryptocurrency and blockchain technology, designed for building and evaluating RAG (Retrieval-Augmented Generation) systems.
Source Distribution
Source
Documents
Share
iqwiki.com
2,379
68.2%
academy.binance.com
581
16.7%
kraken.com
183
5.2%
coinbase.com
125
3.6%
gemini.com
94
2.7%
ethereum.org
74
2.1%
investopedia.com
51
1.5%
Word Count Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/kskada/crypto-education-en-corpus.code-education
Code de l'éducation, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-education.PLMoABPolish Language Model Awareness Benchmark
