datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jp-llm-evaluator-training
Japanese LLM Evaluator Training Dataset
It realeased on NLP2025 Constructing Open-source Large Language Model Evaluator for Japanese
Overview
Japanese LLM Evaluator Training Dataset is a dataset using for training Japanese LLM evaluator, which is focus on evaluate Japanese LLM from mutiple perspectives and meeting diverse evaluation requirements.
Content
The dataset includes 1000 diveser score rubrics. For every score rubrics, we generate 20 different… See the full description on the dataset page: https://huggingface.co/datasets/ku-nlp/jp-llm-evaluator-training.grade-aware-llm-training-data
Grade-Aware LLM Training Dataset
Dataset Description
This dataset contains 1,107,690 high-quality instruction-tuning examples for grade-aware text simplification, designed for fine-tuning large language models to simplify text to specific reading grade levels with precision and semantic consistency.
Dataset Summary
Total Examples: 1,107,690
Task: Text simplification with precise grade-level targeting
Language: English
Grade Range: 1-12+ (precise 2-decimal… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade-aware-llm-training-data.llm_qwen_training_gazeta-datasethigh_educability_training_split
high_educability_training_split
Textos em português selecionados para treinamento: originais de Carolina e Wikipédia classificados nas classes 3 ou 4 pelo educability-norberto-mini-4class-v1, mais as reformulações publicadas vinculadas aos originais elegíveis.
Carregamento
from datasets import load_dataset
ds = load_dataset(
"br-llm-data/high_educability_training_split",
split="train",
streaming=True,
)
registro = next(iter(ds))
Conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/high_educability_training_split.Stratego-Datasetstrategorandomn_educability_training_split
