CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ku-nlp /jp-llm-evaluator-training Japanese LLM Evaluator Training Dataset It realeased on NLP2025 Constructing Open-source Large Language Model Evaluator for Japanese Overview Japanese LLM Evaluator Training Dataset is a dataset using for training Japanese LLM evaluator, which is focus on evaluate Japanese LLM from mutiple perspectives and meeting diverse evaluation requirements. Content The dataset includes 1000 diveser score rubrics. For every score rubrics, we generate 20 different… See the full description on the dataset page: https://huggingface.co/datasets/ku-nlp/jp-llm-evaluator-training.tabular10K<n<100K0 likes24 downloads2y agoHugging Face02yimingwang123 /grade-aware-llm-training-data Grade-Aware LLM Training Dataset Dataset Description This dataset contains 1,107,690 high-quality instruction-tuning examples for grade-aware text simplification, designed for fine-tuning large language models to simplify text to specific reading grade levels with precision and semantic consistency. Dataset Summary Total Examples: 1,107,690 Task: Text simplification with precise grade-level targeting Language: English Grade Range: 1-12+ (precise 2-decimal… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade-aware-llm-training-data.tabulartext-generation1M<n<10M0 likes20 downloads1y agoHugging Face03masterkristall /llm_qwen_training_gazeta-datasettabular1K<n<10K0 likes15 downloads2mo agoHugging Face04br-llm-data /high_educability_training_splitgated high_educability_training_split Textos em português selecionados para treinamento: originais de Carolina e Wikipédia classificados nas classes 3 ou 4 pelo educability-norberto-mini-4class-v1, mais as reformulações publicadas vinculadas aos originais elegíveis. Carregamento from datasets import load_dataset ds = load_dataset( "br-llm-data/high_educability_training_split", split="train", streaming=True, ) registro = next(iter(ds)) Conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/high_educability_training_split.tabulartext-generation1M<n<10M0 likes10 downloads17d agoHugging Face05STRATEGO-LLM-TRAINING /Stratego-Datasettabularn<1K0 likes4 downloads9mo agoHugging Face06STRATEGO-LLM-TRAINING /strategogatedtabular1K<n<10K0 likes1 downloads8mo agoHugging Face07br-llm-data /randomn_educability_training_splitgatedtabular1M<n<10M0 likes1 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.