datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MANTA-1M
Abstract
We introduce MANTA, an automated pipeline that generates high-quality large-scale instruction fine-tuning datasets from massive web corpora while preserving their diversity and scalability. By extracting structured syllabi from web documents and leveraging high-performance LLMs, our approach enables highly effective query-response generation with minimal human intervention. Extensive experiments on 8B-scale LLMs demonstrate that fine-tuning on the MANTA-1M dataset… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/MANTA-1M.KoMT-Bench
KoMT-Bench
Introduction
We present KoMT-Bench, a benchmark designed to evaluate the capability of language models in following instructions in Korean.
KoMT-Bench is an in-house dataset created by translating MT-Bench [1] dataset into Korean and modifying some questions to reflect the characteristics and cultural nuances of the Korean language.
After the initial translation and modification, we requested expert linguists to conduct a thorough review of our benchmark… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/KoMT-Bench.Ko-LongRAG
Abstract
The rapid advancement of large language models (LLMs) significantly enhances long-context Retrieval-Augmented Generation (RAG), yet existing benchmarks focus primarily on English. This leaves low-resource languages without comprehensive evaluation frameworks, limiting their progress in retrieval-based tasks. To bridge this gap, we introduce Ko-LongRAG, the first Korean long-context RAG benchmark. Unlike conventional benchmarks that depend on external retrievers… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/Ko-LongRAG.
