datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Designed-Vocalizations-Dataset
Designed Vocalizations Dataset
Paper · Demo & audio samples
The Designed Vocalizations Dataset supports voice conversion for designed vocalizations
— monster growls, robotic voices, and other sound-designed timbres — an area left
underexplored by benchmarks that focus on natural human speech. It curates diverse raw vocal
sources (speech and animal / non-linguistic sounds) and applies professional vocal-effects
processing to produce corresponding effect-modified variants. A… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset.K-MMBench
K-MMBench
We introduce K-MMBench, a Korean adaptation of the MMBench [1] designed for evaluating vision-language models.
By translating the dev subset of MMBench into Korean and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
K-MMBench consists of questions across 20 evaluation dimensions, such as identity reasoning, image emotion, and attribute recognition, allowing a thorough… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-MMBench.K-MMStar
K-MMStar
We introduce K-MMStar, a Korean adaptation of the MMStar [1] designed for evaluating vision-language models.
By translating the val subset of MMStar into Korean and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
(We observe that there are unanswerable cases (e.g., multiple images required to answer the question but only has a single image, vague questions or options) in the… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-MMStar.K-SEED
K-SEED
We introduce K-SEED, a Korean adaptation of the SEED-Bench [1] designed for evaluating vision-language models.
By translating the first 20 percent of the test subset of SEED-Bench into Korean, and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
K-SEED consists of questions across 12 evaluation dimensions, such as scene understanding, instance identity, and instance attribute… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-SEED.K-DTCBench
K-DTCBench
We introduce K-DTCBench, a newly developed Korean benchmark featuring both computer-generated and handwritten documents, tables, and charts.
It consists of 80 questions for each image type and two questions per image, summing up to 240 questions in total.
This benchmark is designed to evaluate whether vision-language models can process images in different formats and be applicable for diverse domains.
All images are generated with made-up values and statements for… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-DTCBench.offsetbias
Dataset Card for OffsetBias
Dataset Description:
💻 Repository: https://github.com/ncsoft/offsetbias
📜 Paper: OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Dataset Summary
OffsetBias is a pairwise preference dataset intended to reduce common biases inherent in judge models (language models specialized in evaluation). The dataset is introduced in paper OffsetBias: Leveraging Debiased Data for Tuning Evaluators. OffsetBias contains 8,504 samples… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/offsetbias.K-LLaVA-W
K-LLaVA-W
We introduce K-LLaVA-W, a Korean adaptation of the LLaVA-Bench-in-the-wild [1] designed for evaluating vision-language models.
By translating the LLaVA-Bench-in-the-wild into Korean and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
(Since our goal was to build a benchmark exclusively focused in Korean, we change the English texts in images into Korean for localization.)… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-LLaVA-W.offsetbias-NCSOFT
Dataset Card for OffsetBias
Dataset Description:
💻 Repository: https://github.com/ncsoft/offsetbias
📜 Paper: OffsetBias: Leveraging Debiased Data for Tuning Evaluators
Dataset Summary
OffsetBias is a pairwise preference dataset intended to reduce common biases inherent in judge models (language models specialized in evaluation). The dataset is introduced in paper OffsetBias: Leveraging Debiased Data for Tuning Evaluators. OffsetBias contains 8,504 samples… See the full description on the dataset page: https://huggingface.co/datasets/GenRM/offsetbias-NCSOFT.NCSOFT__Llama-VARCO-8B-Instruct-details
Dataset Card for Evaluation run of NCSOFT/Llama-VARCO-8B-Instruct
Dataset automatically created during the evaluation run of model NCSOFT/Llama-VARCO-8B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NCSOFT__Llama-VARCO-8B-Instruct-details.NCSOFT__Llama-VARCO-8B-InstructNCSOFT_offsetbias-PreferenceShareGPTChatbot-4-NCSOFT-VARCO-8B
