multilingual-evaluation
Evaluation-Multilingual-VC
Evaluation-Multilingual-VC
We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon,
Filter languages that support by Whisper Large V3 to evaluate WER automatically,
Only take test set, sort by up votes.
Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows.
Only build first 500 rows for each language
Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.NTEU_Multilingual_Evaluation_Dataset
Dataset Card for NTEU Multilingual Evaluation Dataset
Dataset Summary
This evaluation dataset for Machine Translation was created by the NTEU - Neural Translation for the EU project.
The evaluation dataset includes around 1,000 parallel sentences in the 24 official European languages.
The original NTEU dataset has been cleaned and filtered by removing empty lines and near-duplicates, and it has been augmented with Catalan.
The Catalan version was manually produced by a… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/NTEU_Multilingual_Evaluation_Dataset.ponys-ai-multilingual-companion-evaluation
Ponys.ai Multilingual AI Companion Evaluation Protocols
This public collection contains 20 reusable evaluation protocols for AI companion and character experiences. It covers conversation memory, persona consistency, consent recovery, visual continuity, code switching, and regional language behavior across Japanese, Korean, Latin American Spanish, Brazilian Portuguese, Simplified Chinese, Traditional Chinese, and English.
Each protocol includes structured metadata and a CSV… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-ai-multilingual-companion-evaluation.EvalyxAi-multilingual-rlhf-evaluation-data
Evalyxai Multilingual RLHF Evaluation Dataset
Overview
This dataset contains 60 expert-level evaluation tasks in APEX-Agents format for training and evaluating multilingual AI models.
Dataset Contents
Thai Language: 10 expert-level tasks
Chinese Language: 10 expert-level tasks
Japanese Language: 10 expert-level tasks
Korean Language: 10 expert-level tasks
Vietnamese Language: 10 expert-level tasks
Indonesian Language: 10 expert-level tasks
Total: 60 bilingual… See the full description on the dataset page: https://huggingface.co/datasets/spinxgritone/EvalyxAi-multilingual-rlhf-evaluation-data.multilingual-llm-evaluation
Multilingual LLM Evaluation
A small evaluation dataset for comparing language models across English, Hindi, and Spanish.
Columns
language: language code (en, hi, or es)
question: question provided to the model
expected_answer: reference answer used for scoring
Intended use
This dataset can be used to compare model accuracy, language adherence, and response speed across languages.
Limitations
This is a small demonstration dataset and… See the full description on the dataset page: https://huggingface.co/datasets/userhuggingface4321/multilingual-llm-evaluation.
