renataaraujoe/Bilingual-LLM-Eval-106
Bilingual-LLM-Eval-106 π Overview Bilingual-LLM-Eval-106 is a curated, human-annotated evaluation dataset of 106 LLM response pairs in English and Portuguese, designed to benchmark model performance across multiple quality dimensions. The dataset focuses on realistic and challenging evaluation scenarios, including adversarial prompts, ambiguous queries, and hard negatives. It is intended for: LLM evaluation and benchmarking LLM-as-a-Judge research Multilingualβ¦ See the full description on the dataset page: https://huggingface.co/datasets/renataaraujoe/Bilingual-LLM-Eval-106.
Bilingual-LLM-Eval-106
π Overview
Bilingual-LLM-Eval-106 is a curated, human-annotated evaluation dataset of 106 LLM response pairs in English and Portuguese, designed to benchmark model performance across multiple quality dimensions.
The dataset focuses on realistic and challenging evaluation scenarios, including adversarial prompts, ambiguous queries, and hard negatives.
It is intended for:
- LLM evaluation and benchmarking
- LLM-as-a-Judge research
- Multilingual robustness analysis
- Annotation quality research
π― Objectives
This dataset was built to:
- Design and apply a professional annotation rubric
- Collect diverse prompt types:
- Factual
- Reasoning
- Adversarial
- Ambiguous
- Annotate responses using multi-dimensional quality labels
- Ensure bilingual parity (EN + PT) with equivalent difficulty
- Validate annotation consistency using:
- Cohenβs Kappa
- Bilingual calibration
- Study LLM-as-a-Judge biases through controlled experiments
π Dataset Structure
The dataset consists of annotated CSV files:
annotations_EN_batch1.csvβ English samplesannotations_PT_batch1.csvβ Portuguese samplesannotations_mirrored_batch.csvβ Cross-lingual mirrored examplesannotations_edge_cases.csvβ Adversarial and difficult cases
Each row represents a prompt + response pair with annotations.
π§Ύ Annotation Dimensions
Each response is evaluated across multiple dimensions:
- Faithfulness β factual correctness and grounding
- Relevance β alignment with the prompt
- Fluency β linguistic quality
- Completeness β coverage of required information
- Safety β harmful or risky content
Labels follow a structured rubric inspired by industry standards used in:
- Scale AI
- Anthropic
- DataAnnotation
π Languages
- English (EN)
- Portuguese (PT)
The dataset includes:
- Independent annotations per language
- Mirrored examples for cross-lingual consistency analysis
π§ͺ Research Use Cases
This dataset enables:
π LLM Evaluation
Benchmark models across multiple qualitative dimensions beyond accuracy.
βοΈ LLM-as-a-Judge Analysis
Study bias, inconsistency, and failure modes in automated evaluation systems.
π Multilingual Testing
Compare performance across English and Portuguese under equivalent conditions.
π― Robustness Testing
Evaluate models on:
- Edge cases
- Adversarial prompts
- Ambiguous inputs
π Annotation Quality
- Annotated following a custom-built professional rubric
- Includes hard negatives and adversarial cases
- Validated with:
- Inter-annotator agreement (Cohenβs Kappa)
- Cross-lingual calibration
π Data Format
Typical columns may include:
promptresponselanguagefaithfulnessrelevancefluencycompletenesssafetynotes(optional)
β οΈ Limitations
- Dataset size is intentionally small (106 samples) for high-quality evaluation, not training
- Domain coverage is diverse but not exhaustive
- Some annotations may include subjective judgment despite calibration
π€ Contributions
This dataset was created as an independent research and engineering project focused on LLM evaluation quality and methodology.
β Citation
@dataset{bilingualllmeval1062026, author = {Renata de Araujo}, title = {Bilingual-LLM-Eval-106: A Human-Annotated Benchmark for EnglishβPortuguese LLM Evaluation}, year = {2026}, publisher = {renataaraujoe}, howpublished = {\url{https://huggingface.co/datasets/renataaraujoe/Bilingual-LLM-Eval-106}}, note = {Version 1.0} }
