datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smugri-flores-testsetMultilingual FLORES-based benchmark for Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian. Expanded with Proper Karelian, Ludian, and Veps.
Please, cite the following paper if you use Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian datasets:
@inproceedings{
yankovskaya2023machine,
title={Machine Translation for Low-resource Finno-Ugric Languages},
author={Lisa Yankovskaya and Maali Tars and Andre T{\"a}ttar and Mark… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/smugri-flores-testset.llama-3.1-medprm-reward-test-set🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its scalability is not limited to… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-test-set.gdelt-rag-golden-testset-v2
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v2.sinhala-test-set-50k
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.allenai_dolma_test_set
Dolma
Dolma is a dataset of 3 trillion tokens from a diverse mix of web content, academic publications, code, books, and encyclopedic materials.
More information:
Read Dolma manuscript and its Data Sheet on ArXiv;
Explore the open source tools we created to curate Dolma.
Want to request removal of personal data? Use this form to notify us of documents containing PII about a specific user.
To learn more about the toolkit used to create Dolma, including how to replicate this… See the full description on the dataset page: https://huggingface.co/datasets/maxkaufmann/allenai_dolma_test_set.JudgeBias-DPO-RefFree-testset
JudgeBias-DPO-RefFree-testset
A fixed evaluation benchmark (1,000 samples) for assessing DPO-trained LLM judges on materials science synthesis recipe evaluation in a reference-free setting.
Purpose
This test set supports two evaluation methods:
1. Reward Accuracy (Log-Probability)
Use prompt, chosen, and rejected to compute implicit reward accuracy without generation:
reward_chosen = log P(chosen | prompt)
reward_rejected = log P(rejected | prompt)
accuracy =… See the full description on the dataset page: https://huggingface.co/datasets/iknow-lab/JudgeBias-DPO-RefFree-testset.gdelt-rag-golden-testset-v3
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v3.Mol-LLM-testset
Dataset summary
This dataset includes the evaluation benchmark used in the Mol-LLM paper, covering a broad range of molecular tasks for multimodal molecular language models.
It provides test splits with natural-language instructions, 1D molecular sequences, and labels, enabling fair comparison of generalist molecular LLMs under in-distribution and out-of-distribution settings.
Supported tasks and modalities
Task groups: reaction prediction (FS, RS, RP), property… See the full description on the dataset page: https://huggingface.co/datasets/KU-AGI/Mol-LLM-testset.environmental_registry_test_set
Environmental Registry Test Set
This dataset is the anonymized primary benchmark used for evaluating agentic Portuguese Text-to-SQL over a real PostgreSQL/PostGIS environmental-registry database. The underlying production database is not released, but the benchmark metadata and gold labels are provided for transparency and comparison.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/environmental_registry_test_set.biosciences-golden-testset
Biosciences RAG Golden Test Set
Dataset Description
This dataset contains 12 question-answering pairs for evaluating RAG systems on biomedical research topics. The QA pairs were synthetically generated using the RAGAS framework from 140 source documents spanning knowledge graphs, LLM applications in biomedicine, protein interaction databases, and gene-to-phenotype mapping.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation ground truth… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-golden-testset.gdelt-rag-golden-testset-v4
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v4.gdelt-rag-golden-testset
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset.rede_saude_publica_test_set
Rede Saude Publica Test Set
This dataset is the public-health transfer benchmark for the released Text-to-SQL agent artifact. It is a synthetic Brazilian public-health schema and test set used to measure cross-database generalization: the fine-tuned model was not trained on trajectories from this schema.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/rede_saude_publica_test_set.Ecom-Chatbot-Test-Set
Ecom Chatbot Synthetic Test Set
A 2,000-sample fully synthetic test set for evaluating e-commerce chatbot models fine-tuned on
rescommons/Ecom-Chatbot-Finetuning-Dataset.
Designed for zero-contamination evaluation — all products, orders, customer names, and responses
are synthetically generated and do not overlap with the training data.
Dataset Summary
Split
Samples
test
2,000
Group Distribution
Group
Count
Description
A
667… See the full description on the dataset page: https://huggingface.co/datasets/V1rtucious/Ecom-Chatbot-Test-Set.testsetbuddhist-scholar-test-set
Vietnamese Buddhist Scholar Test Set
Dataset Description
This dataset contains 1008 Vietnamese question-answer pairs focused on Buddhist teachings and literature. The dataset was created to evaluate chatbots' knowledge and understanding of Buddhist concepts, particularly for Vietnamese-speaking users.
Dataset Details
Dataset Summary
Language: Vietnamese
Task: Question Answering, Chatbot Evaluation
Domain: Buddhism, Religious Studies
Size: 1008… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddhist-scholar-test-set.kitrec-test-seta
KitREC Test Dataset - Set A
Evaluation test dataset for the KitREC (Knowledge-Instruction Transfer for Recommendation) cross-domain recommendation system.
Dataset Description
This test dataset is designed for evaluating fine-tuned LLMs on cross-domain recommendation tasks across 10 different user types.
Dataset Summary
Attribute
Value
Candidate Set
Set A (Hybrid (Hard negatives + Random))
Total Samples
30,000
Source Domain
Books
Target Domains… See the full description on the dataset page: https://huggingface.co/datasets/Younggooo/kitrec-test-seta.northeast-languages-test-set
Northeast Languages Test Set
A curated test set of 500 deduplicated sentences per language for evaluating language models on Northeast Indian languages.
Languages
This dataset contains test data for 9 Northeast Indian languages:
Assamese (asm) - 500 sentences
Garo (grt) - 500 sentences
Khasi (kha) - 500 sentences
Kokborok (trp) - 500 sentences
Meitei (mni) - 500 sentences
Mizo (lus) - 500 sentences
Naga (nag) - 500 sentences
Nyishi (njz) - 500 sentences
Pnar (pbv) -… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-languages-test-set.subtitle-summary-testset
Subtitle Summary & Keyword — Test set (YouTube)
방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다.
jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다.
규모
split
회차
레코드(5분)
test
114
519
약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28
스키마
program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.kitrec-test-setb
KitREC Test Dataset - Set B
Evaluation test dataset for the KitREC (Knowledge-Instruction Transfer for Recommendation) cross-domain recommendation system.
Dataset Description
This test dataset is designed for evaluating fine-tuned LLMs on cross-domain recommendation tasks across 10 different user types.
Dataset Summary
Attribute
Value
Candidate Set
Set B (Random (Fair baseline))
Total Samples
30,000
Source Domain
Books
Target Domains
Movies &… See the full description on the dataset page: https://huggingface.co/datasets/Younggooo/kitrec-test-setb.subtitle-summary-testset-sft
Subtitle Summary & Keyword — Test set SFT (YouTube)
jungsanghyun/subtitle-summary-testset를 messages 형식으로 가공한 평가용 버전.
형식 (system 없음, 단일 턴)
입력(user): [이전 요약] {last_summary}\n[자막] {5min_script}
출력(assistant): [요약] {한 문장}\n[키워드] {검색어}
split
examples
test
519
파싱 \[요약\]\s*(.+) / \[키워드\]\s*(.+). chat_template.jinja 동봉.
라이선스
cc-by-nc-4.0, 연구·비상업.
testset
