datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JudgeBias-DPO-RefFree-testset
JudgeBias-DPO-RefFree-testset
A fixed evaluation benchmark (1,000 samples) for assessing DPO-trained LLM judges on materials science synthesis recipe evaluation in a reference-free setting.
Purpose
This test set supports two evaluation methods:
1. Reward Accuracy (Log-Probability)
Use prompt, chosen, and rejected to compute implicit reward accuracy without generation:
reward_chosen = log P(chosen | prompt)
reward_rejected = log P(rejected | prompt)
accuracy =… See the full description on the dataset page: https://huggingface.co/datasets/iknow-lab/JudgeBias-DPO-RefFree-testset.sinhala-test-set-50k
Sinhala Test Set - 50K Sentences
A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.Ecom-Chatbot-Test-Set
Ecom Chatbot Synthetic Test Set
A 2,000-sample fully synthetic test set for evaluating e-commerce chatbot models fine-tuned on
rescommons/Ecom-Chatbot-Finetuning-Dataset.
Designed for zero-contamination evaluation — all products, orders, customer names, and responses
are synthetically generated and do not overlap with the training data.
Dataset Summary
Split
Samples
test
2,000
Group Distribution
Group
Count
Description
A
667… See the full description on the dataset page: https://huggingface.co/datasets/V1rtucious/Ecom-Chatbot-Test-Set.kitrec-test-setb
KitREC Test Dataset - Set B
Evaluation test dataset for the KitREC (Knowledge-Instruction Transfer for Recommendation) cross-domain recommendation system.
Dataset Description
This test dataset is designed for evaluating fine-tuned LLMs on cross-domain recommendation tasks across 10 different user types.
Dataset Summary
Attribute
Value
Candidate Set
Set B (Random (Fair baseline))
Total Samples
30,000
Source Domain
Books
Target Domains
Movies &… See the full description on the dataset page: https://huggingface.co/datasets/Younggooo/kitrec-test-setb.kitrec-test-seta
KitREC Test Dataset - Set A
Evaluation test dataset for the KitREC (Knowledge-Instruction Transfer for Recommendation) cross-domain recommendation system.
Dataset Description
This test dataset is designed for evaluating fine-tuned LLMs on cross-domain recommendation tasks across 10 different user types.
Dataset Summary
Attribute
Value
Candidate Set
Set A (Hybrid (Hard negatives + Random))
Total Samples
30,000
Source Domain
Books
Target Domains… See the full description on the dataset page: https://huggingface.co/datasets/Younggooo/kitrec-test-seta.subtitle-summary-testset
Subtitle Summary & Keyword — Test set (YouTube)
방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다.
jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다.
규모
split
회차
레코드(5분)
test
114
519
약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28
스키마
program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.
