datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IfEvalCode-testsetvideo-SALMONN_2_testset
video-SALMONN 2 Benchmark
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.pySpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialGen-Testset.MMM-datasets-TestsetMultilingual Mutual Reinforcement Effect Mix Datasets
This is a Training set of OIELLM.
This Train set already formatted by OIELLM's format. The test set is in the another page in huggingface.
The MMM support 3 languages (English, Chinese and Japanese). And you must use task instruct words to define kind of task.
Mutual Reinforcement Effect.
OIELLM's input and output
MMM Dataset
The following is input and output format:
{
"input": "In 1953, filming of "On the Waterfront" starring… See the full description on the dataset page: https://huggingface.co/datasets/ganchengguang/MMM-datasets-Testset.SpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/BenjaminChai579/SpatialGen-Testset.testsetVidyaapati-Hindi-Konkani-TestsetKoala-test-setThis dataset is taken from https://github.com/arnav-gudibande/koala-test-set
ddro-testsetsRAG_legal_comparison_test_set
Dataset de Evaluación (Test Set) para Sistemas RAG en el Dominio Legal
Este repositorio contiene el conjunto de pruebas (Test Set / Golden Dataset) diseñado específicamente para auditar, evaluar y comparar el rendimiento de diferentes configuraciones de sistemas de Generación Aumentada por Recuperación (RAG) sobre documentación jurídica y administrativa española y europea.
El dataset se ha construido con el propósito de servir de base para métricas de evaluación RAG (como… See the full description on the dataset page: https://huggingface.co/datasets/AingeruBeOr/RAG_legal_comparison_test_set.environmental_registry_test_set
Environmental Registry Test Set
This dataset is the anonymized primary benchmark used for evaluating agentic Portuguese Text-to-SQL over a real PostgreSQL/PostGIS environmental-registry database. The underlying production database is not released, but the benchmark metadata and gold labels are provided for transparency and comparison.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/environmental_registry_test_set.LL144-Test-Setpostvalid-v2-test-set
LongShOTBench (Test Split)
Benchmark accompanying the NeurIPS 2026 submission
"A Benchmark for Omni-Modal Reasoning in Long Videos."
This dataset is shared anonymously to support double-blind review.
Purpose
LongShOTBench evaluates multimodal LLMs on long-form video understanding
across vision, speech, and non-speech audio, using intent-driven questions
and weighted criterion-level rubrics. Intended for evaluation only, not
training.
License
CC BY-NC-SA 4.0.… See the full description on the dataset page: https://huggingface.co/datasets/anonymsubs/postvalid-v2-test-set.HatEval_2019_Test_Set_Task5Test Set from HatEval (Basile et al, 2019), SemEval-2019 Task 5
testset_ragaskeval-testset
keval_test
The keval-testset is a dataset designed for training and validating the keval model.
The keval model follows the LLM-as-a-judge approach, which evaluates LLMs by assessing their responses to prompts from the ko-bench dataset. In other words, the keval model assigns scores to LLM-generated responses based on predefined evaluation criteria.
The keval-testset serves as a crucial resource for training and validating the keval model, enabling precise benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/davidkim205/keval-testset.CISA-testsetCISA-testset from İbrahim Temo's Memoir: Cross-Individual Sentiment Analysis Test Dataset for Historical Turkish
This test dataset is specifically designed for evaluating Cross-Individual Sentiment Analysis (CISA) performance on historical Turkish texts from İbrahim Temo's memoirs.
📚 Dataset Description
This test dataset contains 200 sentences extracted from the first 66 pages of İbrahim Temo's original memoirs "İttihad ve Terakki Cemiyetinin Teşekkülü ve Hidematı Vataniye ve İnkılâbı Milliye… See the full description on the dataset page: https://huggingface.co/datasets/dbbiyte/CISA-testset.rede_saude_publica_test_set
Rede Saude Publica Test Set
This dataset is the public-health transfer benchmark for the released Text-to-SQL agent artifact. It is a synthetic Brazilian public-health schema and test set used to measure cross-database generalization: the fine-tuned model was not trained on trajectories from this schema.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/rede_saude_publica_test_set.testSetTestSetGenkgrammar-testset
kgrammar-testset
The kgrammar-testset is a dataset designed for training and validating the kgrammar model, which identifies grammatical errors in Korean text and outputs the number of detected errors.
The kgrammar-testset was generated using GPT-4o. To create error-containing documents, a predefined prompt was used to introduce grammatical mistakes into responses when given a question. The dataset is structured to ensure a balanced distribution, consisting of 50% general questions… See the full description on the dataset page: https://huggingface.co/datasets/davidkim205/kgrammar-testset.cv10-uk-testset-clean-zipaGenerated by https://github.com/lingjzhu/zipa (zipa_large_crctc_500000_avg10.pth) and https://github.com/dmort27/epitran
test-lawyerHate_Political_Opponent_2021_Test_SetTest set from "Hate Towards the Political Opponent"(Grimminger et al., 2021)
stereoset_test_setOLID_2019_Test_SetTest set from OLID dataset (Zampieri et al., 2019), SemEval 2019
choreo_concepts_qna_test_set_v0.1This dataset contains the merge of rtweera/user_centric_results_v1 and rtweera/user_centric_results_v2 datasets for creating a unified testing for Choreo Concepts SLM finetuning project.
llama-3.1-medprm-reward-raw-test-setembedding-testsetsubtitle-summary-testset
Subtitle Summary & Keyword — Test set (YouTube)
방송 유튜브 자막 기반 누적 요약 + 검색 키워드 태스크의 평가용 테스트셋입니다.
jungsanghyun/subtitle-summary-dataset의 train/validation과 겹치지 않는 별도 방송으로, 동일 파이프라인(Qwen3-235B, vLLM, 자막-only)으로 생성했습니다.
규모
split
회차
레코드(5분)
test
114
519
약 25시간 · 요약 중앙값 59자 · 도메인: Entertainment 236 · News & Politics 103 · Pets & Animals 87 · Travel & Events 34 · Education 31 · Music 28
스키마
program_name, last_summary, 5min_script(입력) →… See the full description on the dataset page: https://huggingface.co/datasets/jungsanghyun/subtitle-summary-testset.
