open-ended
open-endedness-classifierspatial-reasoning-open-ended-qwen2.5-vl-mcq-sft_0spatial-reasoning-open-ended-qwen2.5-vl-mcq-sft_3qwen3-4b-mocking-diverse-open-endedphi-3.5-mini-instruct-cn-openended-kr0.2-a0.05-creativellama-3.1-8b-instruct-cn-openended-kr0.1-a0.5-creativellama-3.2-3b-instruct-cn-openended-kr0.2-a1.0-creativeqwen3-4b-nervous-diverse-open-ended-zh
Verbalized-Sampling-Open-Ended-QA
Verbalized-Sampling-Open-Ended-QA
This dataset demonstrates how Verbalized Sampling (VS) increases diversity in open-ended question answering while maintaining response quality. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.
Dataset Description
The Open-Ended QA dataset contains diverse responses from state-of-the-art LLMs to open-ended questions across various domains. This dataset evaluates:
Response diversity: Coverage of… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Open-Ended-QA.QUEST-SFT-Data-Open-ended
QUEST SFT Data (Open-ended)
Project Page | Paper | GitHub
Open-ended supervised fine-tuning trajectories for QUEST (tool-using assistant format). Split: train. Columns: messages (list[{role, content}]).
Load
from datasets import load_dataset
ds = load_dataset("osunlp/QUEST-SFT-Data-Open-ended", split="train", streaming=True)
row = next(iter(ds))
print(row.keys())
QUEST Family
Type
Resources
35B checkpoints
RL, MT+SFT, MT, SFT
30B checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-SFT-Data-Open-ended.gpqa-open-ended
GPQA Open-Ended
An open-ended reformulation of the GPQA (Graduate-Level Google-Proof Questions) benchmark. All 546 questions have been converted from multiple-choice to free-response format, preserving the original difficulty and domain expertise requirements while removing the ability to eliminate answers or pattern-match against option structure.
Why open-ended?
MCQ benchmarks have a ceiling problem for scalable oversight research: a non-expert judge who cannot solve… See the full description on the dataset page: https://huggingface.co/datasets/joanvelja/gpqa-open-ended.debug_MathVista_open_ended_to_remove
Dataset Card for "debug_MathVista_open_ended_to_remove"
More Information needed
open-ended
Dataset Summary
EVE-open-ended is a collection of open-ended question-answer pairs focused on Earth Observation (EO). The datasets cover a wide range of EO topics, including, but not limited to satellite imagery analysis, remote sensing techniques, environmental monitoring, LiDAR, etc.
The datasets are designed to facilitate the development and evaluation of large language models (LLMs) in understanding and generating responses related to Earth Observation.
Metrics… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/open-ended.debug_MMMU_open_ended_to_remove
Dataset Card for "debug_MMMU_open_ended_to_remove"
More Information needed
