jams
Datasets
All datasets matching “jams”ExposureQATriviaHG
TriviaHG is an extensive dataset crafted specifically for hint generation in question answering. Unlike conventional datasets, TriviaHG provides 10 hints per question instead of direct answers. This unique approach encourages users to engage in critical thinking and reasoning to derive the solution. Covering diverse question types across varying difficulty levels, the dataset is partitioned into training, validation, and test sets. These subsets facilitate the fine-tuning and… See the full description on the dataset page: https://huggingface.co/datasets/JamshidJDMY/TriviaHG.Parse
A reasoning-focused open-domain Question Answering benchmark for Persian (FA)covering Boolean, Factoid, and Multiple-choice questions with Reasoning + Multi-hop settings.
✨ Highlights
🧠 Designed to evaluate reasoning capabilities of LLMs in a low-resource language
✅ Supports Zero-shot, Few-shot, and Chain-of-Thought (CoT) evaluation
🧪 Includes scripts for automatic evaluation + fine-tuning utilities
👥 Comes with human evaluation interfaces (quality + difficulty… See the full description on the dataset page: https://huggingface.co/datasets/JamshidJDMY/Parse.PlausibleQA
PlausibleQA: A Large-Scale QA Dataset with Answer Plausibility Scores
PlausibleQA is a large-scale QA dataset designed to evaluate and enhance the ability of Large Language Models (LLMs) in distinguishing between correct and highly plausible incorrect answers. Unlike traditional QA datasets that primarily focus on correctness, PlausibleQA provides candidate answers annotated with plausibility scores and justifications.
🗂 Overview
📌 Key Statistics
10… See the full description on the dataset page: https://huggingface.co/datasets/JamshidJDMY/PlausibleQA.WikiHint
WikiHint is a human-annotated dataset designed for automatic hint generation and ranking for factoid questions. This dataset, based on Wikipedia, contains 5,000 hints for 1,000 questions and supports research in hint evaluation, ranking, and generation.
🗂 Overview
1,000 questions with 5,000 manually created hints.
Hints ranked by human annotators based on helpfulness.
Evaluated using LLMs (LLaMA, GPT-4) and human performance studies.
Supports hint ranking and automatic hint… See the full description on the dataset page: https://huggingface.co/datasets/JamshidJDMY/WikiHint.InferentialQA
Inferential Question Answering (Inferential QA)
Inferential Question Answering (Inferential QA) introduces a new class of reasoning QA tasks that challenge models to infer answers from indirect textual evidence rather than extracting them directly from answer-containing passages.
We present QUIT (QUestions requiring Inference from Texts) — a large-scale benchmark of 7,401 questions and 2.4 million passages, designed to evaluate how well modern retrieval-augmented… See the full description on the dataset page: https://huggingface.co/datasets/JamshidJDMY/InferentialQA.
