datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenDebateEvidence-Anonymized
Dataset Card for OpenDebateEvidence (Anonymized)
A collection of evidence used in collegiate and high school debate competitions,
with all debater-identifying columns removed.
This is an anonymized redistribution of
Yusuf5/OpenCaselist. The
argumentative content is byte-for-byte unchanged. 26 of the original 45 columns
have been dropped. See Anonymization for exactly what was
removed and why.
Dataset Details
Dataset Description
This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.reasoning-lite
Anonymous Reasoning Lite
This repository contains data accompanying an anonymous TMLR submission. It provides
1,211,520 sampled reasoning attempts from four model configurations on competition-math
and field-balanced multiple-choice questions. The data omits top-20 alternative-token
distributions while retaining realized-token log probabilities and ranks.
Contents
The same attempts are available in full and metadata-only representations:
Configuration group… See the full description on the dataset page: https://huggingface.co/datasets/AnonymizedTMLRSubmission/reasoning-lite.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.Fictional_Persona_Dialogs_Anonymized-benchmark
Fictional_Persona_Dialogs_Anonymized Benchmark Dataset
Short Summary:
A 68-pair synthetic Question-Answering (QA) dataset derived from anonymized fictional dialogues, specifically designed for rigorous Retrieval-Augmented Generation (RAG) system evaluation. It isolates and demonstrates the critical impact of contextual information on LLM accuracy.
Introduction & Motivation:
This dataset addresses the need for a clean, bias-minimized benchmark to accurately… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Fictional_Persona_Dialogs_Anonymized-benchmark.
