datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lfqalfqa_support_docsSupport documents for building https://huggingface.co/vblagoje/bart_lfqa model
lfqa_idlegal-reasoning-lfqa-merged
Dataset Card for "legal-reasoning-lfqa-merged"
More Information needed
lfqa_test_docslfqa-max-answer-length-512
Dataset Card
Dataset Description
The dataset contains simple, long-form answers to questions and corresponding contexts.
Similar to ELI5 but with context.
This dataset is a filtered version of LLukas22/lfqa_preprocessed,
which in turn is a processed and simplified version of of vblagoje's lfqa_support_docs and lfqa datasets.
I have filtered out overly long answers, based on the number of tokens in the answer using the LED tokenizer.
It can be reproduced with the notebook… See the full description on the dataset page: https://huggingface.co/datasets/stefanbschneider/lfqa-max-answer-length-512.lfqa_squadlegal-reasoning-lfqa-synthetic
Dataset Card for "legal-reasoning-lfqa-synthetic"
More Information needed
lfqa_preprocessed
Dataset Card for "stefanbschneider/lfqa_preprocessed"
Dataset Description
The dataset contains simple, long-form answers to questions and corresponding contexts.
Similar to ELI5 but with context.
This dataset is a filtered version of LLukas22/lfqa_preprocessed,
which in turn is a processed and simplified version of of vblagoje's lfqa_support_docs and lfqa datasets.
This dataset (stefanbschneider/lfqa_preprocessed) has filtered out overly long answers (around 3%).
It can… See the full description on the dataset page: https://huggingface.co/datasets/stefanbschneider/lfqa_preprocessed.lfqa-max-answer-length-1024
Dataset Card
Dataset Description
The dataset contains simple, long-form answers to questions and corresponding contexts.
Similar to ELI5 but with context.
This dataset is a filtered version of LLukas22/lfqa_preprocessed,
which in turn is a processed and simplified version of of vblagoje's lfqa_support_docs and lfqa datasets.
I have filtered out overly long answers, based on the number of tokens in the answer using the LED tokenizer.
It can be reproduced with the notebook… See the full description on the dataset page: https://huggingface.co/datasets/stefanbschneider/lfqa-max-answer-length-1024.lfqa-preprocessed-iteli5_lfqa_bestlfqa_discourseLFQA discourse contains discourse annotations of long-form answers.
- [VALIDITY]: Validity annotations of (question, answer) pairs.
- [ROLE]: Role annotations of valid answer paragraphs.lfqa_preprocessed
Dataset Card for "lfqa_preprocessed"
Dataset Summary
This is a simplified version of vblagoje's lfqa_support_docs and lfqa datasets.
It was generated by me to have a more straight forward way to train Seq2Seq models on context based long form question answering tasks.
Dataset Structure
Data Instances
An example of 'train' looks as follows.
{
"question": "what's the difference between a forest and a wood?",
"answer": "They're used… See the full description on the dataset page: https://huggingface.co/datasets/LLukas22/lfqa_preprocessed.eli5-lfqa
ELI5: Long Form Question Answering
https://arxiv.org/pdf/1907.09190
Items (train): 272_634
Items (val): 1_507
Downloaded from:
https://www.kaggle.com/datasets/trandaiphu/eli5-dataset
Tokens ctxs length distribution summary:
min: 152
mean: 3745.35
median (p50): 3736.00
p90: 4037.00
p95: 4152.00
p99: 4432.00
max: 6530
Tokens answers length distribution summary:
rows: 272_634
min: 8
mean: 401.31
median (p50): 217.00
p90: 820.70
p95: 1230.00
p99:… See the full description on the dataset page: https://huggingface.co/datasets/aitetic/eli5-lfqa.lfqa_expert_pairwise_human_preference_no_reasoningLFQA-HP-1M
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
Overview
LFQA-HP-1M is a large-scale human preference dataset for Long-Form
Question Answering (LFQA).
The dataset is designed to support research in:
Human preference modeling
Pairwise answer evaluation
Fine-grained rubric-based scoring
LLM-as-a-Judge evaluation
Long-form response generation benchmarking
LFQA tasks require multi-sentence, explanatory, reasoning-based answers
rather than… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/LFQA-HP-1M.lfqalfqa_with_supports_subsetattributionBench_lfqa_expertqa_dpo
Dataset Card for "attributionBench_lfqa_expertqa_dpo"
More Information needed
intent-aware-lfqa-intent-implicitbart-finetuned-eli5_lfqa_best_slice-512_2023-12-10_runLFQAeli5_lfqa_best_sliceintent-aware-lfqa-multivieweli5-lfqa-combined
eli5-lfqa-combined
Postprocessed from eli5-lfqa: (ctxs, question+answers[])
Assembled by: dataset_assembler.py
Size: 1.1B
intent-aware-lfqa-baselinelfqa_eval
Dataset Card for "lfqa_eval"
More Information needed
LFQA_eval_dataset_unit_tests_justificationlfqa-textbugger
