datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GutenQA_Paragraphs
📚 GutenQA-Paragraphs
GutenQA-Paragraphs consists on the same 100 Public Domain Narrative Books used in GutenQA. In this version, passages are extracted at the paragraph level.
The GutenQA dataset, as available, is the result of applying the text segmentation method LumberChunker to the GutenQA-Paragraphs.
The dataset is organized into the following columns:
Book Name: The title of the book from which the passage is extracted.
Book ID: A unique integer identifier assigned to each… See the full description on the dataset page: https://huggingface.co/datasets/LumberChunker/GutenQA_Paragraphs.wikipedia-paragraph-sft
Wikipedia Paragraph Supervised Finetuning Dataset
Model Description
This dataset is designed for training language models to generate supervised finetuning data from raw text. It consists of text passages and corresponding question-answer pairs in JSONLines format.
Intended Use
The primary purpose of this dataset is to enable large language models (LLMs) to generate high-quality supervised finetuning data from raw text inputs, useful for creating custom… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-sft.RuBQ_2.0-paragraphs
RuBQ_2.0-paragraphs
For test and dev data see d0rj/RuBQ_2.0
