datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dspy_data_generation_LT
Lithuanian QA Dataset - Generated with DSPy & Gemma2 27B Q4
Introduction This dataset was created using DSPy, a Python framework that simplifies the generation of question and answer (QA) pairs from a given context. The dataset is composed of context, questions, and answers, all in Lithuanian. The context was primarily sourced from the following resources:
Lithuanian Wikipedia (lt.wikipedia.org) Lietuviškoji enciklopedija (vle.lt) Book: Vitalija Skėruvienė, Civilinė Teisė Mokomoji… See the full description on the dataset page: https://huggingface.co/datasets/ArturG9/Dspy_data_generation_LT.QA-text-generation-alpaca-data-cleaned
Dataset Card for mBART-QA-Processed
This dataset consists of tokenized pairs of instructions and contexts designed for fine-tuning Sequence-to-Sequence models (like mBART or T5) on Question Answering tasks.
Dataset Details
Dataset Description
The dataset is a processed version of a Question Answering corpus (SQuAD-like). It has been formatted to follow a specific prompt structure: instruction: {question} input: {context}. The targets (labels) are the direct… See the full description on the dataset page: https://huggingface.co/datasets/SOULAMA/QA-text-generation-alpaca-data-cleaned.
