datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rag_thai_laws
Thai Laws Dataset
This dataset contains Thai law texts from the Office of the Council of State, Thailand.
The dataset has been cleaned and processed by the iApp Team to improve data quality and accessibility. The cleaning process included:
Converting system IDs to integer format
Removing leading/trailing whitespace from titles and text
Normalizing newlines to maintain consistent formatting
Removing excessive blank lines
The cleaned dataset is now available on Hugging Face for easy… See the full description on the dataset page: https://huggingface.co/datasets/iapp/rag_thai_laws.cs50-educational-rag
CS50 Pedagogical RAG Dataset
📜 Dataset Description
This repository contains the data artifacts for the undergraduate thesis, which explores the use of a pedagogical chatbot with Retrieval-Augmented Generation (RAG) for Harvard's CS50: Introduction to Computer Science course.
The project involved several stages of data processing, from raw content collection to the generation and curation of a high-quality evaluation dataset. To ensure full transparency and… See the full description on the dataset page: https://huggingface.co/datasets/dev-jonathanb/cs50-educational-rag.RAGPPI
RAG Benchmark for Protein-Protein Interactions (RAGPPI)
📊 Overview
Retrieving expected therapeutic impacts in protein-protein interactions (PPIs) is crucial in drug development, enabling researchers to prioritize promising targets and improve success rates. While Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) frameworks accelerate discovery, no benchmark exists for identifying therapeutic impacts in PPIs.
RAGPPI is the first factual QA benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Youngseung/RAGPPI.rag_thai_laws
Thai Laws Dataset
This dataset contains Thai law texts from the Office of the Council of State, Thailand.
The dataset has been cleaned and processed by the iApp Team to improve data quality and accessibility. The cleaning process included:
Converting system IDs to integer format
Removing leading/trailing whitespace from titles and text
Normalizing newlines to maintain consistent formatting
Removing excessive blank lines
The cleaned dataset is now available on Hugging Face for easy… See the full description on the dataset page: https://huggingface.co/datasets/atyko/rag_thai_laws.multi-tafseer-quran-rag
Quran Tafseer RAG Dataset
A structured Arabic dataset of Quranic tafseer collected from eight classical and modern tafseer books.The dataset contains verse-aligned tafseer passages designed for Retrieval-Augmented Generation (RAG) systems and Arabic NLP research.
Each record links a Quran verse with its corresponding tafseer explanation from one of the tafseer books and includes rich metadata such as surah information, tafseer source, and embedding-ready text.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/omaressam1111/multi-tafseer-quran-rag.rag-bioask-plDerived from rag-datasets/mini-bioasq.
This is a subset of the above dataset translated to Polish using DeepL. It should help you in finetuning LLMs for RAG purposes.
RAGPPI_Atomics
RAG Benchmark for Protein-Protein Interactions (RAGPPI)
📊 Overview
Retrieving expected therapeutic impacts in protein-protein interactions (PPIs) is crucial in drug development, enabling researchers to prioritize promising targets and improve success rates. While Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) frameworks accelerate discovery, no benchmark exists for identifying therapeutic impacts in PPIs.
RAGPPI is the first factual QA benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Youngseung/RAGPPI_Atomics.
