datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
A-Dataset-for-Complex-Reasoning-over-Textual-Knowledge-Graphs-in-Medicine
RiTeK: Medical Textual Knowledge Graph QA Benchmark
RiTeK is a benchmark for complex reasoning over medical Textual Knowledge Graphs (medical TKGs). It evaluates whether retrieval systems and Large Language Models (LLMs) can answer realistic medical questions by using both relational paths and textual entity descriptions.
Dataset Overview
The benchmark contains three medical graph QA subsets:
Dataset
Directory
Splits
KG file
ADint
Adint
train / dev / test… See the full description on the dataset page: https://huggingface.co/datasets/ChenAI2015/A-Dataset-for-Complex-Reasoning-over-Textual-Knowledge-Graphs-in-Medicine.multidomain-complex-text-pool
Complex Text Pool Dataset
Overview
A curated collection of complex, long-form English texts sampled from 9 diverse domains. Each document has been truncated to a maximum of 4,000 characters, preserving clean sentence boundaries. The dataset is designed to provide challenging, real-world text samples across multiple subject areas.
Categories and Sample Counts
Category
Samples
news
9,999
encyclopedic
10,000
conversational
10,000… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/multidomain-complex-text-pool.synthetic-complex-Text-to-SQL
Synthetic Complex Text-to-SQL
Synthetic Text-to-SQL using multiple joins, WHERE statements, window and aggregate functions on filtered SQL from bigcode/the-stack.
Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_TextsHere presented a partially synthesized dataset, developed utilizing the GPT-4 model, for the purpose of NLG, particulary for the task of hierarchical generation of longer texts from short summaries. The creation of this dataset was undertaken as a component of my thesis paper. It incorporates excerpts from prominent British and American novels, from which plots, summaries, and metadata have been derived using GPT-4 API to facilitate extensive future research.
The metadata included in the… See the full description on the dataset page: https://huggingface.co/datasets/Fleur-roar/Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_Texts.
