datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RDB2G-Bench
RDB2G-Bench
This is an offical dataset of the paper RDB2G-Bench: A Comprehensive Benchmark for Automatic Graph Modeling of Relational Databases.
RDB2G-Bench is a toolkit for benchmarking graph-based analysis and prediction tasks by converting relational database data into graphs.
Our code is available at GitHub.
Overview
RDB2G-Bench provides comprehensive performance evaluation data for graph neural network models applied to relational database tasks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaistdata/RDB2G-Bench.network-packet-flow-header-payloadEach row contains the information of a network packet and its label. The format is given below:
british-museum-rdf-as-csv-2014This repository contains data that was released under the BM license issued in 2014.
acc_rd_s1-gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.rde-72-dataset
RDE-72: Rotating Detonation Engine Spatiotemporal Dataset
Overview
RDE-72 is the first open dataset of full 2D spatial fields from Rotating Detonation Engine (RDE) CFD simulations. It combines 12 high-fidelity OpenFOAM reactingFoam cases with 60 synthetic cases generated via a Conditional Variational Autoencoder (CVAE), enabling neural surrogate modeling and operating envelope mapping.
Key feature: Unlike prior RDE datasets that provide only bulk statistics (mean pressure… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/rde-72-dataset.sinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/genz-slang-pairs-1k.packet-tag-explanationThis dataset contains the packet information and the tags and their corresponding explanation. For more information, visit here.
rdf-query-based-summarizationall_RD_datasets
RD Dataset With References
This dataset contains Arabic terms and their definitions.The data was extracted and combined from the following sources:
https://huggingface.co/datasets/Basma2423/Arabic-Terminologies-and-Definitions
https://data.mendeley.com/datasets/gxr3j4tdk5/3
https://huggingface.co/datasets/MohamedRashad/arabic-roots
https://arai.ksaa.gov.sa/sharedTask2024/
Each entry consists of:
word
definition
Graph2Text_rdf_typetribu-rdc-v.0.1rdf-triplebusinessrdf-triple-based-summarizationrdf-summarization-btestsynthetic_equipment_utilization_datasetrdf-summarization-drestaurant-reviews-timelines
🍽️ Restaurant Reviews with Timelines (Synthetic GPT-4.1 Nano)
Dataset Repository: Programmer-RD-AI/restaurant-reviews-timelines-gpt4nano
📚 Overview
This synthetic dataset comprises over 10,000 restaurant reviews, meticulously generated using OpenAI's GPT-4.1 Nano model. Each review is contextualized within a specific phase of a restaurant's lifecycle, such as:
Opening Hype (Year 1)
Needs Overhaul (Year 4)
New and Improving (Year 2)
Rise and Fall (Year 3)
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/restaurant-reviews-timelines.embedded_faqs_medicarerdf-summarizationrdf-summarizaiton-c
