datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llmops-database
The ZenML LLMOps Database
To learn more about ZenML and our open-source MLOps framework, visit
zenml.io.
Dataset Summary
The LLMOps Database is a comprehensive collection of over 500 real-world
generative AI implementations that showcases how organizations are successfully
deploying Large Language Models (LLMs) in production. The case studies have been
carefully curated to focus on technical depth and practical problem-solving,
with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.gpu-database
GPU Database
Comprehensive GPU specifications database with architecture, manufacturing, API support, performance details, and kernel development specs.
2,824 GPUs across NVIDIA, AMD, and Intel
Part of RightNow — AI-powered code editor for GPU kernel development
Data
Vendor
GPUs
File
NVIDIA
1,286
data/nvidia/all.json
AMD
1,292
data/amd/all.json
Intel
180
data/intel/all.json
All
2,824
data/all-gpus.json
Schema
Each GPU contains up to 55… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/gpu-database.ACCOUNTING_DATABASESlyrics-database
Genius Lyrics
This dataset is a processed, lightweight subset of the original brunokreiner/genius-lyrics dataset. It has been filtered to focus exclusively on English songs with valid artist data, making it optimized for NLP tasks involving English songwriting, lyric generation, or genre classification.
All original credit goes to Bruno Kreiner for scraping and compiling the original data. This version is simply a filtered downsize for ease of use.
Processing & Filtering… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/lyrics-database.database-query-logs-synthetic
Database Query Logs (synthetic)
3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL
Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text,
type, complexity, execution timing, and row-count metadata.
These queries are synthetic
The queries were programmatically generated, not captured from production systems.
They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.danbooru2023-metadata-database
Metadata Database for Danbooru2023
Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023
The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset.
This dataset contains a sqlite db file which have all the tags and posts metadata in it.
The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together)
The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.Synthesis-Properties-Database-for-NanomaterialsThis is a database for nanomaterial synthesis. By fine-tuning the qwen3-14b model, it can extract synthesis steps, synthesis routes, and corresponding product properties (primarily including size, morphology, absorption spectra, and emission spectra) from target paragraphs. This fine-tuning project can be found at https://github.com/ime1452/Synthesis-Properties-Database-for-Nanomaterials.
The database contains two files: dataset.json holds the raw, unprocessed data, while… See the full description on the dataset page: https://huggingface.co/datasets/Kai-gu/Synthesis-Properties-Database-for-Nanomaterials.slices_database
A Database for SLICES Representation of Inorganic Crystal Structures
This database provides inorganic crystal structures encoded as SLICES strings (Simplified Line-Input Crystal-Encoding System), an invertible and invariant string-based representation for crystalline materials.
DATABASE
📚 Bonolota_DB: Emotion-aware Bengali Storytelling Dataset
Bonolota_DB is a modular, offline-first dataset designed for Bengali storytelling engines.It contains emotion-tagged stories written in YAML format, with Markdown support and emoji cues for UI rendering and voice synthesis.
✨ Features
✅ YAML structure for easy parsing
✅ Markdown content for rich display
✅ Emotion tags (pride, hope, nostalgia) for filtering
✅ Emoji support for UI
✅ Bengali language (bn) with… See the full description on the dataset page: https://huggingface.co/datasets/Lavlu118557/DATABASE.database-sql-instruction-dataset
Database & SQL Instruction Dataset
High-quality instruction-response pairs covering PostgreSQL, advanced queries, indexing strategies, and database optimization.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Database topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/database-sql-instruction-dataset.ourdream-database
OurDream Character Database
Complete database of 20,884 AI characters from ourdream.ai, categorized as Women (12,492) and Trans (8,392).
Dataset Structure
database/all_characters.json — Full character database with 20,884 entries
Each entry includes: id, displayId, name, gender, style, age, likeCount, messageCount, tags, shortDescription, thumbUrl, category
Categories
Category
Count
Women
12,492
Trans
8,392
Total
20,884… See the full description on the dataset page: https://huggingface.co/datasets/lcuifer0/ourdream-database.fungi-rag-agent-sft-database
Fungi RAG Agent SFT Database
This dataset contains the supervised fine-tuning database used to train the
JacktheLander/smollm2-1.7b-fungi-rag-agent-distill-lora-gguf adapter.
It was built for the project-owned fungi RAG learning system:
JacktheLander/FunghiResearchAgent
The examples teach a small local language model to follow the system's agent contract:
use rag.search for evidence-backed mycology answers,
use safety.review for wild-mushroom edibility, field-identification… See the full description on the dataset page: https://huggingface.co/datasets/JacktheLander/fungi-rag-agent-sft-database.
