datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-databaseSkyWorld-OSM-Databaseosint-tool-databaseACCOUNTING_DATABASESdatabase-agent-runs
LibreDB Agent Benchmark
8,199 agent runs · 39 open-weight models served locally, plus one hosted model as a control · 6 task surfaces · 110,711 ledger events · 14,008 refused tool calls
This is the complete measurement record behind the paper What Stops a Small Language Model From
Driving a Database Agent. It is not a scored summary: it is every event the server wrote while the
runs happened, released so that every number in the paper can be recomputed, and disagreed with… See the full description on the dataset page: https://huggingface.co/datasets/libredb/database-agent-runs.Disease_Database
Citation
@misc{chen2024codinterpretablemedicalagent,
title={CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis},
author={Junying Chen and Chi Gui and Anningzhe Gao and Ke Ji and Xidong Wang and Xiang Wan and Benyou Wang},
year={2024},
eprint={2407.13301},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.13301},
}
Agentic-SLS-Database
Agentic-SLS-Database
Canonical graph dataset of Inova Mk1 SLS printer entities: jobs, print sessions, print profiles, and objects (STL geometry). Each entity is its own HF config; relationships are encoded as ID references between rows.
Domain-specific datasets (e.g. ppak10/Agentic-SLS-ASTM) reference rows here by ID and may embed frozen snapshots of the referenced state.
Configs
Config
Description
Script
Output
jobs
One row per .s4a print job, with… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-Database.chesssynapse-databasesai-skill-md-database-10kb
AI Skill MD Database 10KB
This Hugging Face Dataset repository contains 10 Codex-style AI skill folders.
Each SKILL.md file is at least 10 KB and includes detailed workflow guidance.
This is a dataset/database repository only. It is not a Hugging Face Space and contains no app runtime.
Layout
skills/<skill-name>/SKILL.md: skill instructions and trigger metadata
skills/<skill-name>/agents/openai.yaml: UI-facing metadata
skills_index.jsonl: searchable index with byte sizes… See the full description on the dataset page: https://huggingface.co/datasets/abersbail/ai-skill-md-database-10kb.openhands-commit-noise-databases
OpenHands Commit Noise Databases
This dataset contains commit-retrieval databases for 12 SWE-bench repositories at five noise ratios: 0%, 25%, 50%, 75%, and 100%.
Each archive expands to noise_NNN/<repository>/ directories containing:
commits.db: SQLite commit records
commits.faiss: normalized inner-product FAISS index
commits.meta.jsonl: FAISS row-to-commit metadata
commits.index_meta.json: embedding and index configuration
The 0% archive is an exact file-level copy of the… See the full description on the dataset page: https://huggingface.co/datasets/dengyixuan/openhands-commit-noise-databases.data_base_nerdatabase_for_ngugeuefinesse-benchmark-databaseWikipedia HF Dataset's modified partial version. no longer maintained.
The Finesse Benchmark Package, which was rely on current dataset, NO LONGER USE this specific dataset. instead, it can now use the existing datasets.Thank you.
plant-pet-toxicity-database
PlantFun Plant-Pet Toxicity Database
This dataset is exported from the GitHub Repository.
Official website: plantfun.app.
Snapshot
Generated at: 2026-02-13T02:09:58Z
Total markdown articles: 1313
Pet Toxicity Reports: 498
Misdiagnosis Case Studies: 408
Dynamic Care Protocols: 407
Detailed Encyclopedia: Explore all 1313 plants on PlantFun
Files
../articles.csv
../articles.jsonl
../manifest.json
Suggested usage
Plant toxicity and pet safety… See the full description on the dataset page: https://huggingface.co/datasets/LeafVibe/plant-pet-toxicity-database.databaseaitrSO-Python_QA-Database_and_SQL_class
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/RazinAleks/SO-Python_QA-Database_and_SQL_class.LMLM-databaseDATABASE
📚 Bonolota_DB: Emotion-aware Bengali Storytelling Dataset
Bonolota_DB is a modular, offline-first dataset designed for Bengali storytelling engines.It contains emotion-tagged stories written in YAML format, with Markdown support and emoji cues for UI rendering and voice synthesis.
✨ Features
✅ YAML structure for easy parsing
✅ Markdown content for rich display
✅ Emotion tags (pride, hope, nostalgia) for filtering
✅ Emoji support for UI
✅ Bengali language (bn) with… See the full description on the dataset page: https://huggingface.co/datasets/Lavlu118557/DATABASE.Word_in_Sentence_Database
WIS database
This database contains a question answer list about text
This database was built using my this workflow:
1- load a raw text file
2- split into paragraphs
3- split paragraphs into sentences
4- for each word, ask question about its position and answer with the position, then ask about the word length and answer with the actual length of the word
5- ask a question about the number of words in the sentence and answer it
6- build a json database using this.
To do this, I… See the full description on the dataset page: https://huggingface.co/datasets/ParisNeo/Word_in_Sentence_Database.retrieval-databasecastbench-leaderboard-databasescraped-trends-database-236dafDatabasefurniswipe-databaseDisease_Database
Citation
@misc{chen2024codinterpretablemedicalagent,
title={CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis},
author={Junying Chen and Chi Gui and Anningzhe Gao and Ke Ji and Xidong Wang and Xiang Wan and Benyou Wang},
year={2024},
eprint={2407.13301},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.13301},
}
databasemalicious-contract-database-neatdatabase-sql-instruction-dataset
Database & SQL Instruction Dataset
High-quality instruction-response pairs covering PostgreSQL, advanced queries, indexing strategies, and database optimization.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Database topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/database-sql-instruction-dataset.ourdream-database
OurDream Character Database
Complete database of 20,884 AI characters from ourdream.ai, categorized as Women (12,492) and Trans (8,392).
Dataset Structure
database/all_characters.json — Full character database with 20,884 entries
Each entry includes: id, displayId, name, gender, style, age, likeCount, messageCount, tags, shortDescription, thumbUrl, category
Categories
Category
Count
Women
12,492
Trans
8,392
Total
20,884… See the full description on the dataset page: https://huggingface.co/datasets/lcuifer0/ourdream-database.fungi-rag-agent-sft-database
Fungi RAG Agent SFT Database
This dataset contains the supervised fine-tuning database used to train the
JacktheLander/smollm2-1.7b-fungi-rag-agent-distill-lora-gguf adapter.
It was built for the project-owned fungi RAG learning system:
JacktheLander/FunghiResearchAgent
The examples teach a small local language model to follow the system's agent contract:
use rag.search for evidence-backed mycology answers,
use safety.review for wild-mushroom edibility, field-identification… See the full description on the dataset page: https://huggingface.co/datasets/JacktheLander/fungi-rag-agent-sft-database.
