datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fragrance-database
FragDB v5.16 — Fragrance Database (Multilingual Sample)
The most comprehensive structured fragrance database available. This is a free sample of FragDB: 140,230 perfumes, 23 languages — 10-row CSV samples at root.
Full dataset: fragdb.net.
What's New in v5.16
Data updated from v5.15 → v5.16 (snapshot 2026-09-19):
Fragrances: 139,501 → 140,230 (+729)
Brands: 8,272 → 8,316 (+44)
Perfumers: 3,116 → 3,126 (+10)
Notes: 2,596 → 2,606 rows in notes.csv (+10)
Companion… See the full description on the dataset page: https://huggingface.co/datasets/FragDBnet/fragrance-database.Databaseslot-database
Slot Machine Database — 5,669 Slots from 58 Providers
Comprehensive dataset of online slot machine metadata covering 58 game providers. Each record includes RTP, volatility, max win multiplier, grid layout, mechanics, themes, features, and bet ranges.
Homepage: slot.report
API: slot.report/api/
Dataset Description
This dataset contains structured metadata for 5,669 online slot machines, making it one of the largest publicly available slot game databases.… See the full description on the dataset page: https://huggingface.co/datasets/slreport/slot-database.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.MRC-psycholinguistic-database
MRC Psycholinguistic Database
This is the complete MRC psycholinguistic database as found on https://websites.psychology.uwa.edu.au/school/mrcdatabase/uwa_mrc.htm.
Usage
This dataset is ideal for training and evaluating machine learning models for English word concreteness.
Acknowledgments
We extend our heartfelt gratitude to all the authors of the original dataset.
License
This dataset is made available under the MIT license.
lyrics-database
Genius Lyrics
This dataset is a processed, lightweight subset of the original brunokreiner/genius-lyrics dataset. It has been filtered to focus exclusively on English songs with valid artist data, making it optimized for NLP tasks involving English songwriting, lyric generation, or genre classification.
All original credit goes to Bruno Kreiner for scraping and compiling the original data. This version is simply a filtered downsize for ease of use.
Processing & Filtering… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/lyrics-database.pima-indians-diabetes-database
Pima Indians Diabetes Dataset Split
This directory contains split datasets of Pima Indians Diabetes Database.
For each splits, we have
Mock data: The mock data is a smaller dataset (10 rows for both train and test) that is used to test the model and data processing code.
Private data: Each private data contains 123-125 rows for training, and 32-33 rows for testing.
horse-racing-database-fr
Huggy Database — French Horse Races (1996-2026)
Official site: https://huggydatabase.com/ · Interactive Query Builder: https://huggydatabase.com/#query
356,371 races · 30 years (1996 → 2026) · 40+ variables per runner · SQL-ready.
A structured French horse-racing archive covering harness trot, mounted trot, flat and jumps, with PMU/PMH odds spread, shoeing, blinkers, going, autostart and pedigree.
🇬🇧 English
What is in this repository
This… See the full description on the dataset page: https://huggingface.co/datasets/huggy-engine/horse-racing-database-fr.pima-indians-diabetes-database-partitions
Pima Indians Diabetes Dataset Split
This directory contains a dataset split for Pima Indians Diabetes Database.
Mock Data
The mock data is a smaller dataset (10 rows) that is used to test the model components.
Private Data
The private data is the remaining data that is used to train the model.
Phone_Timings_Database
📖 TajweedAI: Quranic Phoneme Timing Benchmark (Phases 1, 2 & 3)
📌 Project Overview
TajweedAI evaluates Quranic recitation accuracy by analyzing both pronunciation (phoneme classification) and timing (rule duration evaluation).
This benchmark provides empirical, tempo-normalized duration boundaries for all 70 Quranic phonemes derived from forced alignments (MFA trained on Quranic audio) across 7 master reference reciters:
Sheikh Mahmoud Khalil Al-Husary (Gold… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTamertechno1/Phone_Timings_Database.FunctionKey_Bag_Events_Database
FunctionKey_Bag_Events_Database
tags: event planning, ticket sales forecasting, attendee behavior
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The dataset titled 'FunctionKey_Bag_Events_Database' is a curated collection of information relevant to music festivals and events, focusing on the commercial and operational aspects of the industry. It includes contractual agreements, ticket sales forecasts, attendee behavior patterns… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FunctionKey_Bag_Events_Database.world-company-database
World Company Database — 100K Sample
A sample of 100,000 companies from the S.C.A.L.A. Score database, which contains 244M+ companies across 50+ countries.
Dataset Description
This dataset provides structured company information including:
Company name, location (country, city, province, postal code)
Legal form and registration status
NACE industry sector codes and descriptions
Financial data (revenue, net income, total assets, equity)
Employee count and founding… See the full description on the dataset page: https://huggingface.co/datasets/alf1990mi/world-company-database.master-database-tourismriya_databaseplant-database-2
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jibrand/plant-database-2.ai-tools-database-25k-aitoolbuzzai-models-database
Convly AI Models Database
A continuously updated, hand-verified dataset of 30+ AI language models — specs, licenses, API pricing (USD per 1M tokens), and local-hardware (VRAM) requirements.
Maintained by Convly.ai · Live interactive version: https://convly.ai/models/
Fields
name, slug, convly_url, developer, model_type, modality, parameters, context_window, max_output, license, open_weights, release_date, input_price (USD/1M tokens), output_price (USD/1M tokens)… See the full description on the dataset page: https://huggingface.co/datasets/sakd99/ai-models-database.plant-variety-database
Plant Variety Database
An open dataset that joins cultivar-level seed-catalog data with USDA hardiness zones and per-zone monthly planting calendars — 1,972 varieties × 13 zones × 12 months, fully sourced, CC BY 4.0.
The hero rows aren't the 1,972 varieties (USDA PLANTS already has ~98K species). They're the joins:
20,728 variety × zone planting-calendar entries (indoor sow / transplant / direct sow / harvest windows)
21,880 companion-plant pairings with relationship and reason
2… See the full description on the dataset page: https://huggingface.co/datasets/WindRiverGreens/plant-variety-database.clinical-quad-data-cut-query-backlog-database-lock-decision-error-v0.1Clinical Quad Data Cut Query Backlog Database Lock Decision Error v0.1
Each row is a data cut snapshot.
Core quad
Data cut timingQuery backlogDatabase lock pressureDecision error risk
Target
label_wrong_call_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
Vector_Database_With_Open-SourceVN-Database
Visual Novels Dataset
This dataset contains a collection of visual novel scripts sourced from alpindale/visual-novels.
Converted into Hugging Face and, by extension. datasets-compatible format.
slices_database
A Database for SLICES Representation of Inorganic Crystal Structures
This database provides inorganic crystal structures encoded as SLICES strings (Simplified Line-Input Crystal-Encoding System), an invertible and invariant string-based representation for crystalline materials.
clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1Clinical Quad Data Cut Timing Database Lock Pressure Query Backlog CSR Narrative Drift v0.1
Each row is a trial monthly snapshot.
Core quad
Data cut timingDatabase lock pressureQuery backlogCSR narrative drift
Target
label_regulatory_issue_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1.ReefSense-Databasedream-databaseskills_network_vector_databaseFinetune_Phi3_model_on_DataBaseinfo_about_database_table_relationabhishekgupta56447_anime-offline-database
Anime Offline Database
"35,000+ anime titles with cross-referenced IDs
Dataset Info
Source: Kaggle
Original Size: 9.39 MB
Kaggle Downloads: 45
Files: 1
Files
anime_database.csv
Mirrored from Kaggle
Wallet-Database
Latest Wallet Address Database
Latest All Bitcoin Cash Address Wallet
Latest All Dogecoin Address Wallet
Latest All Litecoin Address Wallet
Latest All DASH Address Wallet
Latest All ZCASH Address Wallet
Latest All ETHEREUM Address Wallet
