datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fragrance-database
FragDB v5.16 — Fragrance Database (Multilingual Sample)
The most comprehensive structured fragrance database available. This is a free sample of FragDB: 140,230 perfumes, 23 languages — 10-row CSV samples at root.
Full dataset: fragdb.net.
What's New in v5.16
Data updated from v5.15 → v5.16 (snapshot 2026-09-19):
Fragrances: 139,501 → 140,230 (+729)
Brands: 8,272 → 8,316 (+44)
Perfumers: 3,116 → 3,126 (+10)
Notes: 2,596 → 2,606 rows in notes.csv (+10)
Companion… See the full description on the dataset page: https://huggingface.co/datasets/FragDBnet/fragrance-database.slot-database
Slot Machine Database — 5,669 Slots from 58 Providers
Comprehensive dataset of online slot machine metadata covering 58 game providers. Each record includes RTP, volatility, max win multiplier, grid layout, mechanics, themes, features, and bet ranges.
Homepage: slot.report
API: slot.report/api/
Dataset Description
This dataset contains structured metadata for 5,669 online slot machines, making it one of the largest publicly available slot game databases.… See the full description on the dataset page: https://huggingface.co/datasets/slreport/slot-database.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.MRC-psycholinguistic-database
MRC Psycholinguistic Database
This is the complete MRC psycholinguistic database as found on https://websites.psychology.uwa.edu.au/school/mrcdatabase/uwa_mrc.htm.
Usage
This dataset is ideal for training and evaluating machine learning models for English word concreteness.
Acknowledgments
We extend our heartfelt gratitude to all the authors of the original dataset.
License
This dataset is made available under the MIT license.
Phone_Timings_Database
📖 TajweedAI: Quranic Phoneme Timing Benchmark (Phases 1, 2 & 3)
📌 Project Overview
TajweedAI evaluates Quranic recitation accuracy by analyzing both pronunciation (phoneme classification) and timing (rule duration evaluation).
This benchmark provides empirical, tempo-normalized duration boundaries for all 70 Quranic phonemes derived from forced alignments (MFA trained on Quranic audio) across 7 master reference reciters:
Sheikh Mahmoud Khalil Al-Husary (Gold… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTamertechno1/Phone_Timings_Database.pima-indians-diabetes-database
Pima Indians Diabetes Dataset Split
This directory contains split datasets of Pima Indians Diabetes Database.
For each splits, we have
Mock data: The mock data is a smaller dataset (10 rows for both train and test) that is used to test the model and data processing code.
Private data: Each private data contains 123-125 rows for training, and 32-33 rows for testing.
horse-racing-database-fr
Huggy Database — French Horse Races (1996-2026)
Official site: https://huggydatabase.com/ · Interactive Query Builder: https://huggydatabase.com/#query
356,371 races · 30 years (1996 → 2026) · 40+ variables per runner · SQL-ready.
A structured French horse-racing archive covering harness trot, mounted trot, flat and jumps, with PMU/PMH odds spread, shoeing, blinkers, going, autostart and pedigree.
🇬🇧 English
What is in this repository
This… See the full description on the dataset page: https://huggingface.co/datasets/huggy-engine/horse-racing-database-fr.pima-indians-diabetes-database-partitions
Pima Indians Diabetes Dataset Split
This directory contains a dataset split for Pima Indians Diabetes Database.
Mock Data
The mock data is a smaller dataset (10 rows) that is used to test the model components.
Private Data
The private data is the remaining data that is used to train the model.
world-company-database
World Company Database — 100K Sample
A sample of 100,000 companies from the S.C.A.L.A. Score database, which contains 244M+ companies across 50+ countries.
Dataset Description
This dataset provides structured company information including:
Company name, location (country, city, province, postal code)
Legal form and registration status
NACE industry sector codes and descriptions
Financial data (revenue, net income, total assets, equity)
Employee count and founding… See the full description on the dataset page: https://huggingface.co/datasets/alf1990mi/world-company-database.master-database-tourismriya_databaseai-models-database
Convly AI Models Database
A continuously updated, hand-verified dataset of 30+ AI language models — specs, licenses, API pricing (USD per 1M tokens), and local-hardware (VRAM) requirements.
Maintained by Convly.ai · Live interactive version: https://convly.ai/models/
Fields
name, slug, convly_url, developer, model_type, modality, parameters, context_window, max_output, license, open_weights, release_date, input_price (USD/1M tokens), output_price (USD/1M tokens)… See the full description on the dataset page: https://huggingface.co/datasets/sakd99/ai-models-database.clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1Clinical Quad Data Cut Timing Database Lock Pressure Query Backlog CSR Narrative Drift v0.1
Each row is a trial monthly snapshot.
Core quad
Data cut timingDatabase lock pressureQuery backlogCSR narrative drift
Target
label_regulatory_issue_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1.plant-variety-database
Plant Variety Database
An open dataset that joins cultivar-level seed-catalog data with USDA hardiness zones and per-zone monthly planting calendars — 1,972 varieties × 13 zones × 12 months, fully sourced, CC BY 4.0.
The hero rows aren't the 1,972 varieties (USDA PLANTS already has ~98K species). They're the joins:
20,728 variety × zone planting-calendar entries (indoor sow / transplant / direct sow / harvest windows)
21,880 companion-plant pairings with relationship and reason
2… See the full description on the dataset page: https://huggingface.co/datasets/WindRiverGreens/plant-variety-database.clinical-quad-data-cut-query-backlog-database-lock-decision-error-v0.1Clinical Quad Data Cut Query Backlog Database Lock Decision Error v0.1
Each row is a data cut snapshot.
Core quad
Data cut timingQuery backlogDatabase lock pressureDecision error risk
Target
label_wrong_call_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
Vector_Database_With_Open-Sourceslices_database
A Database for SLICES Representation of Inorganic Crystal Structures
This database provides inorganic crystal structures encoded as SLICES strings (Simplified Line-Input Crystal-Encoding System), an invertible and invariant string-based representation for crystalline materials.
ReefSense-Databaseskills_network_vector_databaseabhishekgupta56447_anime-offline-database
Anime Offline Database
"35,000+ anime titles with cross-referenced IDs
Dataset Info
Source: Kaggle
Original Size: 9.39 MB
Kaggle Downloads: 45
Files: 1
Files
anime_database.csv
Mirrored from Kaggle
NPM-Weibull-DATABASE-v9_1
NPM-Weibull DATABASE v9_1
Benchmark database of Weibull (k, λ) fits for 12 transformer model entries spanning 7 architectural families (Pythia, OLMo-1/2, LLaMA-3, Mistral, Qwen2.5, Qwen3) — 70M–14B parameters, GeLU and SwiGLU activations, Pre-LN and QK-Norm placements.
📄 Paper: A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions (arXiv:2605.18898)
📦 Source repo (primary): github.com/tiexinding/NPM-Weibull-public — directory database_v9_1/
🔧 Python… See the full description on the dataset page: https://huggingface.co/datasets/TiexinDing/NPM-Weibull-DATABASE-v9_1.WD14-DataBaseglof-databasemy-fonts-databasesnowmobile-track-database
Snowmobile Track Specification Database
Snowmobile studs — the open reference dataset of snowmobile track specifications — 2,600+ model/track combinations across Ski-Doo, Polaris, Arctic Cat, Yamaha, and Lynx, model years 2001–2027. Includes track length, lug height, and internal ply count (1-ply vs 2-ply). Maintained by Fast-Trac Traction, American manufacturer of snowmobile studs since 1989. Compiled from 35+ years of fitment records. For recommended stud size and quantity for… See the full description on the dataset page: https://huggingface.co/datasets/Fasttractraction/snowmobile-track-database.us-doctors-database-sampleUS Doctors and Physicians Database (Sample)
This dataset contains a free sample of US healthcare providers, including their basic details.
📥 Download the Complete 1.4 Million Database:
If you are a healthcare recruiter or B2B marketer looking for the full verified list of 1.4 Million US Doctors (with direct emails, NPI numbers, and phone numbers), you can instantly download it from our official website:
United States Doctor Database
VECTOR_DATABASEmy-database-part1autotrain-data-groceries-database-l
