datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
azerbaijan-court-data
Azerbaijan Court System Dataset
The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations.
Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale.
Quick Start
Load with Hugging Face datasets
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.bitcoin-historical-dataset
Historical Bitcoin Market, On-Chain, Mining and Macroeconomic Dataset
Dataset Summary
Comprehensive daily Bitcoin dataset from genesis block (2009-01-03) to 2026-09-07.
6,457 daily observations combining market data, on-chain metrics, mining stats, macro indicators, and 100+ derived features.
Historical Coverage
Period
Coverage
Reliability
2009-01-03 to 2010-07-17
No market price
Protocol only
2010-07-18 to 2013-04-27
Monthly… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-historical-dataset.historical-gold-prices
Historical Gold Price Dataset
A comprehensive, research-grade collection of gold price observations spanning from 1833 to 2026 — over 193 years of continuous data.
Dataset Summary
This dataset provides clean, structured, machine-readable historical gold data suitable for:
Machine learning and time-series forecasting
Quantitative finance research
Financial data analysis
Economic research
Gold price prediction
Correlation analysis
Financial education
Algorithmic… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/historical-gold-prices.global-asset-market-cap-intelligence
Global Asset Market Capitalization Intelligence Dataset (GAMCID)
What is GAMCID?
GAMCID is a research-grade, machine-learning-ready dataset capturing the historical evolution of global assets ranked by market capitalization. It covers public companies, precious metals, cryptocurrencies, ETFs, and commodities — sourced from CompaniesMarketCap.com.
Why does it exist?
Existing financial datasets typically focus on single asset classes (stocks OR crypto… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/global-asset-market-cap-intelligence.khamsat
Khamsat
A structured Arabic-language dataset collected from khamsat.com, the largest Arabic freelance microservice marketplace.
This dataset is an independent research, not affiliated with khamsat. All content rights reserved to hsoub.com.
Abstract
Pricing in freelance marketplaces is a persistent challenge for both sellers seeking to maximize earnings and buyers seeking fair value.
This dataset was constructed to support data-driven… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/khamsat.Sorting-Algorithms-Performance-Metrics
Sorting Algorithms Benchmark Dataset (Array Size: 1000)
A benchmark dataset comparing execution time, memory usage, and comparison counts of various sorting algorithms (Bubble Sort, Selection Sort, Insertion Sort, Merge Sort, Quick Sort, Heap Sort, Odd-Even Sort) on arrays of size 1000. Each algorithm was run 100 times with randomized inputs to ensure statistical significance.
Dataset Details
Columns
run: Trial number (1-100 per algorithm).
algorithm:… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/Sorting-Algorithms-Performance-Metrics.uniGame
UniGame Dataset
This dataset explores the relationship between gaming habits and academic performance among students. It includes various attributes such as age, educational level, CGPA, gaming habits, and other related factors.
Dataset Details
Dataset Description
This dataset aims to investigate how gaming affects the academic performance of students. It includes information on the respondents' demographics, gaming habits, and academic results.
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/ismail31415/uniGame.Quantum_Gate_Performance_Evaluation
🧪 Quantum Gate Performance Dataset
📘 Title:
Comprehensive Quantum Gate Performance Analysis: A Comparative Study of Noise and No-Noise Effects
📂 Dataset Description:
This repository contains benchmarking results for 13 quantum gates (e.g., H, CNOT, Toffoli) tested under noisy and noise-free conditions, based on 1000 simulation runs per gate configuration. Total 26000 rows and 13 columns.
📊 Features include:
Gate Type
Execution Time
Error Rate
Fidelity… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/Quantum_Gate_Performance_Evaluation.Spoken2TSL
Dataset Description
This dataset is a collection of Turkish to Turkish Sign Language (TSL) grammar version translations. The dataset is designed to facilitate research and development in the field of sign language translation and understanding. It contains pairs of sentences in Turkish and their corresponding TSL translations, which have been curated to follow the grammatical structure of TSL.
Data Collection
The data was collected primarily from the website… See the full description on the dataset page: https://huggingface.co/datasets/ismaildlml/Spoken2TSL.adaption-ethereum-tx-hashes
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-ethereum_tx_hashes
This dataset contains a collection of Ethereum transaction hashes represented as 64-character hexadecimal strings prefixed with '0x'. Each entry corresponds to a unique transaction identifier on the Ethereum blockchain. The data is formatted as plain text completions suitable for training models on blockchain address patterns.
Dataset size
There are… See the full description on the dataset page: https://huggingface.co/datasets/Ismail131/adaption-ethereum-tx-hashes.Dataset_de_pruebaExtraído de https://github.com/anthony-wang/BestPractices/tree/master/data.
Campos:
Formula (string)
T (float64): Temperatura (K)
CP (float64): Capacidad calorífica (J/mol K)
Dataset105blindspots-analysis
DeepSeek-R1-Distill-Qwen-1.5B Blind Spot Analysis
Model Tested
Model: tphage/DeepSeek-R1-Distill-Qwen-1.5BLink: https://huggingface.co/tphage/DeepSeek-R1-Distill-Qwen-1.5B
This is a 1.5B parameter distilled base reasoning model derived from Qwen architecture. It is not fine-tuned for a narrow downstream task.
Objective
The purpose of this dataset is to systematically document prediction errors ("blind spots") of the DeepSeek-R1-Distill-Qwen-1.5B model across… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/blindspots-analysis.FleetVisiontitanic_practica
