datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apogee
Apogée: Crypto Market Candlestick Dataset
Overview
Most traders believe crypto is random, but deep learning scaling laws suggest otherwise. Apogée is an open-source research initiative exploring the scaling laws of crypto market forecasting. While financial markets are often assumed to be unpredictable, modern deep learning suggests that increasing data and compute could uncover measurable predictability.
Our goal is to quantify how many bits of future price movement… See the full description on the dataset page: https://huggingface.co/datasets/duonlabs/apogee.UCI_drugDuET-dataset
DuET TE measurements of 64 human cell types & benchmark datasets for TE and MRL prediction task
This dataset comprises TE datasets for 64 cell types and benchmark datasets for TE and MRL prediction task.
How to setup
First, clone the main repository to your work directory:
$ git clone https://github.com/mogam-ai/DuET.git
$ cd DuET
Then, download the dataset repository into DuET/datasets subdirectory.
# Needs huggingface-cli (pip install huggingface-cli)
$ huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/DuET-dataset.saas-vendor-outage-duration-incident-resolution-time-mttr
How long do SaaS vendor outages last? Incident resolution time per vendor, rebuilt daily
As of 2026-09-23 12:29 UTC. For every incident a vendor posted on its own public status
page with BOTH an opened time and a resolved time, this dataset computes
duration_minutes = resolved_at - started_at
and rolls it up per vendor. It is derived, every day, from the incident table in
saas-vendor-status-pages-outages-incidents-daily; the two are rebuilt by the same job and cannot
disagree.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-outage-duration-incident-resolution-time-mttr.trail_ptw_dumpsl4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.hatecheck-dutch
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-dutch.dutch-colaDutch CoLA is a corpus of linguistic acceptability for Dutch: a dataset consisting of sentences in Dutch, each marked as either acceptable (class 1) or unacceptable (class 0). These sentences are collected from existing descriptions of Dutch grammar (see sources below) with expert-annotated acceptability labels.
Dutch CoLA is part of the group project by students of BA Information Science program at the University of Groningen. List of people involved (alphabetic order):
Abdi, Silvana
Brouwer… See the full description on the dataset page: https://huggingface.co/datasets/GroNLP/dutch-cola.dumpINJEXIS-Duplicate-Prompt-Injection-Dataset
SPML Chatbot Prompt Injection Dataset
Arxiv Paper
Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Duplicate-Prompt-Injection-Dataset.Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/Durgesh111/Cifer-Fraud-Detection-Dataset-AF.dualchem
DualChem
DualChem is a benchmark of 600 expert-curated PhD-level chemistry questions (485 multiple choice, 115 free-form) across 7 subdomains, designed to measure whether LLMs provide dangerous uplift alongside their technical utility. Each item is annotated with an expert-written benign use case, an expert-written harmful use case, and 1–5 severity scores for both.
Dataset Configurations
benchmark_questions (600 items) — the benchmark items: prompt, response type… See the full description on the dataset page: https://huggingface.co/datasets/DualChem-author/dualchem.full-dubai-pulselalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.dedeucebench-results
DedeuceBench Results Repository
This dataset stores submitted runs and an aggregated leaderboard for DedeuceBench. A run consists of a raw results.jsonl file produced by the CLI and a one-line CSV produced by the aggregator. The top-level leaderboard.csv is the append-only global table.
File Layout
leaderboard.csv — global leaderboard table with one row per (model, subset) entry.
runs/YYYY-MM-DD/<route>.<subset>/ — per-run artifacts:… See the full description on the dataset page: https://huggingface.co/datasets/comfortably-dumb/dedeucebench-results.stanislau-shumski-u-bitvakh-i-viaznitsakh-uspaminy-pra-1812-1848-gady-siargei-du
У бітвах і вязьніцах. Успаміны пра 1812–1848 гады
Metadata
Author: Станіслаў Шумскі
Title: У бітвах і вязьніцах. Успаміны пра 1812–1848 гады
Narrator: Сяргей Дубавец
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/stanislau-shumski-u-bitvakh-i-viaznitsakh-uspaminy-pra-1812-1848-gady-siargei-du.quora_duplicate_questionsAdapter by: Aisuko
Only for researching.
DubaiRealEstateSalesInsights
Dubai Real Estate – Exploratory Data Analysis (EDA)
Overview
This project presents an Exploratory Data Analysis (EDA) of residential real-estate listings in Dubai.The goal is to identify key factors influencing property prices using statistical exploration, data cleaning, and visual insights.
The full analysis was performed in Google Colab.The dataset (dubai_real_estate.csv) is hosted on HuggingFace.
Dataset
Source: Kaggle – Dubai Real Estate Listings
File:… See the full description on the dataset page: https://huggingface.co/datasets/pelegelraz/DubaiRealEstateSalesInsights.dungeon-dataset
Brogue Map Dataset
To clone this repo, use:
git clone https://huggingface.co/datasets/DolphinNie/dungeon-dataset
1. Data Explanation
This is the Map dataset from the open-sourced game Brogue. It contains 49,000 train dataset, 14,000 test dataset and 7,000 validation dataset.
Each map is stored in a .csv file. The map is a (32x32) array, which is the map size.
Each cell in the array is a int number ranged from 0 to 13, which represented 14 tiles.
"G_NONE": 0… See the full description on the dataset page: https://huggingface.co/datasets/DolphinNie/dungeon-dataset.Dutch-Government-Data-for-Bias-detectiondumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.dubai-real-estate-dld
Dubai Real Estate — DLD open data, cleaned and aggregated
Registered property sales, rent contracts and the project registry of the
Dubai Land Department (DLD), cleaned and aggregated by
Dubai Data — the same numbers that are published on the portal,
exported from its nightly build.
Data through: 2026-09-17 (DLD publishes with a lag of a few working days)
Updated: weekly from the portal's nightly build
Methodology: https://datadubai.ae/methodology/ — outlier trimming, minimum… See the full description on the dataset page: https://huggingface.co/datasets/datadubai/dubai-real-estate-dld.lich-am-duong-1900-2100
Lịch âm dương 1900 tới 2100
Vietnamese lunisolar calendar, 1900 to 2100
1. Mô tả · Description
Mỗi ngày dương lịch từ 1900 tới 2100 kèm ngày âm lịch, tháng nhuận, can chi ngày và năm, tiết khí đang trong và trực của ngày.
Every solar day from 1900 to 2100 with its lunar date, leap month flag, day and year stem-branch, the solar term in force and the day officer.
Số dòng · Rows: 73,414
Phiên bản · Version: 1.0.0 (2026-09-20)
Mã hoá · Encoding: UTF-8 không BOM… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/lich-am-duong-1900-2100.Duolingo-Spaced-Repetition-Datale-hoi-quy-doi-ngay-duong
Lễ hội Việt Nam quy đổi sang ngày dương, 2026 tới 2030
Vietnamese festivals mapped to solar dates, 2026 to 2030
1. Mô tả · Description
Ngày dương lịch thật của từng lễ hội trong năm năm, kèm thứ trong tuần. Lễ hội âm lịch rơi vào ngày dương khác nhau mỗi năm, và bảng này tính sẵn phần ấy.
The actual solar date of each festival across five years, with the weekday. Lunar festivals land on a different solar date every year, and this table works that out in advance.… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/le-hoi-quy-doi-ngay-duong.kws-recordings-soundailabel
kws-recordings-soundailabel
This repository serves as the KWS audio dataset repo for label_kws training.
Layout
audio/
manifests/
docs/
README.md
dataset_summary.json
Source of Truth
manifests/full_manifest.csv
manifests/kws_multiclass_manifest.csv
manifests/binary_emergency_manifest.csv
manifests/review_queue.csv
For compatibility, the same derived files are also available under data/derived/.
Current Summary
rows with local audio present: 796… See the full description on the dataset page: https://huggingface.co/datasets/dusen0528/kws-recordings-soundailabel.wiki-stance-enafrica-synth-energy-dust-soiling-maintenance-impact-all
Africa Synth Energy Dust Soiling Maintenance Impact All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-dust-soiling-maintenance-impact-all.pump_and_dump_datasetgamebench-2-results
DuelLab GameBench 2 results
This dataset repository is a public mirror of the immutable DuelLab GameBench 2 release gb2-20260807-e64919171e70. DuelLab is the canonical source; this mirror is generated from its release manifest rather than maintained as a separate leaderboard.
Release
Public release: gb2-20260807-e64919171e70
Generated: 2026-08-07T12:33:34Z
Benchmark version: 2-known8-r3
Games: 8
Models: 44
Reasoning variants: 111
Rated matches: 37483
Canonical… See the full description on the dataset page: https://huggingface.co/datasets/DuelLab/gamebench-2-results.
