datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apogee
Apogée: Crypto Market Candlestick Dataset
Overview
Most traders believe crypto is random, but deep learning scaling laws suggest otherwise. Apogée is an open-source research initiative exploring the scaling laws of crypto market forecasting. While financial markets are often assumed to be unpredictable, modern deep learning suggests that increasing data and compute could uncover measurable predictability.
Our goal is to quantify how many bits of future price movement… See the full description on the dataset page: https://huggingface.co/datasets/duonlabs/apogee.Ecommerce_textFinanceIQUCI_drugDuET-dataset
DuET TE measurements of 64 human cell types & benchmark datasets for TE and MRL prediction task
This dataset comprises TE datasets for 64 cell types and benchmark datasets for TE and MRL prediction task.
How to setup
First, clone the main repository to your work directory:
$ git clone https://github.com/mogam-ai/DuET.git
$ cd DuET
Then, download the dataset repository into DuET/datasets subdirectory.
# Needs huggingface-cli (pip install huggingface-cli)
$ huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/DuET-dataset.trail_ptw_dumpssaas-vendor-outage-duration-incident-resolution-time-mttr
How long do SaaS vendor outages last? Incident resolution time per vendor, rebuilt daily
As of 2026-09-22 12:29 UTC. For every incident a vendor posted on its own public status
page with BOTH an opened time and a resolved time, this dataset computes
duration_minutes = resolved_at - started_at
and rolls it up per vendor. It is derived, every day, from the incident table in
saas-vendor-status-pages-outages-incidents-daily; the two are rebuilt by the same job and cannot
disagree.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-outage-duration-incident-resolution-time-mttr.symptom-disease-datasetHVUE
⚠️ DEPRECATED — Use HVUE v2
This dataset (HVUE v1) is deprecated. The original benchmark contains exact and high-sequence-similarity overlap across supervised train/test partitions. Consequently, results obtained using these splits should not be interpreted as estimates of generalization to sequence-independent held-out viruses. HVUE v1 is retained for transparency and reproducibility of the original submission.
Use duttaprat/HVUE-v2 instead, which implements leakage-controlled… See the full description on the dataset page: https://huggingface.co/datasets/duttaprat/HVUE.HVUE-v2
HVUE v2: Human Virome Understanding Evaluation Benchmark
⚠️ This is version 2 of the HVUE benchmark. Version 1 (duttaprat/HVUE) contained train-test sequence overlap due to chunk-level random splitting before clustering. HVUE v2 corrects this with cluster-aware splitting and verified zero leakage. All v1 results should be considered superseded.
Overview
HVUE v2 is a rigorously constructed benchmark for evaluating DNA language models on three epidemiologically… See the full description on the dataset page: https://huggingface.co/datasets/duttaprat/HVUE-v2.l4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.hatecheck-dutch
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-dutch.dutch-colaDutch CoLA is a corpus of linguistic acceptability for Dutch: a dataset consisting of sentences in Dutch, each marked as either acceptable (class 1) or unacceptable (class 0). These sentences are collected from existing descriptions of Dutch grammar (see sources below) with expert-annotated acceptability labels.
Dutch CoLA is part of the group project by students of BA Information Science program at the University of Groningen. List of people involved (alphabetic order):
Abdi, Silvana
Brouwer… See the full description on the dataset page: https://huggingface.co/datasets/GroNLP/dutch-cola.trac-verify-cot
TracGPT labelled CoT grids
One row per 32-slice OASIS MRI grid, with Q1–Q6 / A1–A6 targets.
Images are one store-only zip per split (JPEGs are already compressed).
Unzip into {split}/images/ so CSV paths stay valid.
split
grids
individual slices
train
5577
all slices used in those grids
test
535
all slices used in those grids
Layout
train/
data_labelled.cot.csv
images.zip # unzip → images/grid + images/slices
test/… See the full description on the dataset page: https://huggingface.co/datasets/ducbanh/trac-verify-cot.dumpINJEXIS-Duplicate-Prompt-Injection-Dataset
SPML Chatbot Prompt Injection Dataset
Arxiv Paper
Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Duplicate-Prompt-Injection-Dataset.SQuAD_v1.1_Du_et_al_2017_formattedThe Du. et. al. 2017 paper provides the splits fo the SQuAD v1.1 dataset
in the json format. However, they are formatted differently than the original SQuAD dataset as posted
on huggingface.
So for ease of use in your own code, I'm providing a formatted version of the data with splits that were used in that paper.
The data_preprocessing.py script is also provided for convenience.
NOTE: The 'answers' column is stored as a string. This is because I exported the dataframe as .csv. So the… See the full description on the dataset page: https://huggingface.co/datasets/simpleParadox/SQuAD_v1.1_Du_et_al_2017_formatted.Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/Durgesh111/Cifer-Fraud-Detection-Dataset-AF.MTEB_leaks_and_duplications
LLE MTEB
This dataset lists the presence or absence of leaks and duplicate data in the datasets constituting the MTEB leaderboard (EN & FR).
For more information concerning the methodology and find out what the column names correspond to, please consult the following blog post.To keep things simple, we invite the reader to read the percentages indicated in the text_and_label_test_biased column, which correspond to the proportion of biased data in the test split of the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/MTEB_leaks_and_duplications.en_vi_advanced_sentences
Model description
This data I crawled from these site: https://prep.vn/blog/idiom-theo-chu-de-trong-tieng-anh/ and https://www.enewsdispatch.com/
Idiom site I carefully translation, however, the enews site I use google translate
dusun-dictionary-corpus
Dusun-English-Malay Dictionary and Corpus
A trilingual dataset containing Dusun, English, and Malay words, phrases, and sentences compiled for linguistic research, dictionary development, and machine translation.
This dataset is based on the Dusun language as spoken by the Dusun ethnic group of Sabah, Malaysia. While the Dusun dialect in the dataset shares approximately 99% similarity with standardized Kadazandusun, there may be minor differences in vocabulary, spelling, and… See the full description on the dataset page: https://huggingface.co/datasets/DusunDictionary/dusun-dictionary-corpus.dualchem
DualChem
DualChem is a benchmark of 600 expert-curated PhD-level chemistry questions (485 multiple choice, 115 free-form) across 7 subdomains, designed to measure whether LLMs provide dangerous uplift alongside their technical utility. Each item is annotated with an expert-written benign use case, an expert-written harmful use case, and 1–5 severity scores for both.
Dataset Configurations
benchmark_questions (600 items) — the benchmark items: prompt, response type… See the full description on the dataset page: https://huggingface.co/datasets/DualChem-author/dualchem.Dutch-GOV-Law-wetten.overheid.nl
Dutch GOV Laws
This dataset is created by scraping https://wetten.overheid.nl, I used the Sitemap to get all possible URLS.
It possible some URLS are missing, around 1% gave a 404 or 405 error.
The reason for creating this dataset is I couldn't find any other existing dataset with this data.
So here is this dataset, Enjoy!
Please note this dataset is not complety checked or cleaned, this was a short research project for myself.
full-dubai-pulselalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.chatgpt-dutch-simplification
Dataset Card for ChatGPT Dutch Simplification
Dataset Summary
Created in light of a master thesis by Charlotte Van de Velde as part of the Master of Science in Artificial Intelligence at KU Leuven.
Charlotte is supervised by Vincent Vandeghinste and Bram Vanroy.
The dataset contains Dutch source sentences and aligned simplified sentences, generated with ChatGPT. All splits combined, the dataset
consists of 1267 entries.
Charlotte used gpt-3.5-turbo with the following… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification.Dutch-Government-Data-for-Bias-detectiondedeucebench-results
DedeuceBench Results Repository
This dataset stores submitted runs and an aggregated leaderboard for DedeuceBench. A run consists of a raw results.jsonl file produced by the CLI and a one-line CSV produced by the aggregator. The top-level leaderboard.csv is the append-only global table.
File Layout
leaderboard.csv — global leaderboard table with one row per (model, subset) entry.
runs/YYYY-MM-DD/<route>.<subset>/ — per-run artifacts:… See the full description on the dataset page: https://huggingface.co/datasets/comfortably-dumb/dedeucebench-results.Judgement-De-Identification-Result법원 판결문 비식별 모델의 성능 결과입니다.
SOTA 급 LLM을 활용한 법원 판결문 개인정보 비식별 성능(Few-shot 성능)
모델
정확도
재현율
F1 점수
GPT-4o(2024-08-06)
97.82
99.66
98.74
Qwen2.5-Max
96.46
95.83
96.14
DeepSeek-V3
98.73
98.92
98.81
Gemini-2.0-Flash
99.38
95.78
97.55
7~8B급 sLLM의 파인튜닝 전후 법원 판결문 개인정보 비식별 성능
모델
파인튜닝 전
파인튜닝 후
정확도
재현율
F1 점수
정확도
재현율
F1 점수
EXAONE-3.5-7.8B-Instruct
68.26
67.89
68.08
98.59
94.4896.49
Ministral-8B-Instruct-2410
35.6
4.33
7.72
99.07
98.32
98.70… See the full description on the dataset page: https://huggingface.co/datasets/ducut91/Judgement-De-Identification-Result.Pinyin-Hanzi
汉字语句序列与汉语拼音序列数据集
汉字语句序列与汉语拼音序列数据集,包含多领域文本,可用于训练汉字-汉语拼音互转模型。
