CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01duonlabs /apogee Apogée: Crypto Market Candlestick Dataset Overview Most traders believe crypto is random, but deep learning scaling laws suggest otherwise. Apogée is an open-source research initiative exploring the scaling laws of crypto market forecasting. While financial markets are often assumed to be unpredictable, modern deep learning suggests that increasing data and compute could uncover measurable predictability. Our goal is to quantify how many bits of future price movement… See the full description on the dataset page: https://huggingface.co/datasets/duonlabs/apogee.tabulartime-series-forecastingn<1K1 likes1.3k downloads1y agoHugging Face02Duyacquy /Ecommerce_texttext10K<n<100K0 likes795 downloads1y agoHugging Face03Duxiaoman-DI /FinanceIQtext1K<n<10K49 likes736 downloads3y agoHugging Face04Duyacquy /UCI_drugtabular10K<n<100K0 likes472 downloads1y agoHugging Face05mogam-ai /DuET-dataset DuET TE measurements of 64 human cell types & benchmark datasets for TE and MRL prediction task This dataset comprises TE datasets for 64 cell types and benchmark datasets for TE and MRL prediction task. How to setup First, clone the main repository to your work directory: $ git clone https://github.com/mogam-ai/DuET.git $ cd DuET Then, download the dataset repository into DuET/datasets subdirectory. # Needs huggingface-cli (pip install huggingface-cli) $ huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/DuET-dataset.tabular1M<n<10M0 likes377 downloads9mo agoHugging Face06konkanello /trail_ptw_dumpstabular100M<n<1B0 likes293 downloads2mo agoHugging Face07APProjects /saas-vendor-outage-duration-incident-resolution-time-mttr How long do SaaS vendor outages last? Incident resolution time per vendor, rebuilt daily As of 2026-09-22 12:29 UTC. For every incident a vendor posted on its own public status page with BOTH an opened time and a resolved time, this dataset computes duration_minutes = resolved_at - started_at and rolls it up per vendor. It is derived, every day, from the incident table in saas-vendor-status-pages-outages-incidents-daily; the two are rebuilt by the same job and cannot disagree.… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-outage-duration-incident-resolution-time-mttr.tabulartabular-regression10K<n<100K0 likes293 downloads55m agoHugging Face08dux-tecblic /symptom-disease-datasettexttext-classification1K<n<10K16 likes208 downloads3y agoHugging Face09duttaprat /HVUE ⚠️ DEPRECATED — Use HVUE v2 This dataset (HVUE v1) is deprecated. The original benchmark contains exact and high-sequence-similarity overlap across supervised train/test partitions. Consequently, results obtained using these splits should not be interpreted as estimates of generalization to sequence-independent held-out viruses. HVUE v1 is retained for transparency and reproducibility of the original submission. Use duttaprat/HVUE-v2 instead, which implements leakage-controlled… See the full description on the dataset page: https://huggingface.co/datasets/duttaprat/HVUE.text100K<n<1M0 likes198 downloads28d agoHugging Face10duttaprat /HVUE-v2 HVUE v2: Human Virome Understanding Evaluation Benchmark ⚠️ This is version 2 of the HVUE benchmark. Version 1 (duttaprat/HVUE) contained train-test sequence overlap due to chunk-level random splitting before clustering. HVUE v2 corrects this with cluster-aware splitting and verified zero leakage. All v1 results should be considered superseded. Overview HVUE v2 is a rigorously constructed benchmark for evaluating DNA language models on three epidemiologically… See the full description on the dataset page: https://huggingface.co/datasets/duttaprat/HVUE-v2.texttext-classification1M<n<10M0 likes174 downloads20d agoHugging Face11mayank-dubey-ai /l4-gpu-llm-benchmark-leaderboard 🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB) An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU. 📊 Executive Summary & Key Takeaways ⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.tabulartext-generationn<1K0 likes173 downloads1mo agoHugging Face12Paul /hatecheck-dutch Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-dutch.tabulartext-classification1K<n<10K2 likes114 downloads4y agoHugging Face13GroNLP /dutch-colaDutch CoLA is a corpus of linguistic acceptability for Dutch: a dataset consisting of sentences in Dutch, each marked as either acceptable (class 1) or unacceptable (class 0). These sentences are collected from existing descriptions of Dutch grammar (see sources below) with expert-annotated acceptability labels. Dutch CoLA is part of the group project by students of BA Information Science program at the University of Groningen. List of people involved (alphabetic order): Abdi, Silvana Brouwer… See the full description on the dataset page: https://huggingface.co/datasets/GroNLP/dutch-cola.tabulartext-classification10K<n<100K6 likes99 downloads2y agoHugging Face14ducbanh /trac-verify-cot TracGPT labelled CoT grids One row per 32-slice OASIS MRI grid, with Q1–Q6 / A1–A6 targets. Images are one store-only zip per split (JPEGs are already compressed). Unzip into {split}/images/ so CSV paths stay valid. split grids individual slices train 5577 all slices used in those grids test 535 all slices used in those grids Layout train/ data_labelled.cot.csv images.zip # unzip → images/grid + images/slices test/… See the full description on the dataset page: https://huggingface.co/datasets/ducbanh/trac-verify-cot.imagevisual-question-answering1K<n<10K0 likes90 downloads1mo agoHugging Face15deeponh /dumptabular1M<n<10M0 likes88 downloads11mo agoHugging Face16Sahildhonde-9 /INJEXIS-Duplicate-Prompt-Injection-Dataset SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sahildhonde-9/INJEXIS-Duplicate-Prompt-Injection-Dataset.tabulartext-classification10K<n<100K0 likes87 downloads24d agoHugging Face17simpleParadox /SQuAD_v1.1_Du_et_al_2017_formattedThe Du. et. al. 2017 paper provides the splits fo the SQuAD v1.1 dataset in the json format. However, they are formatted differently than the original SQuAD dataset as posted on huggingface. So for ease of use in your own code, I'm providing a formatted version of the data with splits that were used in that paper. The data_preprocessing.py script is also provided for convenience. NOTE: The 'answers' column is stored as a string. This is because I exported the dataframe as .csv. So the… See the full description on the dataset page: https://huggingface.co/datasets/simpleParadox/SQuAD_v1.1_Du_et_al_2017_formatted.text10K<n<100K1 likes84 downloads2y agoHugging Face18Durgesh111 /Cifer-Fraud-Detection-Dataset-AF 📊 Cifer Fraud Detection Dataset 🧠 Overview The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection. This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/Durgesh111/Cifer-Fraud-Detection-Dataset-AF.tabulartabular-classification10M<n<100M0 likes80 downloads10mo agoHugging Face19lbourdois /MTEB_leaks_and_duplications LLE MTEB This dataset lists the presence or absence of leaks and duplicate data in the datasets constituting the MTEB leaderboard (EN & FR). For more information concerning the methodology and find out what the column names correspond to, please consult the following blog post.To keep things simple, we invite the reader to read the percentages indicated in the text_and_label_test_biased column, which correspond to the proportion of biased data in the test split of the dataset in… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/MTEB_leaks_and_duplications.textn<1K0 likes79 downloads2y agoHugging Face20duwuonline /en_vi_advanced_sentences Model description This data I crawled from these site: https://prep.vn/blog/idiom-theo-chu-de-trong-tieng-anh/ and https://www.enewsdispatch.com/ Idiom site I carefully translation, however, the enews site I use google translate texttranslationn<1K2 likes74 downloads3y agoHugging Face21DusunDictionary /dusun-dictionary-corpus Dusun-English-Malay Dictionary and Corpus A trilingual dataset containing Dusun, English, and Malay words, phrases, and sentences compiled for linguistic research, dictionary development, and machine translation. This dataset is based on the Dusun language as spoken by the Dusun ethnic group of Sabah, Malaysia. While the Dusun dialect in the dataset shares approximately 99% similarity with standardized Kadazandusun, there may be minor differences in vocabulary, spelling, and… See the full description on the dataset page: https://huggingface.co/datasets/DusunDictionary/dusun-dictionary-corpus.texttranslation10K<n<100K3 likes74 downloads1mo agoHugging Face22DualChem-author /dualchem DualChem DualChem is a benchmark of 600 expert-curated PhD-level chemistry questions (485 multiple choice, 115 free-form) across 7 subdomains, designed to measure whether LLMs provide dangerous uplift alongside their technical utility. Each item is annotated with an expert-written benign use case, an expert-written harmful use case, and 1–5 severity scores for both. Dataset Configurations benchmark_questions (600 items) — the benchmark items: prompt, response type… See the full description on the dataset page: https://huggingface.co/datasets/DualChem-author/dualchem.tabularquestion-answering1K<n<10K0 likes74 downloads5mo agoHugging Face23ethux /Dutch-GOV-Law-wetten.overheid.nl Dutch GOV Laws This dataset is created by scraping https://wetten.overheid.nl, I used the Sitemap to get all possible URLS. It possible some URLS are missing, around 1% gave a 404 or 405 error. The reason for creating this dataset is I couldn't find any other existing dataset with this data. So here is this dataset, Enjoy! Please note this dataset is not complety checked or cleaned, this was a short research project for myself. text10K<n<100K3 likes70 downloads1y agoHugging Face24viewit-ai /full-dubai-pulsetabular1M<n<10M0 likes68 downloads2y agoHugging Face25Salesforce /lalm-judge-validation-full-duplex LALM Judge Validation on Full-Duplex Voice Agents Companion dataset for the paper A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents. This repository contains the anonymised ratings, adversarial-defect recall tables, JSON schemas, and analysis scripts used to produce every headline number, table, and figure in that paper. Summary 209 rated stereo sessions: 152 full-duplex agent-client conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.tabularaudio-classification1K<n<10K2 likes62 downloads2mo agoHugging Face26BramVanroy /chatgpt-dutch-simplification Dataset Card for ChatGPT Dutch Simplification Dataset Summary Created in light of a master thesis by Charlotte Van de Velde as part of the Master of Science in Artificial Intelligence at KU Leuven. Charlotte is supervised by Vincent Vandeghinste and Bram Vanroy. The dataset contains Dutch source sentences and aligned simplified sentences, generated with ChatGPT. All splits combined, the dataset consists of 1267 entries. Charlotte used gpt-3.5-turbo with the following… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification.text1K<n<10K5 likes58 downloads3y agoHugging Face27milenamileentje /Dutch-Government-Data-for-Bias-detectiontabulartext-classification1K<n<10K2 likes58 downloads2y agoHugging Face28comfortably-dumb /dedeucebench-results DedeuceBench Results Repository This dataset stores submitted runs and an aggregated leaderboard for DedeuceBench. A run consists of a raw results.jsonl file produced by the CLI and a one-line CSV produced by the aggregator. The top-level leaderboard.csv is the append-only global table. File Layout leaderboard.csv — global leaderboard table with one row per (model, subset) entry. runs/YYYY-MM-DD/<route>.<subset>/ — per-run artifacts:… See the full description on the dataset page: https://huggingface.co/datasets/comfortably-dumb/dedeucebench-results.tabularn<1K0 likes56 downloads1y agoHugging Face29ducut91 /Judgement-De-Identification-Result법원 판결문 비식별 모델의 성능 결과입니다. SOTA 급 LLM을 활용한 법원 판결문 개인정보 비식별 성능(Few-shot 성능) 모델 정확도 재현율 F1 점수 GPT-4o(2024-08-06) 97.82 99.66 98.74 Qwen2.5-Max 96.46 95.83 96.14 DeepSeek-V3 98.73 98.92 98.81 Gemini-2.0-Flash 99.38 95.78 97.55 7~8B급 sLLM의 파인튜닝 전후 법원 판결문 개인정보 비식별 성능 모델 파인튜닝 전 파인튜닝 후 정확도 재현율 F1 점수 정확도 재현율 F1 점수 EXAONE-3.5-7.8B-Instruct 68.26 67.89 68.08 98.59 94.4896.49 Ministral-8B-Instruct-2410 35.6 4.33 7.72 99.07 98.32 98.70… See the full description on the dataset page: https://huggingface.co/datasets/ducut91/Judgement-De-Identification-Result.text1K<n<10K0 likes53 downloads2y agoHugging Face30Duyu /Pinyin-Hanzi 汉字语句序列与汉语拼音序列数据集 汉字语句序列与汉语拼音序列数据集,包含多领域文本,可用于训练汉字-汉语拼音互转模型。 text1M<n<10M1 likes52 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.