datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
platonic-all-experimentsahr999-dataset
AHR999 BTC Hoarding Index Dataset
Open, daily-updated AHR999 BTC hoarding index dataset, self-computed from
Binance BTCUSDT daily closes and published as CSV and JSON.
This Hugging Face repository is a mirror. The canonical dataset endpoints are:
Dashboard: https://ahr999.aix4u.com/
GitHub: https://github.com/RuochenLyu/ahr999-dataset
CSV endpoint: https://ahr999.aix4u.com/datasets/ahr999.csv
JSON endpoint: https://ahr999.aix4u.com/datasets/ahr999.json
Kaggle discovery mirror:… See the full description on the dataset page: https://huggingface.co/datasets/kshift/ahr999-dataset.ExpansionRx_OpenADMET_KSOL
ExpansionRx-OpenADMET KSOL
KSOL dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict KSOL of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
7298
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_KSOL.food17-flaggedASAP_OpenADMET_KSOL
ASAP-OpenADMET KSOL
ASAP_OpenADMET_KSOL dataset from the ASAP Discovery-OpenADMET Antiviral Drug Discovery Challenge [1] [2] [3]. It is intended to be used through
scikit-fingerprints library.
The task is to predict KSOL (kinetic solubility in uM) of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
477
Recommended split
time
Recommended metric
MAE
References
[1]
ASAP Discovery
"ASAP Discovery x OpenADMET Antiviral Drug… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ASAP_OpenADMET_KSOL.cyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/ks121/cyp-challenge-train-test.superkart-sales-datasetsteamreviewsemotional-supportTwitchStreamsKSL-LEX
Dataset Card for KSL-LEX
Dataset Details
Dataset Description
KSL-LEX is a publicly available lexical database of Korean Sign Language (KSL), inspired by the ASL-LEX project (American Sign Language Lexical Database). It is based on the Korean Sign Language Dictionary's everyday signs collection, which contains 3,669 signs, but has been expanded to include more detailed linguistic information, resulting in 6,463 entries (6,289 unique headwords).
The primary… See the full description on the dataset page: https://huggingface.co/datasets/AAILab/KSL-LEX.Indian-LawlexicographicDataSPARQL
A Dataset for Semantic Parsing of Natural Language Into SPARQL for Wikidata
Dataset Summary
The Natural Language to SPARQL (Lexicographic Data) dataset is designed for the task of semantic parsing, specifically converting natural language utterances into SPARQL queries targeting lexicographic data within the Wikidata Knowledge Graph. This dataset was created as part of a Master's thesis at the University of Zurich.
The dataset contains natural language utterances focused… See the full description on the dataset page: https://huggingface.co/datasets/ksennr/lexicographicDataSPARQL.werewolf-bluffers
Ultimate Werewolf Bluffing Structured Datasets (SD1 & SD2)
Dataset Summary
This repository contains two structured datasets, Structured Dataset 1 (SD1) and Structured Dataset 2 (SD2), designed for the study of bluffing behavior in Ultimate Werewolf, a social deduction game. The datasets provide player-level behavioral representations extracted from game transcripts shared by slhleosun and bolinlai.
Each record represents a single player during a single game round… See the full description on the dataset page: https://huggingface.co/datasets/KSBCode/werewolf-bluffers.seed_pest_agri_schemechinese_traditional_chengyudivorce_QA_kornew_simopenassistant-deepseek-coderCIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/ksay1o/CIC-IDS2017.tcm-qnasen_vtwitter_aspect_based_sentiment_analysissentiment_analysisdemo_classificationSummoner-StatisticsIM_catTA_RTMtitle_generationclassification
