datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dlam-ts-project-data-2026
operations_forecasting_2026
Multivariate hourly forecasting for anonymized operations units.
Target
Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit.
Forecast Contract
Frequency: h
Series: 96
Timesteps per series: 4992
Target column: target
Training history length used by the baseline templates: 168
Rollout block length: 24
Required prediction horizon: validation: 336, test: 336… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/dlam-ts-project-data-2026.DLD_Transactionsplue
PLUE
Repository: https://github.com/ju-resplande/PLUE
Paper:
Leaderboard:
Point of Contact:
Portuguese translation of the GLUE benchmark, SNLI, and Scitail using OPUS-MT model and Google Cloud Translation.
The language data in PLUE is Brazilian Portuguese (BCP-47 pt-BR)
Citation Information
@misc{Gomes2020,
author = {GOMES, J. R. S.},
title = {PLUE: Portuguese Language Understanding Evaluation},
year = {2020},
publisher = {GitHub},
journal = {GitHub… See the full description on the dataset page: https://huggingface.co/datasets/dlb/plue.meteora-dlmm-historical-data
Meteora DLMM Historical Data
Decoded Solana mainnet instructions and events from Meteora DLMM (Dynamic Liquidity Market Maker), a concentrated-liquidity DEX where liquidity sits in discrete price bins and the fee rate rises with volatility.
74 tables, 59,575 rows, one row per decoded instruction or event. Program ID LBUZKhRxPF3XUpBCjp4YzTKgLccjZhTSDM9YuVaPwxo.
This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet.
Read… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/meteora-dlmm-historical-data.UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.tmp-DL_Keyslicesdubai-real-estate-dld
Dubai Real Estate — DLD open data, cleaned and aggregated
Registered property sales, rent contracts and the project registry of the
Dubai Land Department (DLD), cleaned and aggregated by
Dubai Data — the same numbers that are published on the portal,
exported from its nightly build.
Data through: 2026-09-17 (DLD publishes with a lag of a few working days)
Updated: weekly from the portal's nightly build
Methodology: https://datadubai.ae/methodology/ — outlier trimming, minimum… See the full description on the dataset page: https://huggingface.co/datasets/datadubai/dubai-real-estate-dld.user_feedbacknormative_evaluation_llms_everyday_dilemmasliveideabench-DLC-250127This is a supplementary dataset to the https://huggingface.co/datasets/6cf/liveideabench, covering additional model test data from December 23, 2024 to January 27, 2025. It includes test results for the following models:
deepseek-v3
deepseek-r1
phi-4
minimax-01
Opus
mistral-nemo
energy-forecasting-filesDLDUM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/UM-DLP-Public-Benchmarking-Dataset.dlgenai-nppe-datasetailearner-researchlab_iot-intrusion-detection-hybrid-ml-dl-dataset
IoT Intrusion Detection: Hybrid ML-DL Dataset
Hybrid ML-DL network traffic data for IoT intrusion and threat detection
Dataset Info
Source: Kaggle
Original Size: 35.78 MB
Kaggle Downloads: 3
Files: 1
Files
final_dataset.csv
Mirrored from Kaggle
InRanker-msmarcodlgenai_ProteinPredictionlogsdlam-ts-project-data-2026
operations_forecasting_2026
Multivariate hourly forecasting for anonymized operations units.
Target
Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit.
Forecast Contract
Frequency: h
Series: 96
Timesteps per series: 4992
Target column: target
Training history length used by the baseline templates: 168
Rollout block length: 24
Required prediction horizon: validation: 336… See the full description on the dataset page: https://huggingface.co/datasets/SabbirSzl/dlam-ts-project-data-2026.spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!… See the full description on the dataset page: https://huggingface.co/datasets/dlfkd0810/spotify-tracks-dataset.dlgenai-nppe-dataset
DL GenAI NPPE - Face Dataset
This dataset mirrors the Kaggle dataset used in the DL GenAI NPPE project:
Source: /kaggle/input/sep-25-dl-gen-ai-nppe-1/face_datasetContents:
train/ and test/ image folders
train.csv and test.csv metadata files
Used for age (regression) and gender (classification) prediction models.
Linked Space: Age & Gender Prediction Space
DLiPathDeepLearningCommonFactor_DLvsIPCAdlgenai-nppe1dlgenai-nppe-logs
NPPE Protein Secondary Structure – Training Logs
This dataset contains training logs generated during model training
for the Kaggle competition SEP-25-DL-GEN-AI-NPPE-2.
Contents
training_logs.csv
epoch: training epoch
total_loss: average training loss per epoch
Models Logged
Bidirectional LSTM (BiLSTM)
Purpose
This dataset is used for:
Experiment tracking
Metric visualization
The corresponding Hugging Face Space consumes the trained model
and… See the full description on the dataset page: https://huggingface.co/datasets/Mrinal2/dlgenai-nppe-logs.datas_nettoyees_model_FRdlgenai_nppe1-datasetdlgenai-nppe-datasetdlgenai-nppe2dlgenai-nppe-logsdlgenai-nppe-dataset
DLGenAI – NPPE-2 Metrics Dataset
This dataset contains training and validation metrics for NPPE-2 models.
Models Included
BiRNN (from scratch)
CNN + BiLSTM hierarchical (Q3 → Q8)
Columns
model
architecture
best_val_loss
epochs
Purpose
Used for TrackIO visualization, experiment comparison, and course submission.
