datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dlam-ts-project-data-2026
operations_forecasting_2026
Multivariate hourly forecasting for anonymized operations units.
Target
Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit.
Forecast Contract
Frequency: h
Series: 96
Timesteps per series: 4992
Target column: target
Training history length used by the baseline templates: 168
Rollout block length: 24
Required prediction horizon: validation: 336, test: 336… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/dlam-ts-project-data-2026.DLD_Transactionsplue
PLUE
Repository: https://github.com/ju-resplande/PLUE
Paper:
Leaderboard:
Point of Contact:
Portuguese translation of the GLUE benchmark, SNLI, and Scitail using OPUS-MT model and Google Cloud Translation.
The language data in PLUE is Brazilian Portuguese (BCP-47 pt-BR)
Citation Information
@misc{Gomes2020,
author = {GOMES, J. R. S.},
title = {PLUE: Portuguese Language Understanding Evaluation},
year = {2020},
publisher = {GitHub},
journal = {GitHub… See the full description on the dataset page: https://huggingface.co/datasets/dlb/plue.dl_dataset_1meteora-dlmm-historical-data
Meteora DLMM Historical Data
Decoded Solana mainnet instructions and events from Meteora DLMM (Dynamic Liquidity Market Maker), a concentrated-liquidity DEX where liquidity sits in discrete price bins and the fee rate rises with volatility.
74 tables, 59,575 rows, one row per decoded instruction or event. Program ID LBUZKhRxPF3XUpBCjp4YzTKgLccjZhTSDM9YuVaPwxo.
This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet.
Read… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/meteora-dlmm-historical-data.UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information about… See the full description on the dataset page: https://huggingface.co/datasets/alibustami/UM-DLP-Public-Benchmarking-Dataset.tmp-DL_Keyslicesdubai-real-estate-dld
Dubai Real Estate — DLD open data, cleaned and aggregated
Registered property sales, rent contracts and the project registry of the
Dubai Land Department (DLD), cleaned and aggregated by
Dubai Data — the same numbers that are published on the portal,
exported from its nightly build.
Data through: 2026-09-17 (DLD publishes with a lag of a few working days)
Updated: weekly from the portal's nightly build
Methodology: https://datadubai.ae/methodology/ — outlier trimming, minimum… See the full description on the dataset page: https://huggingface.co/datasets/datadubai/dubai-real-estate-dld.DLKcaten-fr-nyu-dl-course-corpus
Dataset information
Dataset from the French translation by Loïck Bourdois of the course by Yann Le Cun and Alfredo Canziani from the NYU.More than 3000 parallel data were created. The whole corpus has been manually checked to make sure of the good alignment of the data.
Note that the English data comes from several different people (about 190, see the acknowledgement section below).This has an impact on the homogeneity of the texts (some write in the past tense, others in the… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/en-fr-nyu-dl-course-corpus.user_feedbackcc-tool-merchant-training-datanormative_evaluation_llms_everyday_dilemmasmentalreddit
MentalReddit
This dataset, dlb/mentalreddit, was created by the DeepLearningBrasil team for the pre-training of their MentalBERTa model.
This model secured the first position in the DepSign-LT-EDI@RANLP-2023 shared task, which focused on classifying social media texts into three levels of depression.
Dataset Description
The MentalReddit dataset is a large collection of English-language comments sourced from Reddit. The data was specifically curated to provide a rich… See the full description on the dataset page: https://huggingface.co/datasets/dlb/mentalreddit.liveideabench-DLC-250127This is a supplementary dataset to the https://huggingface.co/datasets/6cf/liveideabench, covering additional model test data from December 23, 2024 to January 27, 2025. It includes test results for the following models:
deepseek-v3
deepseek-r1
phi-4
minimax-01
Opus
mistral-nemo
energy-forecasting-filesDLDDataset-DLG4_RAT
Description
This dataset contains signle site mutation of protein DLG4_RAT amino acid sequence and the correspond mutation effect score from a deep mutation scanning experiment.
Protein Format: AA sequence
Splits
traing: 1325
valid: 159
test: 176
Related paper
The dataset is from Deep generative models of genetic variation capture the effects of mutations.
Label
Label means fitness score of each mutant amino acid sequence based on a deep mutation… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-DLG4_RAT.UM-DLP-Public-Benchmarking-Dataset
UM DLP Public Benchmarking Dataset
Description
The UM DLP Public Benchmarking Dataset is a publicly available collection designed specifically to stress test Data Loss Prevention (DLP) systems, helping identify detection gaps, false positives, and false negatives for ongoing improvement.
This benchmark dataset contains 1,343 manually validated records across six major categories relevant to financial and sensitive data risks:
Financial Data (Account information… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/UM-DLP-Public-Benchmarking-Dataset.dlgenai-nppe2-dataset
NPPE-2 Protein Secondary Structure Dataset
This dataset is used for the NPPE-2 project on
Protein Secondary Structure Prediction.
Source
The original dataset is provided via Kaggle:
https://www.kaggle.com/competitions/sep-25-dl-gen-ai-nppe-2
Description
The dataset contains protein amino-acid sequences and their
corresponding secondary structure annotations:
Q8 (sst8): 8-state labels
Q3 (sst3): 3-state labels derived from Q8
Files include:
train.csv
test.csv… See the full description on the dataset page: https://huggingface.co/datasets/24f1000743/dlgenai-nppe2-dataset.aws-datadlgenai-nppe-datasetdlp-roster-3ef8ab13
Training Roster - DLP-Q3-2026-3ef8ab13
Published roster for the Q3 2026 Data Literacy Program.
Files:
roster.csv: employees to consider for enrollment.
policy.md: assignment metadata policy and business rules.
dlgenai-nppedlp-roster-3efb195d
Training Roster - DLP-Q3-2026-3efb195d
Published roster for the Q3 2026 Data Literacy Program.
Files:
roster.csv: employees to consider for enrollment.
policy.md: assignment metadata policy and business rules.
dl_datasetdlgenai_ProteinPredictionlogsdlam-ts-project-data-2026
operations_forecasting_2026
Multivariate hourly forecasting for anonymized operations units.
Target
Predict the future hourly operational load index for each series_id. Higher values indicate more operational pressure in that unit.
Forecast Contract
Frequency: h
Series: 96
Timesteps per series: 4992
Target column: target
Training history length used by the baseline templates: 168
Rollout block length: 24
Required prediction horizon: validation: 336… See the full description on the dataset page: https://huggingface.co/datasets/SabbirSzl/dlam-ts-project-data-2026.spotify-tracks-dataset
Content
This is a dataset of Spotify tracks over a range of 125 different genres. Each track has some audio features associated with it. The data is in CSV format which is tabular and can be loaded quickly.
Usage
The dataset can be used for:
Building a Recommendation System based on some user input or preference
Classification purposes based on audio features and available genres
Any other application that you can think of. Feel free to discuss!… See the full description on the dataset page: https://huggingface.co/datasets/dlfkd0810/spotify-tracks-dataset.latin_author_dll_id
Latin Author Name to Digital Latin Library ID Concordance
This dataset contains variant names of authors of works in Latin and their corresponding identifier in the Digital Latin Library's Catalog.
The variant names were gathered from the Virtual International Authority File records for the authors. In instances where an author's name has few or no known variant name forms, pseudo-variant names were generated by "misspelling" the name using the following script:
from textblob import… See the full description on the dataset page: https://huggingface.co/datasets/sjhuskey/latin_author_dll_id.
