datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MuSP-Bench
MuSP-Bench
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Contents
data/questions.csv: all 490 questions, accepted answers, and the
response contract for each.
inputs/pdf/without_context/: one context-removed PDF per piece.
inputs/images/: rendered score-page images for every piece.
inputs/abc/: one ABC score per piece.
inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.ProtST-EnzymeCommissionProtST-BinaryLocalizationProtDescribebulgarian_cds_lg
Dataset Overview
A sentence-level corpus drawn from scanned Bulgarian children's text. Each row represents one segmented sentence, its tokenization, the source URL, and its token count.
Data Schema
Column
Type
Description
MainSentencised
string
The raw, sentence-segmented text (in Bulgarian).
TokenisedSent
list[string]
The sentence split into word-tokens (lowercased, stripped).
SourceLink
string (URL)
Origin of the sentence (e.g. a Chitanka book/text ZIP).… See the full description on the dataset page: https://huggingface.co/datasets/milamarcheva/bulgarian_cds_lg.ProtST-GeneOntology-CCsurvey-language-technologies
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
🏆 Best Social Impact Paper Award at ACL 2025
Dataset Summary
This dataset comprises responses from 1,000 individuals from diverse socioeconomic backgrounds, collected to study how socioeconomic status (SES) influences interaction with language technologies, particularly generative AI and large language models (LLMs). Participants shared demographic and socioeconomic data, as well as up to 10… See the full description on the dataset page: https://huggingface.co/datasets/MilaNLProc/survey-language-technologies.ProtST-AAVProtST-ThermostabilityProtST-FluorescenceProtST-BetaLactamasesubloc_templateProtST-StabilityProtST-GeneOntology-MFrussian_keywordsProtST-GeneOntology-BPrussian-indi-alternativefashion-custom-datafashion-updatedautotrain-xxhsn-fsq88fashionmma_trainingbiasly-data
Biasly: An Expert-Annotated Dataset for Subtle Misogyny Detection and Mitigation
This repository contains the dataset presented in the paper, Biasly: An Expert-Annotated Dataset for Subtle Misogyny Detection and Mitigation. The dataset is the first of its kind in that it provides detailed annotations for each misogynsitic instance, including sub-categories of misogyny, a continuous severity score, and potentially a rewritten version of the original datapoit with the misogyny reduced… See the full description on the dataset page: https://huggingface.co/datasets/mila-ai4h/biasly-data.face2profilebible
