datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MuSP-Bench
MuSP-Bench
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Contents
data/questions.csv: all 490 questions, accepted answers, and the
response contract for each.
inputs/pdf/without_context/: one context-removed PDF per piece.
inputs/images/: rendered score-page images for every piece.
inputs/abc/: one ABC score per piece.
inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.scientific-multitask-instructions
Scientific Multitask Instructions
A multi-task scientific instruction-following dataset created for
supervised fine-tuning and preference-optimization experiments.
Dataset summary
The dataset contains 1,576 conversational scientific examples across
eight task types.
Split
Examples
Train
1,260
Validation
158
Test
158
Total
1,576
Task distribution
Task
Examples
Scientific question answering
256
Summarization
220… See the full description on the dataset page: https://huggingface.co/datasets/Miladsaeedi70/scientific-multitask-instructions.mk-alpaca-cleaned
Citation Information
@misc{alpaca,
author = {Rohan Taori and Ishaan Gulrajani and Tianyi Zhang and Yann Dubois and Xuechen Li and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto },
title = {Stanford Alpaca: An Instruction-following LLaMA model},
year = {2023},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/tatsu-lab/stanford_alpaca}},
}```
mk-ultrachat42k
Dataset Card for UltraChat 42k
Dataset Description
The original UltraChat dataset consists of 1.4 million dialogues generated by ChatGPT and covers a wide range of topics. Our goal is to translate 100,000 examples into Macedonian. So far we have translated 42,767 examples.
Dataset Structure
The dataset has one split, suitable for:
Supervised fine-tuning (sft)
The number of examples per split is shown as follows:
train_sft
test_sft
train_gen
test_gen… See the full description on the dataset page: https://huggingface.co/datasets/milanvelinovski/mk-ultrachat42k.
