datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PM4Bench
PM4Bench
Strictly parallel multilingual evaluation for Large Vision-Language Models
Overview
The paper Benchmarking and Boosting Multilingual Capabilities of LVLMs via
OCR-Centric Reinforcement Learning
introduces PM4Bench to separate language effects from dataset variation. Its
content is strictly parallel across ten languages, and its vision setting
renders textual inputs directly into images. Comparing that setting with
interleaved input identifies OCR… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/PM4Bench.gpn-animal-promoter-datasetsong-jury-leaderboardSID-VLN Datasets of Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale.
genomes-brassicales-balanced-v1More info: https://github.com/songlab-cal/gpn
uta-net-songs
Dataset Details
Dataset Description
This dataset contains a processed version of a web scrape I did for uta-net. The raw data is available for download at here.
Uta-Net site mainly lists songs that have been released in Japan officially (Anime OP/EDs) up to 2023-03.
Curated by: KaraKaraWitch
Shared by: KaraKaraWitch
Language(s) (NLP): JA
License: Not Disclosed / Unsure
Stuff not in this dataset:
Character Songs for Anime
Doujin/Indie Works
Dataset Sample… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/uta-net-songs.PaperWritingBench
PaperWritingBench 🎻
PaperWritingBench is the first benchmark designed to evaluate how well autonomous AI research paper writing systems can synthesize raw research materials into submission-ready papers.
[Paper] [Project Page] [Code]
Dataset Structure
This repository contains:
datasets.zip: The full dataset containing cvpr2025 and iclr2025 folders with raw materials.
metadata.json: A JSON file listing metadata for all 200 papers, including venue, paper ID, number… See the full description on the dataset page: https://huggingface.co/datasets/yiwen-song/PaperWritingBench.songcompose_data
[ACL 2025] SongCompose Dataset
This repository hosts the official dataset used in SongComposer, a system designed for aligning lyrics and melody for LLMs-based vocal composition.
🌟 Overview
The dataset includes three types of aligned resources, grouped by language (English and Chinese):
lyric: Unpaired lyrics in English and Chinese
melody: Unpaired melodies (note sequences and durations)
pair: Aligned lyric-melody pairs with note durations, rest durations, structures… See the full description on the dataset page: https://huggingface.co/datasets/Mar2Ding/songcompose_data.multi3wozvnexpress_plain_text_doi_songsongketdata3song_dataset
🎵 Vietnamese Song Lyrics and Word Timestamps Dataset
Dataset Summary
The song_dataset provides high-quality Vietnamese song data, including metadata, full lyrics, and particularly word-level timestamps.
This dataset is optimally designed for tasks such as:
Training and evaluating automatic speech recognition (ASR) models on music.
Lyrics synchronization (Lyrics Alignment / Karaoke generation).
Natural language processing (NLP) analysis on song lyrics.
The data is… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset.Classical-Chinese-Poetry-Songs
Classical Chinese Poetry Songs
Version 1.0 · Chinese poetry-to-song dataset · MIT License
This is the frozen dataset used in Classical Chinese Poetry Song Generation:
A Curated Dataset and Domain Adaptation (working manuscript title).
It contains synthetic songs with vocals and accompaniment, poetic lyrics,
audio-grounded captions, original generation descriptions, and work-level splits.
The upstream song model was identified by generation providers as Suno V6;
this label was… See the full description on the dataset page: https://huggingface.co/datasets/junhao1122/Classical-Chinese-Poetry-Songs.RxnOptBench
RxnOptBench
RxnOptBench is a fixed benchmark release for offline evaluation of reaction-condition recommendation systems. It packages the current testsetV6 benchmark as a Hugging Face dataset repository that is compatible with the Dataset Viewer and Croissant metadata generation workflow required by the NeurIPS 2026 Evaluations and Datasets hosting guidelines.
This review release was generated on 2026-05-03T15:52:23+00:00 with release version v6-neurips2026-review. It contains 4773… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/RxnOptBench.taboo-song
taboo-song
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-song")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
song_lyrics_Ru2En_PostSoviet_alt_anthemsLyrics to songs by seminal Soviet and Russian songwriters, poets, and bands.
Manually translated to English with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric, syllabic, melodic, and other lyrical/performance-catered features, whilst retaining adequate semantic/significational fidelity.
This repo's variant of the dataset was compiled/structured for ORPO-style fine-tuning of LLMs.
The sampling herein constitues a variated… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/song_lyrics_Ru2En_PostSoviet_alt_anthems.Lithium-Battery-IE-Dataset
Lithium-Ion Battery Patent Technical Indicator Dataset (锂离子电池专利技术指标精标数据集)
Introduction (简介)
This repository provides a highly specialized, bilingual (Chinese & English) instruction-tuning dataset designed for Fine-grained Information Extraction (IE) from Lithium-ion battery patents. It is the official data repository for our data paper: [A Dataset of Fine-Grained Technical Indicators from Lithium-Ion Battery Patents for Instruction Tuning of Large Language Models].… See the full description on the dataset page: https://huggingface.co/datasets/SongKun909/Lithium-Battery-IE-Dataset.Lyrical_Ru2En_Poems_Songs_MeterMatched_jsonl_SFT
LYRICAL Russian2English SFT Version:
Meaning+Meter-Matched Russian & Soviet Poems + Songs
Manually Adapted by a Poet-Translator
1776 rows/items and 2 columns
JSONL (JsonLine) version
Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric, syllabic, melodic, and other lyrical and literary features, whilst retaining adequate semantic/significational fidelity.… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_Ru2En_Poems_Songs_MeterMatched_jsonl_SFT.Lyrical_ru2en_v5_songs_poems_MeterMatched_DPO
Meaning+Meter-Matched Russian & Soviet Poems + Songs
Manually Translated by a Poet-Translator from Russian to English
Translations herein faithfully adapt the Source Lyrics' Metered/Rhythmic/Rhyming Patterns
EDITED VARIANT 5
Re-balanced, refined, standardized, and substantially expanded.
JSONL variant
Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_ru2en_v5_songs_poems_MeterMatched_DPO.GeoSQL-Synth
GeoSQL-Synth
GeoSQL-Synth is a supervised fine-tuning dataset for translating natural language questions into PostGIS SQL queries.
Fields
Field
Type
Description
instruction
string
Constant task instruction that asks the model to act as a PostgreSQL/PostGIS expert and generate SQL from the provided schema and question.
input
string
Prompt body containing [Database Schema], [DB_ID], [Schema], and [User Question] sections.
output
string
Target SQL… See the full description on the dataset page: https://huggingface.co/datasets/SongY123/GeoSQL-Synth.Songketdata5limrank-dataPlease refer to https://github.com/SighingSnow/limrank for usage check.
c2x86-leetcode-eval
c2x86 compilation pairs
Pairs of C source files and corresponding x86 assembly.
Columns: c (C code), s (x86 assembly)
Suggested load:
import json
# JSONL files with one object per line
train = [json.loads(l) for l in open('train.jsonl', 'r', encoding='utf-8')]
test = [json.loads(l) for l in open('test.jsonl', 'r', encoding='utf-8')]
song_dataset_chunked
Vietnamese Songs Word-Level Timestamp Dataset (Chunked)
This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper.
Dataset Summary
The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences.
Duration Insights:
Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset_chunked.Songket_Data_1urdu_bollywood_songs_dataset
Bollywood-Inspired Dataset: Movies and Songs
Created by Fahd Mirza = https://www.youtube.com/@fahdmirza
Overview
This dataset is a creative collection of fictional Bollywood movie titles paired with equally fictional song lyrics. Inspired by the rich tradition of Bollywood cinema, where music plays a pivotal role in storytelling, this dataset aims to provide a unique resource for exploring the interplay between movie themes and their musical expressions.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fahdmirzac/urdu_bollywood_songs_dataset.VOCALOID_songsLLM_Broadcaster_Song_IntroductionsGenerAlign
Dataset Card
GenerAlign is collected to help construct well-aligned LLMs in general domains, such as harmlessness, helpfulness, and honesty. It contains 31398 prompts from existed datasets, including:
FLAN
HH-RLHF
FalseQA
UltraChat
ShareGPT
Similar to UltraFeedback, we complete each prompt with responses from different LLMs, including:
Llama-3.1-Nemotron-70B-Instruct-HF
Llama-3.2-3B-Instructgemma-2-27b-it
All responses are annotated by ArmoRM-Llama3-8B-v0.1.
This dataset has… See the full description on the dataset page: https://huggingface.co/datasets/songff/GenerAlign.Song-Interpretation-Dataset
