CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01songjhPKU /PM4Bench PM4Bench Strictly parallel multilingual evaluation for Large Vision-Language Models Overview The paper Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning introduces PM4Bench to separate language effects from dataset variation. Its content is strictly parallel across ten languages, and its vision setting renders textual inputs directly into images. Comparing that setting with interleaved input identifies OCR… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/PM4Bench.imagevisual-question-answering10K<n<100K2 likes4.1k downloads1mo agoHugging Face02songlab /gpn-animal-promoter-datasettext1M<n<10M1 likes981 downloads2y agoHugging Face03vava22684 /song-jury-leaderboardtextn<1K2 likes248 downloads4h agoHugging Face04SongzeLi /SID-VLN Datasets of Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale. tabular1K<n<10K0 likes196 downloads1y agoHugging Face05songlab /genomes-brassicales-balanced-v1More info: https://github.com/songlab-cal/gpn tabular1M<n<10M0 likes189 downloads3y agoHugging Face06KaraKaraWitch /uta-net-songs Dataset Details Dataset Description This dataset contains a processed version of a web scrape I did for uta-net. The raw data is available for download at here. Uta-Net site mainly lists songs that have been released in Japan officially (Anime OP/EDs) up to 2023-03. Curated by: KaraKaraWitch Shared by: KaraKaraWitch Language(s) (NLP): JA License: Not Disclosed / Unsure Stuff not in this dataset: Character Songs for Anime Doujin/Indie Works Dataset Sample… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/uta-net-songs.text100K<n<1M2 likes105 downloads2y agoHugging Face07yiwen-song /PaperWritingBench PaperWritingBench 🎻 PaperWritingBench is the first benchmark designed to evaluate how well autonomous AI research paper writing systems can synthesize raw research materials into submission-ready papers. [Paper] [Project Page] [Code] Dataset Structure This repository contains: datasets.zip: The full dataset containing cvpr2025 and iclr2025 folders with raw materials. metadata.json: A JSON file listing metadata for all 200 papers, including venue, paper ID, number… See the full description on the dataset page: https://huggingface.co/datasets/yiwen-song/PaperWritingBench.imagetext-generationn<1K0 likes91 downloads4mo agoHugging Face08Mar2Ding /songcompose_data [ACL 2025] SongCompose Dataset This repository hosts the official dataset used in SongComposer, a system designed for aligning lyrics and melody for LLMs-based vocal composition. 🌟 Overview The dataset includes three types of aligned resources, grouped by language (English and Chinese): lyric: Unpaired lyrics in English and Chinese melody: Unpaired melodies (note sequences and durations) pair: Aligned lyric-melody pairs with note durations, rest durations, structures… See the full description on the dataset page: https://huggingface.co/datasets/Mar2Ding/songcompose_data.textn<1K2 likes86 downloads1y agoHugging Face09songbo /multi3woztext10K<n<100K0 likes68 downloads3y agoHugging Face10myduy /vnexpress_plain_text_doi_songtext1K<n<10K0 likes68 downloads1y agoHugging Face11Agtian /songketdata3gatedtext10K<n<100K0 likes46 downloads19d agoHugging Face12sunbv56 /song_dataset 🎵 Vietnamese Song Lyrics and Word Timestamps Dataset Dataset Summary The song_dataset provides high-quality Vietnamese song data, including metadata, full lyrics, and particularly word-level timestamps. This dataset is optimally designed for tasks such as: Training and evaluating automatic speech recognition (ASR) models on music. Lyrics synchronization (Lyrics Alignment / Karaoke generation). Natural language processing (NLP) analysis on song lyrics. The data is… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset.text1K<n<10K0 likes45 downloads7mo agoHugging Face13junhao1122 /Classical-Chinese-Poetry-Songs Classical Chinese Poetry Songs Version 1.0 · Chinese poetry-to-song dataset · MIT License This is the frozen dataset used in Classical Chinese Poetry Song Generation: A Curated Dataset and Domain Adaptation (working manuscript title). It contains synthetic songs with vocals and accompaniment, poetic lyrics, audio-grounded captions, original generation descriptions, and work-level splits. The upstream song model was identified by generation providers as Suno V6; this label was… See the full description on the dataset page: https://huggingface.co/datasets/junhao1122/Classical-Chinese-Poetry-Songs.tabulartext-to-audio1K<n<10K0 likes39 downloads2d agoHugging Face14songjhPKU /RxnOptBench RxnOptBench RxnOptBench is a fixed benchmark release for offline evaluation of reaction-condition recommendation systems. It packages the current testsetV6 benchmark as a Hugging Face dataset repository that is compatible with the Dataset Viewer and Croissant metadata generation workflow required by the NeurIPS 2026 Evaluations and Datasets hosting guidelines. This review release was generated on 2026-05-03T15:52:23+00:00 with release version v6-neurips2026-review. It contains 4773… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/RxnOptBench.text1K<n<10K0 likes38 downloads5mo agoHugging Face15bcywinski /taboo-song taboo-song This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-song") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes29 downloads1y agoHugging Face16AlekseyCalvin /song_lyrics_Ru2En_PostSoviet_alt_anthemsLyrics to songs by seminal Soviet and Russian songwriters, poets, and bands. Manually translated to English with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric, syllabic, melodic, and other lyrical/performance-catered features, whilst retaining adequate semantic/significational fidelity. This repo's variant of the dataset was compiled/structured for ORPO-style fine-tuning of LLMs. The sampling herein constitues a variated… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/song_lyrics_Ru2En_PostSoviet_alt_anthems.texttranslation1K<n<10K1 likes27 downloads1y agoHugging Face17SongKun909 /Lithium-Battery-IE-Dataset Lithium-Ion Battery Patent Technical Indicator Dataset (锂离子电池专利技术指标精标数据集) Introduction (简介) This repository provides a highly specialized, bilingual (Chinese & English) instruction-tuning dataset designed for Fine-grained Information Extraction (IE) from Lithium-ion battery patents. It is the official data repository for our data paper: [A Dataset of Fine-Grained Technical Indicators from Lithium-Ion Battery Patents for Instruction Tuning of Large Language Models].… See the full description on the dataset page: https://huggingface.co/datasets/SongKun909/Lithium-Battery-IE-Dataset.texttext-generationn<1K0 likes26 downloads5mo agoHugging Face18AlekseyCalvin /Lyrical_Ru2En_Poems_Songs_MeterMatched_jsonl_SFT LYRICAL Russian2English SFT Version: Meaning+Meter-Matched Russian & Soviet Poems + Songs Manually Adapted by a Poet-Translator 1776 rows/items and 2 columns JSONL (JsonLine) version Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric, syllabic, melodic, and other lyrical and literary features, whilst retaining adequate semantic/significational fidelity.… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_Ru2En_Poems_Songs_MeterMatched_jsonl_SFT.texttranslation1K<n<10K0 likes23 downloads11mo agoHugging Face19AlekseyCalvin /Lyrical_ru2en_v5_songs_poems_MeterMatched_DPO Meaning+Meter-Matched Russian & Soviet Poems + Songs Manually Translated by a Poet-Translator from Russian to English Translations herein faithfully adapt the Source Lyrics' Metered/Rhythmic/Rhyming Patterns EDITED VARIANT 5 Re-balanced, refined, standardized, and substantially expanded. JSONL variant Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_ru2en_v5_songs_poems_MeterMatched_DPO.texttranslation1K<n<10K0 likes21 downloads11mo agoHugging Face20SongY123 /GeoSQL-Synth GeoSQL-Synth GeoSQL-Synth is a supervised fine-tuning dataset for translating natural language questions into PostGIS SQL queries. Fields Field Type Description instruction string Constant task instruction that asks the model to act as a PostgreSQL/PostGIS expert and generate SQL from the provided schema and question. input string Prompt body containing [Database Schema], [DB_ID], [Schema], and [User Question] sections. output string Target SQL… See the full description on the dataset page: https://huggingface.co/datasets/SongY123/GeoSQL-Synth.text10K<n<100K0 likes21 downloads3mo agoHugging Face21Agtian /Songketdata5gatedtext10K<n<100K0 likes21 downloads18d agoHugging Face22songtingyu /limrank-dataPlease refer to https://github.com/SighingSnow/limrank for usage check. text10K<n<100K0 likes19 downloads10mo agoHugging Face23SongTonyLi /c2x86-leetcode-eval c2x86 compilation pairs Pairs of C source files and corresponding x86 assembly. Columns: c (C code), s (x86 assembly) Suggested load: import json # JSONL files with one object per line train = [json.loads(l) for l in open('train.jsonl', 'r', encoding='utf-8')] test = [json.loads(l) for l in open('test.jsonl', 'r', encoding='utf-8')] textn<1K0 likes19 downloads10mo agoHugging Face24sunbv56 /song_dataset_chunked Vietnamese Songs Word-Level Timestamp Dataset (Chunked) This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper. Dataset Summary The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences. Duration Insights: Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset_chunked.tabular10K<n<100K0 likes19 downloads7mo agoHugging Face25Agtian /Songket_Data_1gatedtext1K<n<10K0 likes18 downloads19d agoHugging Face26fahdmirzac /urdu_bollywood_songs_dataset Bollywood-Inspired Dataset: Movies and Songs Created by Fahd Mirza = https://www.youtube.com/@fahdmirza Overview This dataset is a creative collection of fictional Bollywood movie titles paired with equally fictional song lyrics. Inspired by the rich tradition of Bollywood cinema, where music plays a pivotal role in storytelling, this dataset aims to provide a unique resource for exploring the interplay between movie themes and their musical expressions. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fahdmirzac/urdu_bollywood_songs_dataset.textn<1K0 likes17 downloads3y agoHugging Face27raincandy-u /VOCALOID_songstext1K<n<10K0 likes17 downloads2y agoHugging Face28Alienanthony /LLM_Broadcaster_Song_Introductionstext10K<n<100K0 likes17 downloads2y agoHugging Face29songff /GenerAlign Dataset Card GenerAlign is collected to help construct well-aligned LLMs in general domains, such as harmlessness, helpfulness, and honesty. It contains 31398 prompts from existed datasets, including: FLAN HH-RLHF FalseQA UltraChat ShareGPT Similar to UltraFeedback, we complete each prompt with responses from different LLMs, including: Llama-3.1-Nemotron-70B-Instruct-HF Llama-3.2-3B-Instructgemma-2-27b-it All responses are annotated by ArmoRM-Llama3-8B-v0.1. This dataset has… See the full description on the dataset page: https://huggingface.co/datasets/songff/GenerAlign.tabulartext-generation10K<n<100K2 likes16 downloads1y agoHugging Face30jamimulgrave /Song-Interpretation-Datasettext100K<n<1M0 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.