CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xu-song /cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100. Languages To load a language which isn't part of the config, all you need to do is specify the language code in the config. You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/ E.g. dataset = load_dataset("cc100-samples", lang="en") VALID_CODES = [ "am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de", "el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.texttext-generation1M<n<10M6 likes686 downloads2y agoHugging Face02songdj /ScratchMath ScratchMath Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math AIED 2026 — 27th International Conference on Artificial Intelligence in Education Overview ScratchMath is a multimodal benchmark for evaluating whether MLLMs can analyze handwritten mathematical scratchwork produced by real students. Unlike existing math benchmarks that focus on problem-solving accuracy, ScratchMath targets error diagnosis — identifying… See the full description on the dataset page: https://huggingface.co/datasets/songdj/ScratchMath.imagevisual-question-answering1K<n<10K2 likes168 downloads6mo agoHugging Face03asigalov61 /clean-songs-lyrics-dataset Clean Songs Lyrics Dataset 1.53M+ clean songs lyrics with songs titles and artists names Dataset info This is the combined, deduped, cleaned, and sanitized aggregation of three large lyrics datasets Each lyric was deduplicated Each lyric was checked to be in range of 256 bytes <-> 8192 bytes Each lyric was checked for profanities with alt-profanity-check Each lyric was ASCII sanitized for conistency… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/clean-songs-lyrics-dataset.texttext-generation1M<n<10M2 likes147 downloads5mo agoHugging Face04ThatOneShortGuy /SongLyricsDataset contains songs by artists, the names of the songs, the lyrics of the songs, the release date, the cover photo, and the general popularity of the song. imagetext-generation100K<n<1M4 likes145 downloads3y agoHugging Face05Nicolas-BZRD /English_French_Songs_Lyrics_Translation_Original Original Songs Lyrics with French Translation Dataset Summary Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French. Details of the number of songs by language of origin can be found in the table below: Original language Number of songs en 75786 fr 18486 es 1743 it 803 de 691 sw 529 ko 193 id 169 pt 142 no 122 fi 113 sv 70 hr 53 so 43 ca 41 tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.tabulartranslation10K<n<100K16 likes103 downloads3y agoHugging Face06yiwen-song /PaperWritingBench PaperWritingBench 🎻 PaperWritingBench is the first benchmark designed to evaluate how well autonomous AI research paper writing systems can synthesize raw research materials into submission-ready papers. [Paper] [Project Page] [Code] Dataset Structure This repository contains: datasets.zip: The full dataset containing cvpr2025 and iclr2025 folders with raw materials. metadata.json: A JSON file listing metadata for all 200 papers, including venue, paper ID, number… See the full description on the dataset page: https://huggingface.co/datasets/yiwen-song/PaperWritingBench.imagetext-generationn<1K0 likes69 downloads4mo agoHugging Face07SynhalaAI /Sinhala-Song-Lyrics 🎵 SynhalaAI — Ultimate Sinhala Song Lyrics Dataset (Gold Mix) Dataset Description The SynhalaAI Lyrics Corpus is a meticulously engineered, high-fidelity dataset of Sinhala song lyrics. It was designed specifically to train Large Language Models (LLMs) and advanced tokenizers on the poetic, colloquial, and structured linguistic patterns of the Sinhala language. Unlike standard web-scraped datasets that are littered with English guitar chords, metadata, and HTML… See the full description on the dataset page: https://huggingface.co/datasets/SynhalaAI/Sinhala-Song-Lyrics.texttext-generation1K<n<10K2 likes54 downloads3mo agoHugging Face08SpartanCinder /song-lyrics-artist-classifiertexttext-classification10K<n<100K2 likes53 downloads2y agoHugging Face09palette-lab /songgot-tools-ko Songgot Tools KO Synthetic Korean tool-calling data used to post-train Songgot (github.com/hanishkeloth/songgot): 336,602 verified (request, call) pairs over 183,737 tool schemas across 60 Korean service domains. Released 2026-09-11 under Apache 2.0. How it was made Every row was produced by our own open teacher model, Palette-K-Midm, in three stages (harness/teacher_synth.py in the Songgot repo): Schemas: the teacher invents tool schemas for a domain in a given… See the full description on the dataset page: https://huggingface.co/datasets/palette-lab/songgot-tools-ko.text-generation100K<n<1M0 likes48 downloads15d agoHugging Face10rajtripathi /5M-Songs-Lyrics Dataset Summary This dataset contains 50 million rows of song lyrics sourced from a public Kaggle dataset. It has been preprocessed into an instruction–label format suitable for training or fine-tuning generative language models, particularly for music lyric generation tasks. Each row is designed to guide a model to generate song verses in the style of a specific artist and genre, with corresponding real lyric snippets as ground truth. Supported Tasks and Benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/rajtripathi/5M-Songs-Lyrics.texttext-generation1M<n<10M1 likes43 downloads1y agoHugging Face11Song-SW /CEBThe dataset for bias evaluation of LLMs. Github: https://github.com/SongW-SW/CEB text-classification1 likes42 downloads2y agoHugging Face12songjhPKU /cc-arena-dataset CC-Arena Benchmark Dataset Full benchmark datasets for CC-Arena — a framework for evaluating AI coding agents (Claude Code, Cursor, etc.). Quick Start Via CC-Arena CLI (recommended) # Download a specific benchmark python3 -m cc_arena.tasks.downloader download humaneval # Download with limit python3 -m cc_arena.tasks.downloader download bigcodebench --limit 100 # List all available benchmarks python3 -m cc_arena.tasks.downloader list Via… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/cc-arena-dataset.text-generation1K<n<10K1 likes34 downloads6mo agoHugging Face13fr3on /egyptian-songs Egyptian Arabic Songs Dataset 🎵 Dataset Description This dataset contains 3,063 lines of Egyptian Arabic song lyrics with English translations, spanning 280 songs from 1983-2021. The dataset features natural Egyptian Arabic dialect (العامية المصرية) as used in popular music, with automatic genre classification and line-type detection. Languages Source: Egyptian Arabic (ar_EG) - Colloquial dialect in music Target: English (en) Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/egyptian-songs.tabulartranslation1K<n<10K0 likes31 downloads9mo agoHugging Face14SongKun909 /Lithium-Battery-IE-Dataset Lithium-Ion Battery Patent Technical Indicator Dataset (锂离子电池专利技术指标精标数据集) Introduction (简介) This repository provides a highly specialized, bilingual (Chinese & English) instruction-tuning dataset designed for Fine-grained Information Extraction (IE) from Lithium-ion battery patents. It is the official data repository for our data paper: [A Dataset of Fine-Grained Technical Indicators from Lithium-Ion Battery Patents for Instruction Tuning of Large Language Models].… See the full description on the dataset page: https://huggingface.co/datasets/SongKun909/Lithium-Battery-IE-Dataset.texttext-generationn<1K0 likes29 downloads5mo agoHugging Face15bcywinski /taboo-song taboo-song This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-song") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes28 downloads1y agoHugging Face16HeshamHaroon /1milion_token_EGY_songstexttext-generation1K<n<10K9 likes25 downloads2y agoHugging Face17Shakil2448868 /bangla-rabindranath-songs-synth_promptstexttext-generationn<1K0 likes20 downloads2y agoHugging Face18songff /GenerAlign Dataset Card GenerAlign is collected to help construct well-aligned LLMs in general domains, such as harmlessness, helpfulness, and honesty. It contains 31398 prompts from existed datasets, including: FLAN HH-RLHF FalseQA UltraChat ShareGPT Similar to UltraFeedback, we complete each prompt with responses from different LLMs, including: Llama-3.1-Nemotron-70B-Instruct-HF Llama-3.2-3B-Instructgemma-2-27b-it All responses are annotated by ArmoRM-Llama3-8B-v0.1. This dataset has… See the full description on the dataset page: https://huggingface.co/datasets/songff/GenerAlign.tabulartext-generation10K<n<100K2 likes19 downloads1y agoHugging Face19metncelik /turkish-song-lyricstexttext-classification10K<n<100K2 likes17 downloads1y agoHugging Face20LorthGyu /indonesian-song-lyrics Lirik Lagu Indonesia 🎵 Kumpulan lirik lagu Indonesia — lagu wajib nasional, lagu daerah, dan lagu populer klasik, lengkap dengan judul, artis, genre, dan tahun. Kenapa dataset ini ada? Dataset lirik lagu bahasa Indonesia di HF belum ada — padahal lagu nasional & daerah adalah warisan budaya yang public domain dan aman dipakai. Lirik ini bagus buat fine-tune (pola rima, diksi puitis) atau analysis budaya. Isi Field Tipe Contoh title… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-song-lyrics.tabulartext-generationn<1K0 likes13 downloads2mo agoHugging Face21lluccardoner /melodyGPT-song-chords-text-1 melodyGPT song chords dataset This dataset contains the text representation of song chords. Dataset Details Dataset Description This dataset is created by aggregating the chords of each song given by the Chords and Lyrics Dataset. You can see in the dataset folder of the Github repository of melodyGPT notebooks with the code used to do so. Also, the special characters that are not chords are analysed briefly and this information will be used to create… See the full description on the dataset page: https://huggingface.co/datasets/lluccardoner/melodyGPT-song-chords-text-1.texttext-generation100K<n<1M1 likes11 downloads2y agoHugging Face22Shakil2448868 /bangla-songs-synthetic-prompt Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Shakil2448868/bangla-songs-synthetic-prompt.texttext-generation1K<n<10K0 likes8 downloads2y agoHugging Face23Agtian /Songket_DatagatedDataset instruksi penalaran (reasoning) berbahasa Indonesia dari sumber publik untuk riset NLP, evaluasi LLM, dan fine-tuning model. 📊 Spesifikasi Data Proses: Hanya deduplikasi baris (remove duplicate) pada tingkat pesan. Format: Skema konsisten memiliki struktur messages dengan role user dan assistant. Di dalam role assistant, terdapat key reasoning_content (proses berpikir) dan content (jawaban akhir). ⚠️ Disclaimer Hak Cipta: Hak cipta sepenuhnya milik penulis atau sumber publik asli… See the full description on the dataset page: https://huggingface.co/datasets/Agtian/Songket_Data.text-generation100K<n<1M0 likes6 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.