datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100.
Languages
To load a language which isn't part of the config, all you need to do is specify the language code in the config.
You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/
E.g.
dataset = load_dataset("cc100-samples", lang="en")
VALID_CODES = [
"am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de",
"el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.ScratchMath
ScratchMath
Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math
AIED 2026 — 27th International Conference on Artificial Intelligence in Education
Overview
ScratchMath is a multimodal benchmark for evaluating whether MLLMs can analyze handwritten mathematical scratchwork produced by real students. Unlike existing math benchmarks that focus on problem-solving accuracy, ScratchMath targets error diagnosis — identifying… See the full description on the dataset page: https://huggingface.co/datasets/songdj/ScratchMath.clean-songs-lyrics-dataset
Clean Songs Lyrics Dataset
1.53M+ clean songs lyrics with songs titles and artists names
Dataset info
This is the combined, deduped, cleaned, and sanitized aggregation of three large lyrics datasets
Each lyric was deduplicated
Each lyric was checked to be in range of 256 bytes <-> 8192 bytes
Each lyric was checked for profanities with alt-profanity-check
Each lyric was ASCII sanitized for conistency… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/clean-songs-lyrics-dataset.SongLyricsDataset contains songs by artists, the names of the songs, the lyrics of the songs, the release date, the cover photo, and the general popularity of the song.
English_French_Songs_Lyrics_Translation_Original
Original Songs Lyrics with French Translation
Dataset Summary
Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French.
Details of the number of songs by language of origin can be found in the table below:
Original language
Number of songs
en
75786
fr
18486
es
1743
it
803
de
691
sw
529
ko
193
id
169
pt
142
no
122
fi
113
sv
70
hr
53
so
43
ca
41
tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.PaperWritingBench
PaperWritingBench 🎻
PaperWritingBench is the first benchmark designed to evaluate how well autonomous AI research paper writing systems can synthesize raw research materials into submission-ready papers.
[Paper] [Project Page] [Code]
Dataset Structure
This repository contains:
datasets.zip: The full dataset containing cvpr2025 and iclr2025 folders with raw materials.
metadata.json: A JSON file listing metadata for all 200 papers, including venue, paper ID, number… See the full description on the dataset page: https://huggingface.co/datasets/yiwen-song/PaperWritingBench.Sinhala-Song-Lyrics
🎵 SynhalaAI — Ultimate Sinhala Song Lyrics Dataset (Gold Mix)
Dataset Description
The SynhalaAI Lyrics Corpus is a meticulously engineered, high-fidelity dataset of Sinhala song lyrics. It was designed specifically to train Large Language Models (LLMs) and advanced tokenizers on the poetic, colloquial, and structured linguistic patterns of the Sinhala language.
Unlike standard web-scraped datasets that are littered with English guitar chords, metadata, and HTML… See the full description on the dataset page: https://huggingface.co/datasets/SynhalaAI/Sinhala-Song-Lyrics.song-lyrics-artist-classifiersonggot-tools-ko
Songgot Tools KO
Synthetic Korean tool-calling data used to post-train Songgot (github.com/hanishkeloth/songgot): 336,602 verified
(request, call) pairs over 183,737 tool schemas across 60 Korean service domains. Released 2026-09-11 under
Apache 2.0.
How it was made
Every row was produced by our own open teacher model, Palette-K-Midm, in three stages
(harness/teacher_synth.py in the Songgot repo):
Schemas: the teacher invents tool schemas for a domain in a given… See the full description on the dataset page: https://huggingface.co/datasets/palette-lab/songgot-tools-ko.5M-Songs-Lyrics
Dataset Summary
This dataset contains 50 million rows of song lyrics sourced from a public Kaggle dataset. It has been preprocessed into an instruction–label format suitable for training or fine-tuning generative language models, particularly for music lyric generation tasks.
Each row is designed to guide a model to generate song verses in the style of a specific artist and genre, with corresponding real lyric snippets as ground truth.
Supported Tasks and Benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/rajtripathi/5M-Songs-Lyrics.CEBThe dataset for bias evaluation of LLMs. Github: https://github.com/SongW-SW/CEB
cc-arena-dataset
CC-Arena Benchmark Dataset
Full benchmark datasets for CC-Arena — a framework for evaluating AI coding agents (Claude Code, Cursor, etc.).
Quick Start
Via CC-Arena CLI (recommended)
# Download a specific benchmark
python3 -m cc_arena.tasks.downloader download humaneval
# Download with limit
python3 -m cc_arena.tasks.downloader download bigcodebench --limit 100
# List all available benchmarks
python3 -m cc_arena.tasks.downloader list
Via… See the full description on the dataset page: https://huggingface.co/datasets/songjhPKU/cc-arena-dataset.egyptian-songs
Egyptian Arabic Songs Dataset 🎵
Dataset Description
This dataset contains 3,063 lines of Egyptian Arabic song lyrics with English translations, spanning 280 songs from 1983-2021. The dataset features natural Egyptian Arabic dialect (العامية المصرية) as used in popular music, with automatic genre classification and line-type detection.
Languages
Source: Egyptian Arabic (ar_EG) - Colloquial dialect in music
Target: English (en)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/egyptian-songs.Lithium-Battery-IE-Dataset
Lithium-Ion Battery Patent Technical Indicator Dataset (锂离子电池专利技术指标精标数据集)
Introduction (简介)
This repository provides a highly specialized, bilingual (Chinese & English) instruction-tuning dataset designed for Fine-grained Information Extraction (IE) from Lithium-ion battery patents. It is the official data repository for our data paper: [A Dataset of Fine-Grained Technical Indicators from Lithium-Ion Battery Patents for Instruction Tuning of Large Language Models].… See the full description on the dataset page: https://huggingface.co/datasets/SongKun909/Lithium-Battery-IE-Dataset.taboo-song
taboo-song
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-song")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
1milion_token_EGY_songsbangla-rabindranath-songs-synth_promptsGenerAlign
Dataset Card
GenerAlign is collected to help construct well-aligned LLMs in general domains, such as harmlessness, helpfulness, and honesty. It contains 31398 prompts from existed datasets, including:
FLAN
HH-RLHF
FalseQA
UltraChat
ShareGPT
Similar to UltraFeedback, we complete each prompt with responses from different LLMs, including:
Llama-3.1-Nemotron-70B-Instruct-HF
Llama-3.2-3B-Instructgemma-2-27b-it
All responses are annotated by ArmoRM-Llama3-8B-v0.1.
This dataset has… See the full description on the dataset page: https://huggingface.co/datasets/songff/GenerAlign.turkish-song-lyricsindonesian-song-lyrics
Lirik Lagu Indonesia 🎵
Kumpulan lirik lagu Indonesia — lagu wajib nasional, lagu daerah, dan lagu populer klasik, lengkap dengan judul, artis, genre, dan tahun.
Kenapa dataset ini ada?
Dataset lirik lagu bahasa Indonesia di HF belum ada — padahal lagu nasional & daerah adalah warisan budaya yang public domain dan aman dipakai. Lirik ini bagus buat fine-tune (pola rima, diksi puitis) atau analysis budaya.
Isi
Field
Tipe
Contoh
title… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-song-lyrics.melodyGPT-song-chords-text-1
melodyGPT song chords dataset
This dataset contains the text representation of song chords.
Dataset Details
Dataset Description
This dataset is created by aggregating the chords of each song given by the Chords and Lyrics Dataset.
You can see in the dataset folder of the Github repository of melodyGPT notebooks with the code used to do so.
Also, the special characters that are not chords are analysed briefly and this information will be used to create… See the full description on the dataset page: https://huggingface.co/datasets/lluccardoner/melodyGPT-song-chords-text-1.bangla-songs-synthetic-prompt
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Shakil2448868/bangla-songs-synthetic-prompt.Songket_DataDataset instruksi penalaran (reasoning) berbahasa Indonesia dari sumber publik untuk riset NLP, evaluasi LLM, dan fine-tuning model.
📊 Spesifikasi Data
Proses: Hanya deduplikasi baris (remove duplicate) pada tingkat pesan.
Format: Skema konsisten memiliki struktur messages dengan role user dan assistant. Di dalam role assistant, terdapat key reasoning_content (proses berpikir) dan content (jawaban akhir).
⚠️ Disclaimer
Hak Cipta: Hak cipta sepenuhnya milik penulis atau sumber publik asli… See the full description on the dataset page: https://huggingface.co/datasets/Agtian/Songket_Data.
