datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
news_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。
synthetic-football-commentary-qwen
Synthetic Passionate Football Commentary
Dataset Summary
This dataset contains synthetic conversational data designed to fine-tune large language models for creative writing and persona adoption. Specifically, it trains models to act as a passionate football commentator. The data pairs factual football match events with highly dramatic, emotional, and tactical commentary.
Data Generation
Base Data: The raw input features (Minute, Match, Team, Player, Action)… See the full description on the dataset page: https://huggingface.co/datasets/Alpaczyk/synthetic-football-commentary-qwen.cricket-commentary-dataset
Cricket Commentary Dataset
Description
A curated dataset of cricket commentary examples for fine-tuning language models to generate exciting sports commentary.
Dataset Structure
Each example contains:
instruction: Task description
input: Match situation (batsman, bowler, action, result)
output: Professional commentary text
Example
{
"instruction": "Generate exciting cricket commentary for this moment",
"input": "Batsman: Kohli, Bowler: Starc… See the full description on the dataset page: https://huggingface.co/datasets/siva-gunasehkaran/cricket-commentary-dataset.
