datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
news_commentary
Dataset Card for OPUS News-Commentary
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/news_commentary.OPUS_News-Commentaryparallel-sentences-news-commentary
Dataset Card for Parallel Sentences - News Commentary
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the News-Commentary dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-news-commentary.game_commentary_sftfootball-commentary-dataset
⚽ YallaShoot Football Commentary Dataset
A curated dataset of football (soccer) match commentary, player mentions, and match events — built to power NLP models for the Arabic and global football community.
📌 Dataset Description
This dataset contains structured football match commentary collected from live match feeds, covering top leagues including:
🏴 English Premier League (EPL)
🇪🇸 La Liga
🏆 UEFA Champions League
🌍 Arab World Leagues (Saudi Pro… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/football-commentary-dataset.swiss-law-commentary
Swiss Law Commentary Dataset
Open-access, AI-generated legal commentary on 8 Swiss federal laws.
Source: openlegalcommentary.ch
Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG
Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest)
Languages: DE, FR, IT, EN
Articles exported: 215
License: CC BY-SA 4.0
Built with ♥ by Jonas Hertner
Files
One JSONL file per law (e.g., or.jsonl, zgb.jsonl).
Schema
Field
Type
Description
law… See the full description on the dataset page: https://huggingface.co/datasets/voilaj/swiss-law-commentary.news_commentary
Dataset Card for "news_commentary"
More Information needed
pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.news_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。
news-commentary-v18.de-ensynthetic-football-commentary-qwen
Synthetic Passionate Football Commentary
Dataset Summary
This dataset contains synthetic conversational data designed to fine-tune large language models for creative writing and persona adoption. Specifically, it trains models to act as a passionate football commentator. The data pairs factual football match events with highly dramatic, emotional, and tactical commentary.
Data Generation
Base Data: The raw input features (Minute, Match, Team, Player, Action)… See the full description on the dataset page: https://huggingface.co/datasets/Alpaczyk/synthetic-football-commentary-qwen.swiss-law-commentary
Swiss Law Commentary Dataset
Open-access, AI-generated legal commentary on 8 Swiss federal laws.
Source: openlegalcommentary.ch
Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG
Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest)
Languages: DE, FR, IT, EN
Articles exported: 215
License: CC BY-SA 4.0
Built with ♥ by Jonas Hertner
Files
One JSONL file per law (e.g., or.jsonl, zgb.jsonl).
Schema
Field
Type
Description
law… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/swiss-law-commentary.MultilingualPoetry-CommentaryPairs
Multilingual Poetry Commentary Pairs
多语言诗歌评论对数据集
📊 数据集概述
数据集名称: PoetryMTEB/MultilingualPoetry-CommentaryPairs
维度数量: 7
总数据条数: 14063
总语言数量: 120
数据格式: Parquet
分割: 全部为test集
🎯 数据集用途
本数据集用于诗歌文本分析、文学评论生成、文本相似度计算等任务。每个诗歌-评论对包含:
诗歌文本
专业的文学分析评论
元数据信息
📁 数据集结构
维度说明
数据集按照诗歌分析的七个维度进行组织:
维度(中文)
维度(英文)
描述
语言风格
language_style
分析诗歌的语言特点、风格特征
艺术与修辞手法
artistic_techniques
分析诗歌的艺术技巧、修辞手法
意象与象征
imagery_symbolism
分析诗歌的意象、象征意义
情感脉络… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/MultilingualPoetry-CommentaryPairs.cricket-commentary-ball-datanews_commentarycricket-commentaryCricket-Commentarynews_commentary-en-ar-translationnews-commentary-eng-arz
Dataset details
In this version of the News Commentary dataset, Standard Arabic text segments are converted into Egyptian Arabic (ARZ) using GPT-4.1-Mini.
We calculated the semantic similarity between the English source and the Egyptian Arabic target and selected the 500 segments with the highest scores for the test split,
while the train split comprises the remaining 83.2K segments.
∙ Dataset columns
"english": original English text
"arabic": original Standard… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/news-commentary-eng-arz.afcon2025-commentary
AFCON 2025 Match Commentary Dataset
High-quality French match commentary training data for the Africa Cup of Nations 2025.
Dataset Details
Size: 2,000 training examples
Language: French
Format: JSONL (chat template)
Use Case: Fine-tuning LLMs for realistic African football match commentary
Event Distribution
82% General commentary
10% Goals
5% Substitutions
2% Penalties
1% Cards (yellow/red)
Teams Covered
Morocco, Senegal, Egypt, Nigeria, Côte… See the full description on the dataset page: https://huggingface.co/datasets/oxmo88/afcon2025-commentary.news-commentary-en-arThis is a filtered version of the English-to-Arabic News Commentary dataset available at data.statmt.org/news-commentary.
The filtering process includes removing duplicates, language detection, and semantic filtering based on similarity (>0.70) between the source and translation.
The filtering script is available at data-processing.ipynb.
generic_covas_commentary_v2Cricket-Commentary-SampleFootball-Commentarynews-commentary-cs-defleets-of-wwii-design-commentary
Fleets of World War II — warship classes & design commentary
The structured database inside Fleets of World War II: Design History and Analysis
for Every Ship of Every Navy (Richard Worth, Nimble Books, ISBN 9781608881604). Two layers:
config
rows
what
ship_classes
1009
one record per warship class — nation, ship type, class name, ships in class, key specs, and Worth's design commentary (why it was built that way, its strengths/flaws)
section_overview
216
the Nation… See the full description on the dataset page: https://huggingface.co/datasets/wfzimmerman/fleets-of-wwii-design-commentary.FIFA_commentarycricket-commentary-dataset
Cricket Commentary Dataset
Description
A curated dataset of cricket commentary examples for fine-tuning language models to generate exciting sports commentary.
Dataset Structure
Each example contains:
instruction: Task description
input: Match situation (batsman, bowler, action, result)
output: Professional commentary text
Example
{
"instruction": "Generate exciting cricket commentary for this moment",
"input": "Batsman: Kohli, Bowler: Starc… See the full description on the dataset page: https://huggingface.co/datasets/siva-gunasehkaran/cricket-commentary-dataset.generic_situational_event_commentary_v2generic_covas_commentary
