datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
news_commentary
Dataset Card for OPUS News-Commentary
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/news_commentary.OPUS_News-Commentaryparallel-sentences-news-commentary
Dataset Card for Parallel Sentences - News Commentary
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the News-Commentary dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-news-commentary.news_commentary
Dataset Card for "news_commentary"
More Information needed
news-commentary-v18.de-enMultilingualPoetry-CommentaryPairs
Multilingual Poetry Commentary Pairs
多语言诗歌评论对数据集
📊 数据集概述
数据集名称: PoetryMTEB/MultilingualPoetry-CommentaryPairs
维度数量: 7
总数据条数: 14063
总语言数量: 120
数据格式: Parquet
分割: 全部为test集
🎯 数据集用途
本数据集用于诗歌文本分析、文学评论生成、文本相似度计算等任务。每个诗歌-评论对包含:
诗歌文本
专业的文学分析评论
元数据信息
📁 数据集结构
维度说明
数据集按照诗歌分析的七个维度进行组织:
维度(中文)
维度(英文)
描述
语言风格
language_style
分析诗歌的语言特点、风格特征
艺术与修辞手法
artistic_techniques
分析诗歌的艺术技巧、修辞手法
意象与象征
imagery_symbolism
分析诗歌的意象、象征意义
情感脉络… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/MultilingualPoetry-CommentaryPairs.news_commentarycricket-commentaryCricket-Commentarynews_commentary-en-ar-translationnews-commentary-eng-arz
Dataset details
In this version of the News Commentary dataset, Standard Arabic text segments are converted into Egyptian Arabic (ARZ) using GPT-4.1-Mini.
We calculated the semantic similarity between the English source and the Egyptian Arabic target and selected the 500 segments with the highest scores for the test split,
while the train split comprises the remaining 83.2K segments.
∙ Dataset columns
"english": original English text
"arabic": original Standard… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/news-commentary-eng-arz.Cricket-Commentary-Sampleafcon2025-commentary
AFCON 2025 Match Commentary Dataset
High-quality French match commentary training data for the Africa Cup of Nations 2025.
Dataset Details
Size: 2,000 training examples
Language: French
Format: JSONL (chat template)
Use Case: Fine-tuning LLMs for realistic African football match commentary
Event Distribution
82% General commentary
10% Goals
5% Substitutions
2% Penalties
1% Cards (yellow/red)
Teams Covered
Morocco, Senegal, Egypt, Nigeria, Côte… See the full description on the dataset page: https://huggingface.co/datasets/oxmo88/afcon2025-commentary.news-commentary-en-arThis is a filtered version of the English-to-Arabic News Commentary dataset available at data.statmt.org/news-commentary.
The filtering process includes removing duplicates, language detection, and semantic filtering based on similarity (>0.70) between the source and translation.
The filtering script is available at data-processing.ipynb.
generic_covas_commentary_v2Football-Commentarynews-commentary-cs-deFIFA_commentarygeneric_situational_event_commentary_v2baseball-commentary-datasetgeneric_covas_commentarynews_commentary
Dataset Card for OPUS News-Commentary
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/hfxunlp/news_commentary.hindi-tts-digital-commentary-no-demucsgeneric_situational_event_commentary_v1hindi-tts-digital-commentary
