CoolFace
Datasetpublic

k-chirkunov/jeeran_semeval_2016_arabic_adaptation_coarse

Jeeran — SemEval-2016 Task 5 (Arabic) adaptation, coarse categories The k-chirkunov/jeeran_semeval_2016_arabic_adaptation dataset with its 48 fine-grained Arabic aspect labels collapsed onto 14 coarse families. Everything else — targets, character offsets, polarity, the human-annotated spans, the train/test split — is carried over unchanged. Contents split reviews sentences opinions train 43,365 69,524 181,406 test 10,838 17,594 45,225 The… See the full description on the dataset page: https://huggingface.co/datasets/k-chirkunov/jeeran_semeval_2016_arabic_adaptation_coarse.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes28downloads
Dataset Card

Jeeran — SemEval-2016 Task 5 (Arabic) adaptation, coarse categories

The `k-chirkunov/jeeran_semeval_2016_arabic_adaptation` dataset with its 48 fine-grained Arabic aspect labels collapsed onto 14 coarse families. Everything else — targets, character offsets, polarity, the human-annotated spans, the train/test split — is carried over unchanged.

Contents

splitreviewssentencesopinions
train43,36569,524181,406
test10,83817,59445,225

The datasets view is flat: one row per opinion, with the sentence text repeated across the opinions it contains. SemEval SB1 XML is under semeval_xml/.

Fields

fielddescription
review_id, sentence_idsentence_id is "{review_id}:{n}", consecutive from 0
textthe sentence
targetopinion target expression — verbatim substring of text
from / tocharacter offsets of target in text
categorycoarse family, one of the 14 below
category_finethe original fine label this was mapped from (48 values)
polaritypositive / negative / neutral
span, span_from, span_tothe human-annotated evaluative span the target came from

Coarse taxonomy

#familycovers
1تقييم عام والجودةgeneral or holistic praise/criticism, a personal opinion or impression, a rating/comparison, overall quality, product/goods quality, device/equipment quality, advertising/promotion, notices/announcements, or collected opinions/surveys.
2توفر وتنوع البضاعةvariety or diversity of products/goods/models, product availability, stock, shortage, or abundance.
3الاحترافية والكفاءةprofessionalism, skill, expertise or experience, dedication and commitment to work, responsibility and accountability, honesty and integrity, or cheating/fraud.
4المعاملة والاهتمامhow staff or management TREAT and ATTEND TO customers, visitors, or patients — politeness, respect, helpfulness, reception, bedside manner — and the establishment's attentiveness, responsiveness, effort, and concern to meet people's needs, comfort, and wishes (including whether customers end up satisfied or pleased).
5السعرprice, cost, being expensive or cheap, or value for money.
6النظافةcleanliness, hygiene, tidiness, or care for the cleanliness of the place.
7المكانlocation, view, accessibility, surroundings, or the size/area/space of the place.
8الوقتpunctuality, appointments, delays, waiting time, or working/opening hours.
9الازدحام وعدد الموظفينcrowding or congestion, or the number of employees/staff.
10النظام والحمايةrules, organization, order, procedures, discipline, safety, security, protection, or guards.
11التوصية والتحذيرadvice, guidance, a recommendation or suggestion, or a warning about danger, risk, or harm.
12تهنئةcongratulations, thanks, appreciation, or blessing.
13تخصصsomething intended for a specific group, a specialization, or a dedicated section.
14توترstress, anxiety, tension, confusion, or discomfort.

Distribution

familyopinionsshare
تقييم عام والجودة93,78441.38%
السعر23,82210.51%
المعاملة والاهتمام22,1819.79%
توفر وتنوع البضاعة17,9447.92%
المكان17,9217.91%
الاحترافية والكفاءة14,7706.52%
النظافة11,4965.07%
الوقت6,7772.99%
النظام والحماية5,5782.46%
التوصية والتحذير4,0521.79%
الازدحام وعدد الموظفين3,0141.33%
تخصص2,5431.12%
تهنئة2,2621.00%
توتر4870.21%

Mapping provenance

The fine → coarse mapping is coarsen_aspect_category from the project's evaluation_coarse/prompts_override.py (verified identical to the ASPECT_CATEGORY_MERGE in evaluation/eval_utils.py), applied verbatim — no relabelling by hand or by model.

Two details worth knowing:

  • —The fine dataset contains 80 distinct category strings, not 48: 30 are whitespace variants ("السعر ", " الخبرة"). The mapping strips before lookup, so they normalize onto their 48 canonical labels; 365 rows were affected.
  • —1,677 opinions (0.73%) were dropped — fine label unknown or empty, which belong to no coarse family. Sentences and reviews left empty by that were dropped too, and sentence ids renumbered to stay gap-free.

Verification

Every released opinion satisfies text[from:to] == target, and the target lies within its own span. Checked on the CSVs, on the XML after reparsing, and after the Hub round-trip — 0 exceptions. No review or sentence id appears in both splits.

Caveats inherited from the fine dataset

target and category_fine are model-generated (Gemma); span boundaries and polarity are human. Opinions with no explicit target noun phrase were already dropped upstream, which removed speech-act labels preferentially — so نصيحة/ذم-derived families (التوصية والتحذير) are thinner here than in the source corpus.

Licensing

The underlying reviews were collected from Jeeran; this repository does not assert a license over them. Consult the source terms before redistribution or commercial use.