CoolFace
Datasetpublic

uznlp-uz/uzMED-ABSA

uzMED-ABSA Dataset Summary uzMED-ABSA is an Uzbek medical-domain dataset for aspect-based sentiment analysis. The current release contains 7,500 annotated aspect-level rows in a single TSV file. Each row includes: an Uzbek medical or healthcare-related text sample an aspect term and aspect category a normalized aspect label an opinion expression linked to the aspect a sentiment label and numeric polarity score sentiment intensity and context type negation… See the full description on the dataset page: https://huggingface.co/datasets/uznlp-uz/uzMED-ABSA.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes8downloads
Dataset Card

uzMED-ABSA

Dataset Summary

uzMED-ABSA is an Uzbek medical-domain dataset for aspect-based sentiment analysis. The current release contains 7,500 annotated aspect-level rows in a single TSV file.

Each row includes:

  • an Uzbek medical or healthcare-related text sample
  • an aspect term and aspect category
  • a normalized aspect label
  • an opinion expression linked to the aspect
  • a sentiment label and numeric polarity score
  • sentiment intensity and context type
  • negation, sarcasm, and writing-system flags

The dataset covers 35 healthcare aspect categories, including doctor competence, doctor communication, nurse service, reception process, waiting time, diagnosis, treatment, medication, laboratory tests, equipment, hygiene, emergency care, surgery, inpatient care, privacy, safety, telemedicine, access, and overall recommendation.

Supported Tasks

  • aspect category classification
  • aspect term extraction
  • opinion expression extraction
  • sentiment classification
  • polarity scoring
  • negation detection
  • sarcasm / irony detection
  • writing-system analysis for Uzbek text

Files

  • uzmedabsa.tsv: main file in tab-separated format

Dataset Structure

Columns

ColumnTypeDescription
idstringStable sample identifier, e.g. MED-0001
textstringUzbek medical-domain text sample
aspect_termstringAspect term mentioned in the text
aspect_categorycategoricalHuman-readable aspect category in Uzbek
aspect_labelcategoricalNormalized aspect category code
opinion_expressionstringOpinion phrase or span related to the aspect
sentimentcategoricalSentiment label: POS, NEG, NEU, or MIX
polarity_scoreintegerPolarity score from 1 to 5
intensitycategoricalSentiment intensity: past, oʻrta, or yuqori
context_typecategoricalContext type, such as praise, complaint, contrast, direct evaluation, or irony
negationbinary0 or 1, whether negation is present
sarcasmbinary0 or 1, whether sarcasm or irony is present
scriptcategoricalWriting system label: lotin, kirill, or aralash
row_numberintegerSource row number from the original workbook

Tagsets

Sentiment Labels

LabelMeaningDescription
POSPositiveFavourable evaluation, satisfaction, or approval
NEGNegativeComplaint, dissatisfaction, harm, or negative evaluation
NEUNeutralFactual or emotionally neutral statement
MIXMixedContains both positive and negative signals

Polarity Score

ScoreMeaning
1Strong negative
2Negative
3Neutral or mixed
4Positive
5Strong positive

Intensity Labels

LabelMeaning
pastLow intensity
oʻrtaMedium intensity
yuqoriHigh intensity

Script Labels

LabelMeaning
lotinUzbek Latin script
kirillUzbek Cyrillic script
aralashMixed Latin/Cyrillic or mixed-script text

Aspect Categories

Aspect categoryAspect labelCount
Shifokor malakasiDOCTOR#COMPETENCE225
Shifokor muloqotiDOCTOR#COMMUNICATION225
Shifokorning muomalasiDOCTOR#ATTITUDE225
Hamshira xizmatiNURSE#SERVICE225
Qabul jarayoniRECEPTION#PROCESS225
Kutish vaqtiWAITING#TIME225
Tashxis aniqligiDIAGNOSIS#ACCURACY225
Davolash samaradorligiTREATMENT#EFFECTIVENESS225
Dori-darmon tavsiyasiMEDICATION#RECOMMENDATION225
Dori haqida tushuntirishMEDICATION#EXPLANATION225
Laboratoriya tahlillariLAB#TESTS210
Tibbiy jihozlarEQUIPMENT#QUALITY210
Klinik muhitCLINIC#ENVIRONMENT210
Tozalik va gigiyenaHYGIENE#CLEANLINESS210
Tinchlik va shovqinENVIRONMENT#NOISE210
Xodimlarning tezkorligiSTAFF#RESPONSIVENESS210
Favqulodda yordamEMERGENCY#CARE210
Ogʻriqni boshqarishPAIN#MANAGEMENT210
Jarrohlik/amaliyot sifatiSURGERY#QUALITY210
Statsionar parvarishINPATIENT#CARE210
Chiqaruv/discharge maʻlumotiDISCHARGE#INFORMATION210
Parvarish koordinatsiyasiCARE#COORDINATION210
Narx va toʻlovPRICE#PAYMENT210
ShaffoflikBILLING#TRANSPARENCY210
Sugʻurta/imtiyozINSURANCE#BENEFIT210
Maxfiylik va etik muomalaPRIVACY#ETHICS210
Bemor xavfsizligiPATIENT#SAFETY210
Bolalar/bogʻcha-pediatriya xizmatiPEDIATRICS#SERVICE210
Ayollar salomatligiWOMENS_HEALTH#SERVICE210
Onlayn yozilishONLINE#APPOINTMENT210
TelemeditsinaTELEMEDICINE#SERVICE210
Call-markazCALL_CENTER#SERVICE210
Manzil va kirish qulayligiACCESS#LOCATION210
Ovqatlanish/statsionar servisINPATIENT_FOOD#SERVICE210
Umumiy tavsiyaOVERALL#RECOMMENDATION210

Statistics

Overview

  • Rows: 7,500
  • Columns: 14
  • Split: train only
  • Aspect categories: 35
  • Aspect labels: 35
  • Unique aspect terms: 1,288
  • Unique text samples: 5,725
  • Unique IDs: 5,786
  • IDs with multiple rows: 1,477 (multi-aspect annotations)
  • Extra rows from repeated IDs: 1,714
  • Exact duplicate rows: 0
  • Repeated (text, aspect_term, aspect_label) annotations: 5
  • Text length: 39-307 characters, mean 136.4
  • Text length: 5-39 words, mean 17.1

Sentiment Distribution

LabelCount
POS3,052
NEG2,940
NEU1,003
MIX505

Polarity Distribution

ScoreCount
11,007
21,933
31,508
42,104
5948

Sentiment by Polarity Score

Sentiment12345
MIX0050500
NEG1,0071,933000
NEU001,00300
POS0002,104948

Intensity Distribution

IntensityCount
oʻrta4,521
yuqori1,685
past1,294

Script Distribution

ScriptCount
lotin6,040
kirill801
aralash659

Negation

ValueCount
0 (no negation)6,676
1 (negation)824

Sarcasm

ValueCount
0 (not sarcastic)7,298
1 (sarcastic)202

Context Type Distribution

Context typeCount
kontrast1,603
bevosita1,234
sabab-oqibat372
kontrast/koʻp aspektli253
koʻp aspektli221
shikoyat208
umumiy199
kontrastli koʻp-aspektli post185
qiyosiy171
minnatdorchilik168
tajriba asosida161
tajriba158
kontrastli154
bevosita baho154
aniq142
aralash137
neytral kuzatuv126
kontrastli baho117
tavsifiy100
salbiy tajriba98
taqqoslash96
kinoyali82
umumiy baho79
inkorli76
inkor76
bevosita tajriba70
neytral70
shartli kontekst68
koʻp-aspektli67
koʻp-aspektli baho66
baholovchi63
aniq baho61
maqtov60
aralash baho57
taqqoslashli baho52
sabab-oqibatli baho51
tajriba asosidagi baho49
kinoyaviy baho47
neytral/kuzatuv44
oddiy38
inkorli baho37
koʻp-aspektli post30
shartli24
kinoya22
koʻp aspektli post20
kinoyali shikoyat19
kinoyaviy16
sababiy16
savol15
inkorli-kontrastli baho13
tajriba bayoni12
kuzatuv10
kinoyali baho10
maʻlumot beruvchi9
faktik8
kinoyali kontekst6

Normalization

  • Column names were converted to English snake_case.
  • Text fields are UTF-8 encoded and Unicode-normalized.
  • Whitespace was stripped and collapsed to single spaces.
  • Uzbek apostrophe letters were normalized to ʻ (U+02BB).
  • Sentiment and aspect labels were canonicalized.
  • Numeric flags and polarity scores are stored as integers.
  • Embedded tabs and newlines were removed from cell values before TSV export.

Usage

python
from datasets import load_dataset

dataset = load_dataset(
    "csv",
    data_files={"train": "uzmedabsa.tsv"},
    delimiter="\t",
    encoding="utf-8",
)

print(dataset["train"][0])

Or directly from the Hub:

python
from datasets import load_dataset

dataset = load_dataset("uznlp-uz/uzMED-ABSA")
print(dataset["train"][0])

Citation

If you use uzMED-ABSA in your research, please cite this dataset.

License

This dataset is released under the CC-BY-4.0 license.