voidbeholder/medAbbreviationsRU
Dataset for acronym disambiguatuion on Russian-language biomedical texts Dataset Details This dataset forms part of the Master's dissertation carried out at Saint-Peterburg State University, Department of Computational and Applied Linguistics (to be defended on the 18.06.2024)Dissertation title: Automatic acronym disambiguation based on the Russian medical corpusAuthor: Polina GousyatskayaThis is the first attempt of acronym disambiguation on Russian material and… See the full description on the dataset page: https://huggingface.co/datasets/voidbeholder/medAbbreviationsRU.
Dataset for acronym disambiguatuion on Russian-language biomedical texts
Dataset Details
This dataset forms part of the Master's dissertation carried out at Saint-Peterburg State University, Department of Computational and Applied Linguistics (to be defended on the 18.06.2024) Dissertation title: Automatic acronym disambiguation based on the Russian medical corpus Author: Polina Gousyatskaya This is the first attempt of acronym disambiguation on Russian material and the first dataset of the kind available for the task.
Dataset Description
This dataset is aimed at automatic acronym disambiguation based on the Russian medical corpus and suited for text classification task. The dataset is structured in a basic tabular format, easily readable as pandas DataFrames.
Contents and structure
The dataset contains 75 ambiguous acronyms, all pertaining to biomedical domain. For each acronym sense we scraped a number of contexts (this number depends on the availability of each sense in the Internet texts, but all senses are represented by the sufficient number of contexts to ensure successful classification.) Contexts representing each sense of the ambiguous acronym are labeled.
Therefore, the table structure looks like this:
NB! The number of senses for each acronym varies throughout the dataset. There are 50 binary acronyms: АДГ, АЕ, АГ, АНФ, АТК, БА, БАТ, ВМС, ВПС, ВСА, ГБ, ГИП, ГР, ДФА, КФК, КК, КОА, КСР, ЛП, ЛТГ, МА, МКБ, МЛД, МО, МОС, МПА, МПБ, НОВ, ПА, ПДП, ПФ, ПГБ, ПГД, ПИР, ПЛР, ПВ, РФ, РПГА, РСК, РТ, САД, СЕ, СКО, СМА, ССС, СВЧ, ТТГ, УФР, УЗТ, ЭОП. 16 acronyms with three senses: АД, АГК, АР, БКК, ДОК, МС, МВЛ, НС, ОАА, ПМП, ПНП, ППГ, СА, СИ, СКП, ВКК. 3 acronyms with four senses: ДЭ, ЭД, ЭС. 2 acronyms with five senses: АС, ОВ. 2 acronyms with six senses: ДК, СД. 1 acronym with seven senses: СП. 1 acronym with nine senses: ПД. 1 acronym with eleven senses: ПГ.
Original use
In the paper we used SVM and RuBioBERT to classify contexts of the ambiguous acronyms to their respective num_senses. The SVM implementation reaches 93% accuracy and F1, RuBioBERT - 0.976% accuracy and F1.
Contacts
Polina Gousyatskaya, Saint-Petersburg, Russia polinagousyatskaya@gmail.com Telegram: @voidbeholder
