gsarti/magpie
The MAGPIE corpus is a large sense-annotated corpus of potentially idiomatic expressions (PIEs), based on the British National Corpus (BNC). Potentially idiomatic expressions are like idiomatic expressions, but the term also covers literal uses of idiomatic expressions, such as 'I leave work at the end of the day.' for the idiom 'at the end of the day'. This version of the dataset reflects the filtered subset used by Dankers et al. (2022) in their investigation on how PIEs are represented by NMT models. Authors use 37k samples annotated as fully figurative or literal, for 1482 idioms that contain nouns, numerals or adjectives that are colours (which they refer to as keywords). Because idioms show syntactic and morphological variability, the focus is mostly put on nouns. PIEs and their context are separated using the original corpus’s word-level annotations.
Fix task tags (#2)
Fix `license` metadata (#1)
Update README.md
Update README.md
Update magpie.py
Delete keywords.tsv
Update magpie.py
Update magpie.py
Update keywords.tsv
Update magpie.py
Update magpie.py
Update magpie.py
Update magpie.py
Update magpie.py
Update magpie.py
Upload keywords.tsv
Upload magpie.tsv
Create magpie.py
initial commit
