CoolFace
Datasetpublicgated

MikCil/disambiguation-pretraining-v2

Disambiguation Pretraining v2 Disambiguation Pretraining v2 is an English-Italian corpus for pretraining language models on lexical semantics and word-sense disambiguation. It connects lemmas, definitions, attested contexts, contrasting senses, parts of speech, and subject domains through varied natural-language formulations. Version 2 contains 3,693,921 rows, compared with 900,009 in the first release. The increase comes from new semantic tasks and controlled reuse of source… See the full description on the dataset page: https://huggingface.co/datasets/MikCil/disambiguation-pretraining-v2.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes8downloads

MikCil/disambiguation-pretraining-v2 · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.