CoolFace
Datasetpublic

imvladikon/english_news_weak_ner

Large Weak Labelled NER corpus Dataset Summary The dataset is generated through weak labelling of the scraped and preprocessed news corpus (bloomberg's news). so, only to research purpose. In order of the tokenization, news were splitted into sentences using nltk.PunktSentenceTokenizer (so, sometimes, tokenization might be not perfect) Usage from datasets import load_dataset articles_ds = load_dataset("imvladikon/english_news_weak_ner"… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/english_news_weak_ner.

sourceHugging Faceupdated 3y agoView on Hugging Face
5likes216downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
imvladikon/english_news_weak_ner · CoolFace