datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bc5cdr[Bio Creative 5 CDR NER dataset](https://academic.oup.com/database/article/doi/10.1093/database/baw032/2630271?login=true)conll2003[CoNLL 2003 NER dataset](https://aclanthology.org/W03-0419/)ontonotes5[ontonotes5 NER dataset](https://aclanthology.org/N06-2015/)customer-support-on-twitter-conversationTNews-classification
Dataset Card for "TNews-classification"
More Information needed
tweetner7[TweetNER7](TBA)wnut2017[WNUT 2017 NER dataset](https://aclanthology.org/W17-4418/)wikineural[wikineural](https://aclanthology.org/2021.findings-emnlp.215/)mit_restaurant[mit_restaurant NER dataset](https://groups.csail.mit.edu/sls/downloads/)bionlp2004[BioNLP2004 NER dataset](https://aclanthology.org/W04-1213.pdf)mit_movie_triviaMIT Moviemultinerd[MultiNERD](https://aclanthology.org/2022.findings-naacl.60/)tnews
Dataset Card for "tnews"
More Information needed
fin[FIN NER dataset](https://aclanthology.org/U15-1010.pdf)clue-tnewstnewstweebank_ner[Tweebank NER](https://arxiv.org/abs/2201.07281)t_newsbtc[BTC](https://aclanthology.org/C16-1111/)TNews
TNews
An MTEB dataset
Massive Text Embedding Benchmark
Short Text Classification for News
Task category
t2c
Domains
None
Reference
https://www.cluebenchmarks.com/introduce.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("TNews")
evaluator = mteb.MTEB([task])
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TNews.tnewsdemo_inctnewstnews_COTsaber-wa-chat-hi
Synthetic Hinglish chats between retailer and sales person
This is synthetic conversation in Hinglish generated by gpt-4o with temperature is 0.7 and top-n is 0.6. Each conversation contains at least 8 turns.
