datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mozi_general_instructions_3mSources are listed below:
Chinese General Instruction 2000k BELLE https://huggingface.co/datasets/BelleGroup/train_2M_CN
English generic instruction 52k alpaca-gpt4 https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
Chinese generic dialog instructions 800k BELLE https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M
English Universal Dialog Instruction 94k sharegpt_vicuna https://huggingface.co/datasets/jeffwan/sharegpt_vicuna
Chinese-English-Japanese Universal Command 49k… See the full description on the dataset page: https://huggingface.co/datasets/BNNT/mozi_general_instructions_3m.PatentMatchIPQA
QA evaluation dataset in intellectual property
The IPQA contains questions in seven languages, and the 100 data items include 35 each in Chinese and English, and 6 each in Spanish, Japanese, German, French, and Russian.
IPQuizThe IPQuiz dataset is used to assess a model's understanding of intellectual property-related concepts and regulations.IPQuiz is a multiple-choice question-response dataset collected from publicly available websites around the world in a variety of languages. For each question, the model needs to select an answer from a candidate list.
source:
http://epaper.iprchn.com/zscqb/h5/html5/2023-04/21/content_27601_7600799.htm… See the full description on the dataset page: https://huggingface.co/datasets/BNNT/IPQuiz.bn_news_summarization
Bengali Abstractive News Summarization (BANS)
Dataset Summary
Nowadays news or text summarization becomes very popular in the NLP field. Both the extractive and abstractive approaches of summarization are implemented in different languages. A significant amount of data is a primary need for any summarization. For the Bengali language, there are only a few datasets are available. Our dataset is made for Bengali Abstractive News Summarization (BANS) purposes. As abstractive… See the full description on the dataset page: https://huggingface.co/datasets/sustcsenlp/bn_news_summarization.bnn
