CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BNNT /mozi_general_instructions_3mSources are listed below: Chinese General Instruction 2000k BELLE https://huggingface.co/datasets/BelleGroup/train_2M_CN English generic instruction 52k alpaca-gpt4 https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM Chinese generic dialog instructions 800k BELLE https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M English Universal Dialog Instruction 94k sharegpt_vicuna https://huggingface.co/datasets/jeffwan/sharegpt_vicuna Chinese-English-Japanese Universal Command 49k… See the full description on the dataset page: https://huggingface.co/datasets/BNNT/mozi_general_instructions_3m.text1M<n<10M3 likes78 downloads3y agoHugging Face02BNNT /PatentMatchtext1K<n<10K3 likes76 downloads3y agoHugging Face03BNNT /IPQA QA evaluation dataset in intellectual property The IPQA contains questions in seven languages, and the 100 data items include 35 each in Chinese and English, and 6 each in Spanish, Japanese, German, French, and Russian. textquestion-answeringn<1K2 likes61 downloads3y agoHugging Face04thedeba /bnnews Bengali News Corpus (2023–2026) Dataset Summary The Bengali News Corpus (2023–2026) is a large-scale, high-quality monolingual Bengali dataset consisting of 295,020 cleaned news articles scraped from Prothom Alo (https://www.prothomalo.com), Bangladesh's largest Bengali-language daily newspaper. The dataset spans over 3.5 years of comprehensive news reporting (January 2023 to August 2026) across various domains including National News, Politics, World News… See the full description on the dataset page: https://huggingface.co/datasets/thedeba/bnnews.texttext-classification100K<n<1M0 likes59 downloads1mo agoHugging Face05BNNT /IPQuizThe IPQuiz dataset is used to assess a model's understanding of intellectual property-related concepts and regulations.IPQuiz is a multiple-choice question-response dataset collected from publicly available websites around the world in a variety of languages. For each question, the model needs to select an answer from a candidate list. source: http://epaper.iprchn.com/zscqb/h5/html5/2023-04/21/content_27601_7600799.htm… See the full description on the dataset page: https://huggingface.co/datasets/BNNT/IPQuiz.text1K<n<10K2 likes42 downloads3y agoHugging Face06sustcsenlp /bn_news_summarization Bengali Abstractive News Summarization (BANS) Dataset Summary Nowadays news or text summarization becomes very popular in the NLP field. Both the extractive and abstractive approaches of summarization are implemented in different languages. A significant amount of data is a primary need for any summarization. For the Bengali language, there are only a few datasets are available. Our dataset is made for Bengali Abstractive News Summarization (BANS) purposes. As abstractive… See the full description on the dataset page: https://huggingface.co/datasets/sustcsenlp/bn_news_summarization.textsummarization10K<n<100K1 likes30 downloads4y agoHugging Face07rejauldu /bn-number-v1 Bengali Numeric Dataset (bn-name-v1) Dataset Summary The Bengali Numeric Dataset provides a collection of Bengali number words and arithmetic question–answer pairs.It is designed for language modeling, tokenization experiments, and fine-tuning models on Bengali numerals.This dataset includes mappings of integers (0–100, 1000, 1 lakh, 1 crore) into their Bengali text representation and simple arithmetic expressions. This version (v1) contains train and validation splits… See the full description on the dataset page: https://huggingface.co/datasets/rejauldu/bn-number-v1.textquestion-answering100K<n<1M0 likes10 downloads1y agoHugging Face08Xmansh1t /bnntext1K<n<10K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.