CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lots-of-LoRAs /task982_pib_translation_tamil_bengali Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task982_pib_translation_tamil_bengali Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task982_pib_translation_tamil_bengali.texttext-generation1K<n<10K0 likes120 downloads2y agoHugging Face02Lots-of-LoRAs /task1009_pib_translation_bengali_hindi Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1009_pib_translation_bengali_hindi Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1009_pib_translation_bengali_hindi.texttext-generation1K<n<10K0 likes120 downloads2y agoHugging Face03nahid-hub /B-CORE-bengali-corpus B-CORE: Bangla Pretraining Corpus B-CORE (Bengali Context-aware Optimized and Refined Entities) is a large-scale, rigorously curated Bangla monolingual corpus for language model pretraining, comprising 16.5 million documents (4.32 billion tokens, 52GB (20.8 GB Compressed)). It is among the largest and most carefully curated Bangla pretraining corpora available, constructed through a reproducible multi-stage pipeline. B-CORE was used to pretrain the BnLM-F and BnLM-C Bengali… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/B-CORE-bengali-corpus.texttext-generation10M<n<100M0 likes97 downloads2mo agoHugging Face04mir178 /shangkhachil-bengali-public-domain Bengali Public-Domain Literature 101 complete works by 21 authors, 11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09. Where these texts are read https://shangkhachil.com — the reading site this corpus was built for. Free, no account, 246 works by 28 authors. The complete text of every work in this file can be read there. This file is the text. The site is the part a JSONL cannot be: Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.tabulartext-generationn<1K0 likes72 downloads16d agoHugging Face05rejauldu /bengali-wikipedia 📚 Bengali Wikipedia Language Modeling Dataset (For GPT-2 Training) 📝 Dataset Summary This dataset contains a large Bengali text corpus collected from Bengali Wikipedia.It is cleaned, sentence-segmented, and formatted for next-token prediction language modeling tasks such as GPT-2 training. It includes train and validation splits, suitable for transformer-based Bengali language models. 📊 Dataset Details Property Value Language Bengali (বাংলা)… See the full description on the dataset page: https://huggingface.co/datasets/rejauldu/bengali-wikipedia.texttext-generation100K<n<1M0 likes69 downloads11mo agoHugging Face06rishiraj /bengalichat Dataset Card for Bengali Chat We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved. Dedicated towards addressing this problem, I release 2 new datasets rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for supervised fine-tuning (SFT) to… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/bengalichat.texttext-generation10K<n<100K4 likes43 downloads3y agoHugging Face07Lots-of-LoRAs /task1493_bengali_geopolitical_hate_speech_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1493_bengali_geopolitical_hate_speech_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1493_bengali_geopolitical_hate_speech_binary_classification.texttext-generation1K<n<10K0 likes33 downloads2y agoHugging Face08sayurio /pratilipi-bengali-webscrape Pratilipi Bengali Literature Archive Overview This repository contains a large-scale text dataset scraped from bengali.pratilipi.com, a leading storytelling and self-publishing platform for Bengali literature. The primary goal of this archive is to preserve a vast collection of purely human-written Bengali fiction, serials, poems, and essays, creating a distinct record of human creativity and storytelling. Purpose and Usage This dataset is published… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/pratilipi-bengali-webscrape.imagetext-generation1K<n<10K1 likes33 downloads6mo agoHugging Face09smolify /smolified-bengali-local-food-guide 🤏 smolified-bengali-local-food-guide Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-bengali-local-food-guide. 📦 Asset Details Origin: Smolify Foundry (Job ID: 638d3b25) Records: 1050 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes31 downloads6mo agoHugging Face10Lots-of-LoRAs /task1490_bengali_personal_hate_speech_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1490_bengali_personal_hate_speech_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1490_bengali_personal_hate_speech_binary_classification.texttext-generation1K<n<10K0 likes30 downloads2y agoHugging Face11Lots-of-LoRAs /task1497_bengali_book_reviews_sentiment_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1497_bengali_book_reviews_sentiment_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1497_bengali_book_reviews_sentiment_classification.texttext-generation1K<n<10K0 likes30 downloads2y agoHugging Face12OdiaGenAI /all_combined_bengali_252k Dataset Card for all_combined_bengali_252K Dataset Summary This dataset is a mix of Bengali instruction sets translated from open-source instruction sets: Dolly, Alpaca, ChatDoctor, Roleplay GSM In this dataset Bengali instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM) Languages Bengali Dataset Structure JSON Data Fields output (string) data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.texttext-generation100K<n<1M10 likes29 downloads3y agoHugging Face13Lots-of-LoRAs /task995_pib_translation_bengali_english Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task995_pib_translation_bengali_english Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task995_pib_translation_bengali_english.texttext-generation1K<n<10K0 likes27 downloads2y agoHugging Face14Lots-of-LoRAs /task1496_bengali_reviews_sentiment_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1496_bengali_reviews_sentiment_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1496_bengali_reviews_sentiment_classification.texttext-generation1K<n<10K0 likes27 downloads2y agoHugging Face15Lots-of-LoRAs /task1494_bengali_hate_speech_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1494_bengali_hate_speech_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1494_bengali_hate_speech_classification.texttext-generation1K<n<10K0 likes26 downloads2y agoHugging Face16RohanSardar /smolified-bengali-physics-teacher 🤏 smolified-bengali-physics-teacher Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model RohanSardar/smolified-bengali-physics-teacher. 📦 Asset Details Origin: Smolify Foundry (Job ID: f26e2704) Records: 9493 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by RohanSardar. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes24 downloads6mo agoHugging Face17spedrox-sac /bengali_chat_conv Bengali Chat Conversation Dataset Dataset Overview The Bengali Chat Conversation dataset contains a collection of conversational exchanges in Bengali. Each entry consists of a question and its corresponding answer, covering a wide range of topics including technology, health, environment, education, and more. This dataset can be used for various natural language processing (NLP) tasks such as language modeling, question-answering systems, and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/spedrox-sac/bengali_chat_conv.texttext-generation1K<n<10K1 likes20 downloads2y agoHugging Face18Lots-of-LoRAs /task1492_bengali_religious_hate_speech_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1492_bengali_religious_hate_speech_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1492_bengali_religious_hate_speech_binary_classification.texttext-generation1K<n<10K0 likes19 downloads2y agoHugging Face19rishiraj /bengali-poemstexttext-generation1K<n<10K1 likes18 downloads2y agoHugging Face20abirmondalind /soda_bengali_smalltexttext-generation1K<n<10K0 likes18 downloads3mo agoHugging Face21Lots-of-LoRAs /task1062_pib_translation_marathi_bengali Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1062_pib_translation_marathi_bengali Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1062_pib_translation_marathi_bengali.texttext-generation1K<n<10K0 likes17 downloads2y agoHugging Face22wong132 /bengali-hindi-number-blindspot Blind Spots of Frontier Models: Bengali & Hindi Number Word-to-Digit Conversion Summary This dataset documents a critical blind spot in small open-source language models: failure to correctly convert Bengali and Hindi number words into their digit equivalents. Bengali and Hindi share the South Asian number system (hazar/হাজার, lakh/লাখ, crore/কোটি), and all three tested models consistently fail at this fundamental conversion step. IMPORTANT: Arithmetic calculation errors… See the full description on the dataset page: https://huggingface.co/datasets/wong132/bengali-hindi-number-blindspot.texttext-generationn<1K0 likes17 downloads7mo agoHugging Face23shuvo-halder /Shikha-Bengali-Alpaca Shikha-Bengali-Alpaca Dataset Description Shikha-Bengali-Alpaca হলো একটি বাংলা ইনস্ট্রাকশন-ফলোয়িং (Instruction-Following) ডেটাসেট। এই ডেটাসেটটি মূলত লার্জ ল্যাঙ্গুয়েজ মডেলকে (LLM) বাংলায় আরও ভালোভাবে প্রশ্নের উত্তর দিতে এবং মানুষের নির্দেশনা বুঝতে শেখানোর জন্য তৈরি করা হয়েছে। এই প্রজেক্টটির মূল ভিত্তি হলো বিখ্যাত 'Alpaca' ডেটাসেট, যা পরবর্তীতে পরিমার্জন ও নতুন ডেটা সংযুক্তির মাধ্যমে এই 'Shikha' (শিখা) ভার্সনটি তৈরি করা হয়েছে। Language: Bengali (বাংলা) Creator:… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-halder/Shikha-Bengali-Alpaca.texttext-generation10K<n<100K0 likes16 downloads2d agoHugging Face24Lots-of-LoRAs /task1491_bengali_political_hate_speech_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1491_bengali_political_hate_speech_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1491_bengali_political_hate_speech_binary_classification.texttext-generation1K<n<10K0 likes14 downloads2y agoHugging Face25Lots-of-LoRAs /task996_pib_translation_english_bengali Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task996_pib_translation_english_bengali Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task996_pib_translation_english_bengali.texttext-generation1K<n<10K0 likes13 downloads2y agoHugging Face26jojo-ai-mst /Roleplay-Bengali RolePlay-Bengali Roleplay-Bengali Dataset is a dataset for roleplaying in the Bengali language for Large Language Model. The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, it can be found at this github… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Bengali.texttext-generation1K<n<10K0 likes11 downloads2y agoHugging Face27Lots-of-LoRAs /task1061_pib_translation_bengali_marathi Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1061_pib_translation_bengali_marathi Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1061_pib_translation_bengali_marathi.texttext-generation1K<n<10K0 likes11 downloads2y agoHugging Face28worldjit /bengali-sft-v1 Bengali SFT Dataset (bengali-sft-v1) এটি একটি ছোট বাংলা Instruction-Response ডেটাসেট, Supervised Fine-Tuning (SFT) এর জন্য তৈরি। Dataset Summary ভাষা: বাংলা (Bengali) ফরম্যাট: instruction + output উদাহরণ সংখ্যা: ১২০টি Dataset Structure Column Description instruction ব্যবহারকারীর প্রশ্ন/নির্দেশ output উত্তর/রেসপন্স How to use from datasets import load_dataset ds = load_dataset("worldjit/bengali-sft-v1") texttext-generationn<1K1 likes10 downloads1mo agoHugging Face29Akshayana26 /Bengali_bhasini_datasettexttranslation1K<n<10K1 likes3 downloads2y agoHugging Face30Prithhh2005 /smolified-bengali-local-food-guide 🤏 smolified-bengali-local-food-guide Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-bengali-local-food-guide. 📦 Asset Details Origin: Smolify Foundry (Job ID: 638d3b25) Records: 1050 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes2 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.