datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task982_pib_translation_tamil_bengali
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task982_pib_translation_tamil_bengali
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task982_pib_translation_tamil_bengali.task1009_pib_translation_bengali_hindi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1009_pib_translation_bengali_hindi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1009_pib_translation_bengali_hindi.B-CORE-bengali-corpus
B-CORE: Bangla Pretraining Corpus
B-CORE (Bengali Context-aware Optimized and Refined Entities) is a large-scale, rigorously curated Bangla monolingual corpus for language model pretraining, comprising 16.5 million documents (4.32 billion tokens, 52GB (20.8 GB Compressed)). It is among the largest and most carefully curated Bangla pretraining corpora available, constructed through a reproducible multi-stage pipeline.
B-CORE was used to pretrain the BnLM-F and BnLM-C Bengali… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/B-CORE-bengali-corpus.shangkhachil-bengali-public-domain
Bengali Public-Domain Literature
101 complete works by 21 authors,
11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09.
Where these texts are read
https://shangkhachil.com — the reading site this corpus was built for. Free, no
account, 246 works by 28 authors. The complete text of
every work in this file can be read there.
This file is the text. The site is the part a JSONL cannot be:
Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.bengali-wikipedia
📚 Bengali Wikipedia Language Modeling Dataset (For GPT-2 Training)
📝 Dataset Summary
This dataset contains a large Bengali text corpus collected from Bengali Wikipedia.It is cleaned, sentence-segmented, and formatted for next-token prediction language modeling tasks such as GPT-2 training.
It includes train and validation splits, suitable for transformer-based Bengali language models.
📊 Dataset Details
Property
Value
Language
Bengali (বাংলা)… See the full description on the dataset page: https://huggingface.co/datasets/rejauldu/bengali-wikipedia.bengalichat
Dataset Card for Bengali Chat
We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved.
Dedicated towards addressing this problem, I release 2 new datasets rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for supervised fine-tuning (SFT) to… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/bengalichat.task1493_bengali_geopolitical_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1493_bengali_geopolitical_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1493_bengali_geopolitical_hate_speech_binary_classification.pratilipi-bengali-webscrape
Pratilipi Bengali Literature Archive
Overview
This repository contains a large-scale text dataset scraped from bengali.pratilipi.com, a leading storytelling and self-publishing platform for Bengali literature. The primary goal of this archive is to preserve a vast collection of purely human-written Bengali fiction, serials, poems, and essays, creating a distinct record of human creativity and storytelling.
Purpose and Usage
This dataset is published… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/pratilipi-bengali-webscrape.smolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
task1490_bengali_personal_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1490_bengali_personal_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1490_bengali_personal_hate_speech_binary_classification.task1497_bengali_book_reviews_sentiment_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1497_bengali_book_reviews_sentiment_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1497_bengali_book_reviews_sentiment_classification.all_combined_bengali_252k
Dataset Card for all_combined_bengali_252K
Dataset Summary
This dataset is a mix of Bengali instruction sets translated from open-source instruction sets:
Dolly,
Alpaca,
ChatDoctor,
Roleplay
GSM
In this dataset Bengali instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Bengali
Dataset Structure
JSON
Data Fields
output (string)
data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.task995_pib_translation_bengali_english
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task995_pib_translation_bengali_english
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task995_pib_translation_bengali_english.task1496_bengali_reviews_sentiment_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1496_bengali_reviews_sentiment_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1496_bengali_reviews_sentiment_classification.task1494_bengali_hate_speech_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1494_bengali_hate_speech_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1494_bengali_hate_speech_classification.smolified-bengali-physics-teacher
🤏 smolified-bengali-physics-teacher
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model RohanSardar/smolified-bengali-physics-teacher.
📦 Asset Details
Origin: Smolify Foundry (Job ID: f26e2704)
Records: 9493
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by RohanSardar.
Generated via Smolify.ai.
bengali_chat_conv
Bengali Chat Conversation Dataset
Dataset Overview
The Bengali Chat Conversation dataset contains a collection of conversational exchanges in Bengali. Each entry consists of a question and its corresponding answer, covering a wide range of topics including technology, health, environment, education, and more. This dataset can be used for various natural language processing (NLP) tasks such as language modeling, question-answering systems, and conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/spedrox-sac/bengali_chat_conv.task1492_bengali_religious_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1492_bengali_religious_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1492_bengali_religious_hate_speech_binary_classification.bengali-poemssoda_bengali_smalltask1062_pib_translation_marathi_bengali
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1062_pib_translation_marathi_bengali
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1062_pib_translation_marathi_bengali.bengali-hindi-number-blindspot
Blind Spots of Frontier Models: Bengali & Hindi Number Word-to-Digit Conversion
Summary
This dataset documents a critical blind spot in small open-source language models: failure to correctly convert Bengali and Hindi number words into their digit equivalents. Bengali and Hindi share the South Asian number system (hazar/হাজার, lakh/লাখ, crore/কোটি), and all three tested models consistently fail at this fundamental conversion step.
IMPORTANT: Arithmetic calculation errors… See the full description on the dataset page: https://huggingface.co/datasets/wong132/bengali-hindi-number-blindspot.Shikha-Bengali-Alpaca
Shikha-Bengali-Alpaca
Dataset Description
Shikha-Bengali-Alpaca হলো একটি বাংলা ইনস্ট্রাকশন-ফলোয়িং (Instruction-Following) ডেটাসেট। এই ডেটাসেটটি মূলত লার্জ ল্যাঙ্গুয়েজ মডেলকে (LLM) বাংলায় আরও ভালোভাবে প্রশ্নের উত্তর দিতে এবং মানুষের নির্দেশনা বুঝতে শেখানোর জন্য তৈরি করা হয়েছে।
এই প্রজেক্টটির মূল ভিত্তি হলো বিখ্যাত 'Alpaca' ডেটাসেট, যা পরবর্তীতে পরিমার্জন ও নতুন ডেটা সংযুক্তির মাধ্যমে এই 'Shikha' (শিখা) ভার্সনটি তৈরি করা হয়েছে।
Language: Bengali (বাংলা)
Creator:… See the full description on the dataset page: https://huggingface.co/datasets/shuvo-halder/Shikha-Bengali-Alpaca.task1491_bengali_political_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1491_bengali_political_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1491_bengali_political_hate_speech_binary_classification.task996_pib_translation_english_bengali
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task996_pib_translation_english_bengali
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task996_pib_translation_english_bengali.Roleplay-Bengali
RolePlay-Bengali
Roleplay-Bengali Dataset is a dataset for roleplaying in the Bengali language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Bengali.task1061_pib_translation_bengali_marathi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1061_pib_translation_bengali_marathi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1061_pib_translation_bengali_marathi.bengali-sft-v1
Bengali SFT Dataset (bengali-sft-v1)
এটি একটি ছোট বাংলা Instruction-Response ডেটাসেট, Supervised Fine-Tuning (SFT) এর জন্য তৈরি।
Dataset Summary
ভাষা: বাংলা (Bengali)
ফরম্যাট: instruction + output
উদাহরণ সংখ্যা: ১২০টি
Dataset Structure
Column
Description
instruction
ব্যবহারকারীর প্রশ্ন/নির্দেশ
output
উত্তর/রেসপন্স
How to use
from datasets import load_dataset
ds = load_dataset("worldjit/bengali-sft-v1")
Bengali_bhasini_datasetsmolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
