datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Myanmar-English-general-text-translation
🇲🇲-🇬🇧 Myanmar-English General Text Translation Dataset
📚 Dataset Overview
This dataset is a high-quality, parallel corpus designed for training robust and accurate Myanmar-English Machine Translation (MT) models. It focuses on General Domain texts, covering a wide range of everyday scenarios, literature, conversations, and descriptive narratives.
Our primary goal in creating this dataset is to provide a clean, reliable resource to enhance the performance of… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/Myanmar-English-general-text-translation.text_testdata_general_jailbreak_dt
Dataset Card for "testdata_general_jailbreak_dt"
More Information needed
GeneralTextCorpus
Mixed Content Dataset
Description:This dataset contains a diverse collection of text from multiple domains, including general knowledge, cooking, articles, and more. Each entry typically includes text content along with metadata such as source, title, and language.
The dataset is structured to support research, analysis, or training of NLP models on varied textual content.
Data Structure:Each item typically contains:
id: Unique identifier
text: Main text content
meta: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/ademchaoua/GeneralTextCorpus.GENERAL_PURPOSE_TEXTGeneral-Astronomy-TextbookGeneral_Assembly_Votes_and_Resolution_textsGeneral_Text_Data
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Aniket Kumar]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [Indian Language]
License: [Free to use]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/Threatthriver/General_Text_Data.llm_general_textcarigold_general_chat_text_datasetText data from Carigold forum replies based on General Chat section (https://carigold.com/forum/forums/general-chat.174/)
Language = Malay + English mixed
General_Knowledge_Text-Image_Pair_Corpus
ID
King-IM-104
Quantity
2,000,000 Sets
Image Specification
2K
Text Specification
Includes labels, descriptions in both Chinese and English
Description
Product Features: This corpus includes data from 23 categories such as cuisine, landscapes, architecture, cities, countryside, health, sports, medical, automobiles, backgrounds, finance, education, oil paintings, illustrations, watercolors, travel, fashion, romance, animals, plants, space… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/General_Knowledge_Text-Image_Pair_Corpus.general-text-understanding-datageneral-instruction-text-setgeneral-instruction-text-setrgeneral-text-learning-samplesGeneralTextSetgeneral-text-representation-datageneral-text-training-corpusGeneralTextArchivegeneral-meaning-textsGeneralText-Recentgeneral-text-reasoning-samples
