datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
0xKobolds
0xKobolds — AI Coding Agent Sessions
A growing public dataset of real coding agent sessions from building 0xKobold — an open-source AI assistant framework built on pi with multi-agent orchestration, hot-reload skills, and local LLM support via Ollama.
What this is
Every session in this dataset is an unedited, redacted trace of me working with AI to build and debug 0xKobold. This includes:
Architecture design — multi-agent orchestration, event bus, extension… See the full description on the dataset page: https://huggingface.co/datasets/moikapy/0xKobolds.moi-myanmar-articles-lines
MOI Myanmar Articles Dataset - Lines (DatarrX/moi-myanmar-articles-lines)
Dataset Description
The MOI Myanmar Articles - Lines dataset is a derivative corpus created from the official articles published on the Ministry of Information (MOI) website of the Republic of the Union of Myanmar.
Unlike the main dataset (moi-myanmar-articles), which contains full-length article texts, this dataset has been systematically split line-by-line (sentence-by-sentence). This… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/moi-myanmar-articles-lines.moi-myanmar-articles
MOI Myanmar Articles Dataset (DatarrX/moi-myanmar-articles)
Dataset Description
The MOI Myanmar Articles dataset is a collection of official articles extracted directly from the Ministry of Information (MOI) website of the Republic of the Union of Myanmar. This dataset is curated exclusively to foster the growth, research, and development of the Myanmar (Burmese) language within the fields of Natural Language Processing (NLP) and Machine Learning (ML).… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/moi-myanmar-articles.moi-news-articles-dataset
MOI News & Article Dataset 🇲🇲
This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research.
The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ).
🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.MoistWeb-25k
MoistWeb 💦 - 25k
Samples from HuggingFaceFW/fineweb that contain the word 'moist'
rok_mois_opendata_reports
