datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.pao-sentences-dataset
Pa'O Sentences Dataset is a text corpus for Pa'O language (ပအိုဝ်ႏ) containing structured line-by-line sentences designed for NLP, LLM pre-training, and machine translation.
📝 Pa'O Sentences Dataset (ပအိုဝ်ႏ လိက်လာႏငေါဝ်းရဲဉ်ႏ ရွမ်ခြွဉ်းဗူႏ)
📌 Project Summary (ထာꩻမာꩻခြပ်ရဲဉ်ႏ နပ်ထွားရဲပ်အအဲဉ်ႏ)
The Pa'O Sentences Dataset is an open-source textual corpus developed to support Natural Language Processing (NLP), Large Language Model (LLM) pre-training, Machine… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-sentences-dataset.satsec-decomposition
SatSec Grounded Objective-Decomposition Dataset
Version 2.0 is a leakage-controlled replacement for the original dataset used in
A Controlled Candidate-Set Benchmark for Offline Satellite-Security Plan
Decomposition
(DOI 10.48550/arXiv.2607.26371).
It contains 24 authored full decompositions and 83 mechanically derived next-step rows
across 24 cases.
There are 82 train rows and 25 test rows; the six test cases never occur in training.
Important v2 correction
The… See the full description on the dataset page: https://huggingface.co/datasets/paolocmo/satsec-decomposition.diffusion-generated-text
Diffusion-Generated Text Benchmark
17,565 cleaned responses from three diffusion language model families and 21 generation settings
This benchmark supports research on diffusion-generated language, machine-generated text detection, and robustness across model families and decoding configurations. It includes outputs from DiffusionGemma, LLaDA-8B-Instruct, and LLaDA2-mini with varied generation lengths and block sizes.
Benchmark composition
Generator… See the full description on the dataset page: https://huggingface.co/datasets/paoche11/diffusion-generated-text.
