datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipfs_laos_laws
Laws of Laos
Research snapshot of official legislation collected from Lao Official Gazette.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-22
Coverage
catalog-backed incomplete
Source
Lao Official Gazette
Collector
scrapers/collect_la.py
Laws / instruments
5111
Articles
34506
Language
lo
Jurisdiction
Laos
License
la-official-gazette
Contents… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_laos_laws.temporal-laom-checkpointsipfs_laos_laws_ir
Laos legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_laos_laws (revision cec2cb5088d0567b8ce0698fd6c7a323d983fc0d) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Laos prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_laos_laws_ir.MMDU
📢 News
[06/13/2024] 🚀 We release our MMDU benchmark and MMDU-45k instruct tunning data to huggingface.
💎 MMDU Benchmark
To evaluate the multi-image multi-turn dialogue capabilities of existing models, we have developed the MMDU Benchmark. Our benchmark comprises 110 high-quality multi-image multi-turn dialogues with more than 1600 questions, each accompanied by detailed long-form answers. Previous benchmarks typically involved only single images or a small number of… See the full description on the dataset page: https://huggingface.co/datasets/laolao77/MMDU.laos-voice-dataset-v2LaoBenchMATOfficial implement of Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning
Github Repo: https://github.com/Liuziyu77/Visual-RFT/tree/main/Visual-ARFT
More infomation refer to our Github Repo and paper.
lao_stt_training_data
Lao Speech-to-Text Training Data
ຊຸດຂໍ້ມູນນີ້ຖືກຈັດກຽມຂຶ້ນມາເພື່ອໃຊ້ສຳລັບການເທຣນ ແລະ ປັບແຕ່ງ (Fine-tuning) ໂມເດວ Speech-to-Text (ເຊັ່ນ OpenAI Whisper) ສຳລັບພາສາລາວ.
ໂຄງສ້າງຂອງຂໍ້ມູນ (Dataset Structure)
Train set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ train/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ train.csv
Validation set: ໄຟລ໌ສຽງຢູ່ໃນໂຟນເດີ validation/ ແລະ ມີການ Mapping ຂໍ້ຄວາມໃນ validation.csv
ຮູບແບບຂໍ້ມູນໃນໄຟລ໌ CSV:
audio: ເສັ້ນທາງໄປຫາໄຟລ໌ສຽງ (e.g., train/audio25000.wav)… See the full description on the dataset page: https://huggingface.co/datasets/KitTzk/lao_stt_training_data.ipynbViRFT_COCO_base65ViRFT_COCO_8_cate_4_shotlao-speech-datasetnews-lao-classification
News_lao_Classification
Deduplicated copy of kornwtp/news-lao-classification.
Splits
split
rows
test
3,062
train
9,161
validation
3,062
laouenan-notable-peopleLaouenan, M., Bhargava, P., Eymeoud, J.-B., Gergaud, O., Plique, G., & Wasmer, E. (2023). A Brief History of Human Time - Cross-verified Dataset. data.sciencespo. doi: 10.21410/7E4/RDAG3O
OpenDao-LaoZiAgent
老子智能体 SKILL
基于 LangChain + Chroma 向量数据库的 RAG 问答系统,以春秋时期思想家老子(李耳)的身份回答关于《道德经》和道家思想的问题。
全免费方案 · 无需 API Key · 启动即跑
功能
💬 问道老子 — 以老子身份回答,引用《道德经》原文,返回来源
📖 版本对比 — 对比同一章节的多版本原文(帛书甲本、王弼本等)
🧭 主题导航 — 按思想主题检索(道论、德论、无为而治等)
对话记忆 — 支持多轮对话,保持上下文
技术栈
组件
选型
说明
Embedding
BAAI/bge-base-zh-v1.5
中文优化,本地加载,免费
LLM
Qwen/Qwen2.5-72B-Instruct
HuggingFace 免费推理 API
向量库
Chroma
轻量级,嵌入式
框架
LangChain
RAG 管线
前端
Gradio
Web 聊天界面
无需任何 API Key,启动即用。
快速开始… See the full description on the dataset page: https://huggingface.co/datasets/aaroncxxx/OpenDao-LaoZiAgent.lao-asr-thesis-datasetseniortalk
SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors
Introduction
SeniorTalk is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on elderly aged 75 to 85. This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), speaker dirazation (SD), speech editing and other… See the full description on the dataset page: https://huggingface.co/datasets/laolaiduowangshi/seniortalk.Image_Restoration_Datasetsnews-lao-clustering
news-lao-clustering
Deduplicated copy of kornwtp/news-lao-clustering,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/news-lao-clustering
Deduplicated on: 2026-09-04
Task type: clustering
Splits: test, validation
What changed
A text appearing under more than one gold cluster is collapsed to ONE row carrying the competing cluster ids in a new conflict_label column (NULL elsewhere). labels on that row holds the first-seen value as a placeholder --… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/news-lao-clustering.sib200-lao-clustering
sib200-lao-clustering
Deduplicated copy of kornwtp/sib200-lao-clustering,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/sib200-lao-clustering
Deduplicated on: 2026-09-04
Task type: clustering
Splits: test, validation
What changed
A text appearing under more than one gold cluster is collapsed to ONE row carrying the competing cluster ids in a new conflict_label column (NULL elsewhere). labels on that row holds the first-seen value as a… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/sib200-lao-clustering.ViRFT_COCOk-vaultalt-tha-lao-bitextmining
alt-tha-lao-bitextmining
Deduplicated copy of kornwtp/alt-tha-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-tha-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-tha-lao-bitextmining.sib200-lao-classification
SIB200_lao_Classification
Deduplicated copy of kornwtp/sib200-lao-classification.
Splits
split
rows
test
204
train
701
validation
99
alt-vie-lao-bitextmining
alt-vie-lao-bitextmining
Deduplicated copy of kornwtp/alt-vie-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-vie-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-vie-lao-bitextmining.laos-speech-datasetalt-fil-lao-bitextmining
alt-fil-lao-bitextmining
Deduplicated copy of kornwtp/alt-fil-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-fil-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-fil-lao-bitextmining.alt-mya-lao-bitextmining
alt-mya-lao-bitextmining
Deduplicated copy of kornwtp/alt-mya-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-mya-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-mya-lao-bitextmining.alt-ind-lao-bitextmining
alt-ind-lao-bitextmining
Deduplicated copy of kornwtp/alt-ind-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-ind-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-ind-lao-bitextmining.alt-zsm-lao-bitextmining
alt-zsm-lao-bitextmining
Deduplicated copy of kornwtp/alt-zsm-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/alt-zsm-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/alt-zsm-lao-bitextmining.
