datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMDU
📢 News
[06/13/2024] 🚀 We release our MMDU benchmark and MMDU-45k instruct tunning data to huggingface.
💎 MMDU Benchmark
To evaluate the multi-image multi-turn dialogue capabilities of existing models, we have developed the MMDU Benchmark. Our benchmark comprises 110 high-quality multi-image multi-turn dialogues with more than 1600 questions, each accompanied by detailed long-form answers. Previous benchmarks typically involved only single images or a small number of… See the full description on the dataset page: https://huggingface.co/datasets/laolao77/MMDU.MIA-DPODataset for paper 'MIA-DPO'
We release dpo data of LLaVa-v1.5-7B. You can construct more data using our code in github.
lao_pairs_final
🇱🇦 Lao SFT Pairs Final
A cleaned and merged Lao-language instruction-tuning dataset for supervised fine-tuning (SFT) of large language models — specifically built to improve Lao language capability in models like Gemma 4.
Dataset Summary
Split
File
Examples
Train
lao_train_final.jsonl
57,088
Validation
lao_val_final.jsonl
2,978
Total
60,066
Data Sources
This dataset merges two sources:
1. Lao continuation corpus (32.5%)
Real… See the full description on the dataset page: https://huggingface.co/datasets/AOYPSK/lao_pairs_final.
