datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViMed-PET-part1
Dataset description for three years: 2017, 2018, 2019
This dataset contains data from three years (2017, 2018, 2019). Each year has several month folders, which are named as THANG {month}.
Each year folder is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract all data folders.
Folder structure after extraction
Each folder named THANG {month} of a year is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2807/ViMed-PET-part1.ViMed-PET-part2
Dataset description for year 2023
This dataset contains data from 8 months: January to September, except August, stored in the following folders respectively:
THANG 1
THANG 2
...
THANG 7
THANG 9
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/Duc2305/ViMed-PET-part2.ViMed-PET-part3
Dataset description for year 2023
This dataset contains data from three months: October, November, and December, stored in the following folders respectively:
THANG 10
THANG 11
THANG 12
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders, corresponding to 2… See the full description on the dataset page: https://huggingface.co/datasets/dacthai2k/ViMed-PET-part3.ViMU
ViMU: Benchmarking Video Metaphorical Understanding
Qi Li, Xinchao Wang*
*Corresponding author
xML Lab, National University of Singapore
Our GitHub repository contains the evaluation scripts for ViMU, a benchmark for video metaphorical understanding. The code evaluates multimodal models on four tasks:
Open-ended interpretation (OE)
Evidence grounding (EG)
Rhetoric mechanism identification (RM)
Social value signal identification (SV)
Directory Structure
Expected… See the full description on the dataset page: https://huggingface.co/datasets/LIQIIIII/ViMU.vimgolf-public-challenges-inspect-evalViMix-14M
ViMix-14M: A Curated Multi-Source Video-Text Dataset
Dataset Description
ViMix-14M is a large-scale video-text dataset containing ~14 million video-text pairs with multi-granularity captions, designed to address the data bottleneck in text-to-video generation.
Text-to-video generation has surged in interest since Sora, yet open-source models still face a data bottleneck: there is no large, high-quality, easily obtainable video–text corpus. Existing public datasets… See the full description on the dataset page: https://huggingface.co/datasets/TimingYang/ViMix-14M.vi_mednli
Dataset Summary
Vietnamese-version of MedNLI. The data has been used as a benchmark for evaluating a Vietnamese Biomedical-domain Transformer model.
Citation
Please cite this paper if you use this dataset:
@misc{vipubmed,
doi = {10.48550/ARXIV.2210.05598},
url = {https://arxiv.org/abs/2210.05598},
author = {Phan, Long and Dang, Tai and Tran, Hieu and Phan, Vy and Chau, Lam D. and Trinh, Trieu H.},
keywords = {Computation and Language (cs.CL), Artificial… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/vi_mednli.vi_math_school_full
Vietnamese School Math Dataset
vi_math_school_full is a Vietnamese-language dataset containing 8,656
school-level mathematics problems and related explanations.
Main Uses
Vietnamese math question answering
Mathematical reasoning and problem solving
Instruction tuning / supervised fine-tuning
Educational AI applications
Evaluation of Vietnamese language models on school mathematics
Language
Vietnamese
Domain
School mathematics… See the full description on the dataset page: https://huggingface.co/datasets/ngquocvinh/vi_math_school_full.stillwarm-kv-cache-artifact
A downloadable KV-cache save file — with the honest math
One llama-server slot save: the first 8,192 Llama-tokens of Frankenstein
(public domain), prefilled by Qwen2.5-7B-Instruct Q4_K_M (Apache-2.0 model —
chosen over Llama specifically for artifact licensing) and saved with a
stillwarm sidecar.
This file is USELESS unless your setup matches the sidecar exactly:
field
value
llama.cpp build
b9871 (ef2d770117db45b05aa7ecd1b0acca36370c5470) — advisory: ±5 weeks measured… See the full description on the dataset page: https://huggingface.co/datasets/vimalnakrani/stillwarm-kv-cache-artifact.vi_math_problem_crawl
Dataset Card for Vietnamese Elementary Math Knowledge and Workbook
Dataset Summary
The data includes information about elementary school math knowledge in Vietnam, as well as exercises compiled from books. This is a crawlable dataset that can be trained for text generation tasks.
Supported Tasks and Leaderboards
Languages
The majority of the data is in Vietnamese, but there is still some English from some bilingual workbooks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hllj/vi_math_problem_crawl.sycophancy-induction-resultssometimesanotion__Qwen2.5-14B-Vimarckoso-v3-model_stock-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-model_stock
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-model_stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-model_stock-details.vim-command-pocket
NickIBrody/vim-command-pocket
Vim Command Pocket Dataset is a small offline seed built from official Vim help pages.
It is intentionally lightweight and is meant to serve as the starting point for a larger Vim corpus.
Source pages
https://vimhelp.org/usr_02.txt.html
https://vimhelp.org/quickref.txt.html
Split sizes
{
"total_examples": 747,
"train": 599,
"validation": 74,
"test": 74,
"sources": [
"https://vimhelp.org/quickref.txt.html"… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/vim-command-pocket.vimmrc2.0
ViMMRC 2.0
The Vietnamese Multiple-choice reading comprehension dataset version 2 (ViMMRC 2.0)
The dataset is freely available for research purposes only. Users need to sign the data agreement before receiving the dataset.
More information, please visit the NLP@UIT research group: https://nlp.uit.edu.vn/
The original Github for the dataset (including source code): https://github.com/sonlam1102/vimmrc2
Usage
from datasets import load_dataset
train =… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vimmrc2.0.Vims_cleanvimqa-generated-answers-pass1
Vi-MQA - Pass 1 Generated Answers & Evaluation
This repo contains the Pass 1 outputs and evaluation results for the Vi-MQA Dataset from the VMLU Benchmark Suite with a total of 4,762 records.
Folder Structure
1. Model Outputs (raw_outputs/)
Contains the formatted outputs from the 3 models evaluated in Pass 1:
results_pass1_gemma.jsonl (Gemma 4 31B IT)
results_pass1_llama.jsonl (Llama 4 Scout)
results_pass1_qwen.jsonl (Qwen3 32B)
2.… See the full description on the dataset page: https://huggingface.co/datasets/nygdon/vimqa-generated-answers-pass1.VimGPT-InsertDelete-354k-PythonCodeVi_MetaMathQAVi_MathInstructsometimesanotion__Qwen2.5-14B-Vimarckoso-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v2-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v2
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v2-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-details.multi_news_vietnamese_vimssft-fiscal-fr-demo
Vimen SFT French Tax Law (Demonstration Sample)
Vimen, expert data for European AI
This is a demonstration sample of 15 prompt/response pairs. It is not a training dataset and not licensed for any use. It exists to show Vimen's production and review methodology on a regulated, non-machine-verifiable domain. Production datasets are built to order. Access to the attached file is granted manually, on request, for inspection only. See the License section.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vimen/sft-fiscal-fr-demo.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-IF-Variant-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-IF-Variant
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-IF-Variant
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-IF-Variant-details.pairrm-llama3-preference-datasetllm_judge_preferencesViMedAQAbfsi_QA_alpaca_dataset
BFSI Alpaca Dataset
Dataset Summary
This dataset contains synthetic Banking, Financial Services, and Insurance (BFSI) customer support conversations in Alpaca format.[file:40]Each record is a short, standardized query–response pair designed for training lightweight call center assistants that prioritize safety and compliance.[file:40]
Languages
English
Use Cases
Training or fine-tuning small language models for:
Loan, EMI, and disbursement queries… See the full description on the dataset page: https://huggingface.co/datasets/VimalAntony/bfsi_QA_alpaca_dataset.
