datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vehicle-mixed-traffic-detection
Visaitech Mixed-Traffic Vehicle Detection Dataset (v0.1)
Dashcam frames annotated for pedestrian / 2-wheeler / 3-wheeler / 4-wheeler
detection in South Asian mixed traffic, a class taxonomy general-purpose
COCO-trained detectors don't cover (COCO has no concept of an auto-rickshaw
or motorcycle-vs-bicycle-as-one-class "2-wheeler" grouping tuned for how
this traffic actually mixes on the road).
This is an early v0.1 release: 293 annotated frames from 6 source videos,
published… See the full description on the dataset page: https://huggingface.co/datasets/visaitech/vehicle-mixed-traffic-detection.gsm8k-thai
gsm8k-thai
This dataset is a Thai translation of the GSM8k benchmark (https://huggingface.co/datasets/openai/gsm8k), a dataset of grade school math word problems. The translation was performed using Claude 3.5 Sonnet. It is intended for evaluating the performance of language models on mathematical reasoning in the Thai language. The split of training and test data follows the original GSM8k dataset.
Annotations
source: claude-3.5-sonnet
language: en -> th… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/gsm8k-thai.nitibench
👩🏻⚖️ NitiBench: A Thai Legal Benchmark for RAG
[📄 Technical Report] | [👨💻 Github Repository]
This dataset provides the test data for evaluating LLM frameworks, such as RAG or LCLM. The benchmark consists of two datasets:
NitiBench-CCL
NitiBench-Tax
🏛️ NitiBench-CCL
Derived from the WangchanX-Legal-ThaiCCL-RAG Dataset, our version includes an additional preprocessing step in which we separate the reasoning process from the final answer. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/nitibench.VisAssistpaper-visavisa-datasetwiki-visafineweb-visaus-work-visa-salary-dataset
US Work Visa & Salary Disclosures (H-1B, PERM, LCA)
14.5M salary disclosure records from official US Department of Labor and USCIS visa filings (H-1B LCAs, PERM) — employer, job title, wage and worksite, spanning FY2008 to the latest fiscal year.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-work-visa-salary-dataset
Formats & how to load
Native… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-work-visa-salary-dataset.VISA
Visual Abstraction: A Plug-and-Play Approach for Text-Visual Retrieval
This Hugging Face model repository corresponds to the GitHub project:👉 XLearning-SCU/2025-ICML-VISA
Please visit the GitHub repository for full implementation details, code, and additional resources.
Usage
The processed directory contains intermediate files for datasets used in this project. These files are preprocessed and ready for use in experiments and evaluations.
Intermediate File… See the full description on the dataset page: https://huggingface.co/datasets/XLearning-SCU/VISA.digital-nomad-visa-data
GlobeNomad Visa Dataset
Long-stay, remote-work and nomad visa records — one per programme — each sourced to an
official government page and carrying the date it was last checked.
A country may hold several records. Thailand publishes four: the DTV, the education visa, the
retirement route and the Non-B. Group by country_slug, not by slug — slug is the record key.
See CHANGELOG.md if you are holding a file from before 2026-08-27, when this
was one row per country.
Free to use… See the full description on the dataset page: https://huggingface.co/datasets/globenomad/digital-nomad-visa-data.VisAlign
VisAlign: Dataset for Measuring the Alignment between AI and Humans in Visual Perception
This is the test set of VisAlign (NeurIPS 2023 Datasets and Benchmarks Track), a dataset for measuring the degree of alignment between AI models and humans in visual perception. It contains 900 images across 8 categories.
Ground-truth labels and per-image categories are withheld, and filenames are anonymized IDs — to evaluate your model, submit your predictions to the VisAlign Leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/VisAlign.visaadvisor-country-hub-v0-1-0
VisaAdvisor Country Hub v0.1.0
العربية
This repository is a discovery mirror of the VisaAdvisor Country Hub v0.1.0 public foundation pre-release. Read the bilingual Country Hub landing page. The canonical, citable release is archived on Zenodo under DOI 10.5281/zenodo.21858354. The full source, schema, governance documents and reproducible build scripts are available in the public GitHub repository.
The release preserves the distinction between official claims, editorial… See the full description on the dataset page: https://huggingface.co/datasets/visaadvisor1/visaadvisor-country-hub-v0-1-0.visayan-ethnohistory
Visayan Ethnohistory and Matutino Lineage Studies
Author: Karl Romeo Soland y LacsonORCID iD: 0009-0008-0902-4945Wikidata Item: Q140312355Deutsche Nationalbibliothek (DNB) GND/PND Records:
Record ID 1 (PIZ): 1401722016
Record ID 2 (PIZ): 1401575749
Date: June 2026
Dataset: Matutino/visayan-ethnohistory
Dieses Dataset enthält grundlegende Forschungsarbeiten von Karl Romeo Soland y Lacson zur Matutino-Linie aus Anilao (heute Barrio Cabutungan, Sara, Iloilo, Panay… See the full description on the dataset page: https://huggingface.co/datasets/Matutino/visayan-ethnohistory.unilink-visa-handbook 1|# UNILINK Visa Handbook Dataset
2|
3|> A neutral, citable corpus of visa & immigration facts across 8 jurisdictions (AU/UK/US/CA/NZ/JP/HK/MY), compiled and structured by **UNILINK Education** (licensed education & migration agent, MARN 1687552 / QEAC G167) from official government sources.
4|
5|[](https://creativecommons.org/licenses/by/4.0/)
6|[. It comprises 5,127 legal sections extracted from 35 Thai legislations (primarily focusing on Corporate and Commercial Law).
It is designed to be used as a Context Pool (Knowledge Base) for Retrieval-Augmented Generation (RAG) pipelines. Researchers and developers can load this dataset to… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/nitibench-statute.visa-approval-refusal-rates
Visa approval and refusal rates: Schengen consulates and US nationalities
Three government datasets, normalised across years and made usable. The numbers
are not mine — they are the European Commission's and the US State Department's.
What is mine is the reconciliation: the EU publishes one spreadsheet per year with
country labels that drift between them, and the US publishes PDFs.
Maintained at visachances.com, which is built from
these files.
What's here… See the full description on the dataset page: https://huggingface.co/datasets/sagegar/visa-approval-refusal-rates.document-photo-requirements
Verified Document Photo Requirements Dataset
A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact.
Dataset summary
Version: 1.0.0
Release date: 2026-08-14
Latest source review represented: 2026-08-11
Records: 18 (12 passport, 5 visa, 1 national ID)
Coverage: 14 countries or regions
Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.ViSA_LlavaOV_700KThis dataset was presented in the paper Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents.
Code: https://github.com/HITsz-TMG/ViSA
ViSA_LlavaOV_80KThis dataset was presented in the paper Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents.
Code: https://github.com/HITsz-TMG/ViSA
VIS-APP-Bench
anchor_tasks_web — Dataset
This dataset is part of the benchmark presented in the paper VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents.
Project Page | GitHub Repository | Paper
A web-app generation benchmark. Each task is a multi-page UI taken from a public Figma community file. For every task we ship the textual page descriptions, the rendered mockup PNGs, the Figma node structure, the per-page click-annotations, and the distilled data-testid "anchors"… See the full description on the dataset page: https://huggingface.co/datasets/JunJiaGuo/VIS-APP-Bench.passport-visa-photo-specs
Passport & Visa Photo Specifications (100 Countries, 248 Document Formats)
Machine-readable photo requirements for passports, visas, ID cards, residence permits, and driving licences across 100 countries: physical dimensions (mm), pixel dimensions, DPI, background colour, government portal file-size caps, and a citation to the issuing-authority source for every spec.
Files
passport_photo_specs.csv - flat table, one row per country+document (248 rows)… See the full description on the dataset page: https://huggingface.co/datasets/Whiterocket/passport-visa-photo-specs.visargs
Dataset Card for VisArgs Benchmark
Dataset Summary
Data from: Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
@article{chung2024selective,
title={Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding},
author={Chung, Jiwan and Lee, Sungjae and Kim, Minseo and Han, Seungju and Yousefpour, Ashkan and Hessel, Jack and Yu, Youngjae},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/jiwan-chung/visargs.VisATB[WWW 2026] (VisATB) Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty
Paper Link: https://arxiv.org/abs/2403.04343
Project Link: https://github.com/YanqiDai/VisATB
JUSTNLP2025-L-Summ-formatted
JUSTNLP20205-L-SUMM Formatted Data
This repository provides a filtered and formatted dataset used to train and validate the model prior to submitting.
Data Filtering
We explore the relationship between the length of a judgment and the length of its summarization, measured in characters and words. When plotted on a log scale, summarization length shows a strong correlation with judgment length.
To reduce noise that could affect model performance, we remove samples where… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/JUSTNLP2025-L-Summ-formatted.south-america-travel-planning-index-2026
South America Travel Planning Index 2026
Version 1.1 is a source-linked travel-planning dataset covering all 12 sovereign South American countries. It combines editorial trip-length ranges, buffer-day guidance, gateways, route intensity, seasonality, signature experiences, official tourism links, and entry-check links with a separate registry of 44 official-source records.
Canonical record
Version DOI: https://doi.org/10.5281/zenodo.21765168
Concept DOI for all… See the full description on the dataset page: https://huggingface.co/datasets/visaadvisor1/south-america-travel-planning-index-2026.thai-gazette-evidence-retrieval
Thai Royal Gazette Evidence Retrieval
A retrieval dataset over notices from the Thai Royal Gazette
(ราชกิจจานุเบกษา ratchakitcha 'Royal Gazette'). Each question links to the
documents that answer it. Each link gives the exact character span of the
evidence in the document. The dataset contains questions, documents, and
relevance judgments. It contains no generated answers.
Layout
The dataset uses the MTEB retrieval layout of one repository and three configs.… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/thai-gazette-evidence-retrieval.
