datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visa-anomaly-detection
VisA — Visual Anomaly Dataset
Mirror of the VisA (Visual Anomaly) dataset for research use. Staged as a proxy/pretraining
dataset for CoRe's Situational Control paint-inspection work (core-lab/situational-control).
Source
Original repo: https://github.com/amazon-science/spot-diff
Original download: https://amazon-visual-anomaly.s3.us-west-2.amazonaws.com/VisA_20220922.tar
License: CC BY 4.0 (confirmed in the source repo README)
Contents
10,821… See the full description on the dataset page: https://huggingface.co/datasets/imaadd05/visa-anomaly-detection.vehicle-mixed-traffic-detection
Visaitech Mixed-Traffic Vehicle Detection Dataset (v0.1)
Dashcam frames annotated for pedestrian / 2-wheeler / 3-wheeler / 4-wheeler
detection in South Asian mixed traffic, a class taxonomy general-purpose
COCO-trained detectors don't cover (COCO has no concept of an auto-rickshaw
or motorcycle-vs-bicycle-as-one-class "2-wheeler" grouping tuned for how
this traffic actually mixes on the road).
This is an early v0.1 release: 293 annotated frames from 6 source videos,
published… See the full description on the dataset page: https://huggingface.co/datasets/visaitech/vehicle-mixed-traffic-detection.mirage_mvtec_visadol-visas-database
DOL Visas Database (H-1B LCA + PERM)
Every H-1B/H-1B1/E-3 Labor Condition Application and PERM permanent labor
certification application disclosed by the DOL Office of Foreign Labor
Certification, FY2015 to present, as a single queryable DuckDB database.
8,812,639 rows across 2 tables.
Table
Description
Row Count
Column Count
Date Range
lca
H-1B/H-1B1/E-3 Labor Condition Applications, one row per application per disclosure file, FY2015-present
7,479,697
110
FY2015 to… See the full description on the dataset page: https://huggingface.co/datasets/Nason/dol-visas-database.visaOriginal dataset:
@inproceedings{zou2022spot,
title={Spot-the-difference self-supervised pre-training for anomaly detection and segmentation},
author={Zou, Yang and Jeong, Jongheon and Pemula, Latha and Zhang, Dongqing and Dabeer, Onkar},
booktitle={European Conference on Computer Vision},
pages={392--408},
year={2022},
organization={Springer}
}
gsm8k-thai
gsm8k-thai
This dataset is a Thai translation of the GSM8k benchmark (https://huggingface.co/datasets/openai/gsm8k), a dataset of grade school math word problems. The translation was performed using Claude 3.5 Sonnet. It is intended for evaluating the performance of language models on mathematical reasoning in the Thai language. The split of training and test data follows the original GSM8k dataset.
Annotations
source: claude-3.5-sonnet
language: en -> th… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/gsm8k-thai.nitibench
👩🏻⚖️ NitiBench: A Thai Legal Benchmark for RAG
[📄 Technical Report] | [👨💻 Github Repository]
This dataset provides the test data for evaluating LLM frameworks, such as RAG or LCLM. The benchmark consists of two datasets:
NitiBench-CCL
NitiBench-Tax
🏛️ NitiBench-CCL
Derived from the WangchanX-Legal-ThaiCCL-RAG Dataset, our version includes an additional preprocessing step in which we separate the reasoning process from the final answer. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/nitibench.VisA-2KHigh-resolution industrial image anomaly detection dataset VisA-2K.For more information, see HiAD.
Download
huggingface-cli download --repo-type dataset XimiaoZhang/VisA-2K --local-dir VisA-2K --resume-download
VisAssistCapacitacao_Visao_Computacional
Visão Geral
Este repositório contém as atividades práticas e teóricas do curso de Capacitação em Visão Computacional. O curso aborda fundamentos de processamento digital de imagens, técnicas de filtragem, segmentação, extração de características e aplicações em aprendizado de máquina.
Estrutura do Repositório
O repositório está organizado em pastas por atividade, cada uma contendo:
Enunciado da atividade em PDF
Notebook Jupyter (quando aplicável)
README com… See the full description on the dataset page: https://huggingface.co/datasets/arvoredossaberes/Capacitacao_Visao_Computacional.visa-datasetpaper-visawiki-visavisafineweb-visaus-work-visa-salary-dataset
US Work Visa & Salary Disclosures (H-1B, PERM, LCA)
14.5M salary disclosure records from official US Department of Labor and USCIS visa filings (H-1B LCAs, PERM) — employer, job title, wage and worksite, spanning FY2008 to the latest fiscal year.
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/us-work-visa-salary-dataset
Formats & how to load
Native… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-work-visa-salary-dataset.VISA
Visual Abstraction: A Plug-and-Play Approach for Text-Visual Retrieval
This Hugging Face model repository corresponds to the GitHub project:👉 XLearning-SCU/2025-ICML-VISA
Please visit the GitHub repository for full implementation details, code, and additional resources.
Usage
The processed directory contains intermediate files for datasets used in this project. These files are preprocessed and ready for use in experiments and evaluations.
Intermediate File… See the full description on the dataset page: https://huggingface.co/datasets/XLearning-SCU/VISA.digital-nomad-visa-data
GlobeNomad Visa Dataset
Long-stay, remote-work and nomad visa records — one per programme — each sourced to an
official government page and carrying the date it was last checked.
A country may hold several records. Thailand publishes four: the DTV, the education visa, the
retirement route and the Non-B. Group by country_slug, not by slug — slug is the record key.
See CHANGELOG.md if you are holding a file from before 2026-08-27, when this
was one row per country.
Free to use… See the full description on the dataset page: https://huggingface.co/datasets/globenomad/digital-nomad-visa-data.visaadvisor-country-hub-v0-1-0
VisaAdvisor Country Hub v0.1.0
العربية
This repository is a discovery mirror of the VisaAdvisor Country Hub v0.1.0 public foundation pre-release. Read the bilingual Country Hub landing page. The canonical, citable release is archived on Zenodo under DOI 10.5281/zenodo.21858354. The full source, schema, governance documents and reproducible build scripts are available in the public GitHub repository.
The release preserves the distinction between official claims, editorial… See the full description on the dataset page: https://huggingface.co/datasets/visaadvisor1/visaadvisor-country-hub-v0-1-0.VisAlign
VisAlign: Dataset for Measuring the Alignment between AI and Humans in Visual Perception
This is the test set of VisAlign (NeurIPS 2023 Datasets and Benchmarks Track), a dataset for measuring the degree of alignment between AI models and humans in visual perception. It contains 900 images across 8 categories.
Ground-truth labels and per-image categories are withheld, and filenames are anonymized IDs — to evaluate your model, submit your predictions to the VisAlign Leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/VisAlign.unilink-visa-handbook 1|# UNILINK Visa Handbook Dataset
2|
3|> A neutral, citable corpus of visa & immigration facts across 8 jurisdictions (AU/UK/US/CA/NZ/JP/HK/MY), compiled and structured by **UNILINK Education** (licensed education & migration agent, MARN 1687552 / QEAC G167) from official government sources.
4|
5|[](https://creativecommons.org/licenses/by/4.0/)
6|[ GND/PND Records:
Record ID 1 (PIZ): 1401722016
Record ID 2 (PIZ): 1401575749
Date: June 2026
Dataset: Matutino/visayan-ethnohistory
Dieses Dataset enthält grundlegende Forschungsarbeiten von Karl Romeo Soland y Lacson zur Matutino-Linie aus Anilao (heute Barrio Cabutungan, Sara, Iloilo, Panay… See the full description on the dataset page: https://huggingface.co/datasets/Matutino/visayan-ethnohistory.nitibench-statute
📜 NitiBench-Statute: Thai Legal Corpus for RAG
Part of the NitiBench Project
This dataset contains the complete corpus of legal sections used in the NitiBench benchmark (CCL and Tax subset). It comprises 5,127 legal sections extracted from 35 Thai legislations (primarily focusing on Corporate and Commercial Law).
It is designed to be used as a Context Pool (Knowledge Base) for Retrieval-Augmented Generation (RAG) pipelines. Researchers and developers can load this dataset to… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/nitibench-statute.defectforge-visa-synthetic
DefectForge VisA Synthetic Defects
Synthetic defect images and generation-time masks for the VisA pcb1 and capsules
objects. Each object is generated from only 10 real anomalous training images, while the
frozen high-shot test partition is never visible to generation, filtering, or quality
reference sets.
繁中摘要:這是 VisA pcb1/capsules 的少樣本工業瑕疵合成資料。每個物件只用
10 張真實瑕疵訓練圖;mask 是生成時使用的標註,不是模型事後預測。資料同時提供
filtered 與 unfiltered 版本,並公開 provenance 與 test SHA-256 blocklist。
What is… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/defectforge-visa-synthetic.Nemotron-Content-VISafe-v1
Dataset Description:
VISafe is a Vietnamese-language AI safety evaluation probe dataset for testing model and guardrail behavior on safety-critical prompts. The current validated build contains 3,212 text probes in Vietnamese across jailbreak, toxicity, misinformation, prompt injection, Vietnam-specific political sensitivity, cybercrime, over-refusal, and privacy categories.
The dataset combines translated probes from established English safety benchmarks with Vietnamese-native… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-VISafe-v1.VisAnalog
VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
VisAnalog is a diagnostic benchmark for visual concept transfer on natural images.
Each example follows an analogy pattern: infer the transformation from pair1_source
to pair1_target, transfer that concept to pair2_source, and answer a multiple-choice
question about the expected pair2_target.
Dataset Structure
The uploaded split is test with 617 examples.
Main columns:
pair1_source: first source… See the full description on the dataset page: https://huggingface.co/datasets/zli99/VisAnalog.visa-approval-refusal-rates
Visa approval and refusal rates: Schengen consulates and US nationalities
Three government datasets, normalised across years and made usable. The numbers
are not mine — they are the European Commission's and the US State Department's.
What is mine is the reconciliation: the EU publishes one spreadsheet per year with
country labels that drift between them, and the US publishes PDFs.
Maintained at visachances.com, which is built from
these files.
What's here… See the full description on the dataset page: https://huggingface.co/datasets/sagegar/visa-approval-refusal-rates.document-photo-requirements
Verified Document Photo Requirements Dataset
A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact.
Dataset summary
Version: 1.0.0
Release date: 2026-08-14
Latest source review represented: 2026-08-11
Records: 18 (12 passport, 5 visa, 1 national ID)
Coverage: 14 countries or regions
Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.ViSA_LlavaOV_700KThis dataset was presented in the paper Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents.
Code: https://github.com/HITsz-TMG/ViSA
VISAEBench
