CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01visaitech /vehicle-mixed-traffic-detection Visaitech Mixed-Traffic Vehicle Detection Dataset (v0.1) Dashcam frames annotated for pedestrian / 2-wheeler / 3-wheeler / 4-wheeler detection in South Asian mixed traffic, a class taxonomy general-purpose COCO-trained detectors don't cover (COCO has no concept of an auto-rickshaw or motorcycle-vs-bicycle-as-one-class "2-wheeler" grouping tuned for how this traffic actually mixes on the road). This is an early v0.1 release: 293 annotated frames from 6 source videos, published… See the full description on the dataset page: https://huggingface.co/datasets/visaitech/vehicle-mixed-traffic-detection.imageobject-detectionn<1K1 likes1.1k downloads2mo agoHugging Face02VISAI-AI /gsm8k-thai gsm8k-thai This dataset is a Thai translation of the GSM8k benchmark (https://huggingface.co/datasets/openai/gsm8k), a dataset of grade school math word problems. The translation was performed using Claude 3.5 Sonnet. It is intended for evaluating the performance of language models on mathematical reasoning in the Thai language. The split of training and test data follows the original GSM8k dataset. Annotations source: claude-3.5-sonnet language: en -> th… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/gsm8k-thai.text1K<n<10K0 likes538 downloads2y agoHugging Face03VISAI-AI /nitibench 👩🏻‍⚖️ NitiBench: A Thai Legal Benchmark for RAG [📄 Technical Report] | [👨‍💻 Github Repository] This dataset provides the test data for evaluating LLM frameworks, such as RAG or LCLM. The benchmark consists of two datasets: NitiBench-CCL NitiBench-Tax 🏛️ NitiBench-CCL Derived from the WangchanX-Legal-ThaiCCL-RAG Dataset, our version includes an additional preprocessing step in which we separate the reasoning process from the final answer. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/nitibench.textsentence-similarity1K<n<10K8 likes409 downloads10mo agoHugging Face04gaoCleo /VisAssisttabular10K<n<100K0 likes228 downloads11mo agoHugging Face05MrLight /paper-visaimage100K<n<1M1 likes227 downloads2y agoHugging Face06dddraxxx /visa-datasetimage1K<n<10K0 likes216 downloads2y agoHugging Face07MrLight /wiki-visaimage10K<n<100K0 likes200 downloads2y agoHugging Face08MrLight /fineweb-visaimage10K<n<100K3 likes172 downloads2y agoHugging Face09zalizedata /us-work-visa-salary-dataset US Work Visa & Salary Disclosures (H-1B, PERM, LCA) 14.5M salary disclosure records from official US Department of Labor and USCIS visa filings (H-1B LCAs, PERM) — employer, job title, wage and worksite, spanning FY2008 to the latest fiscal year. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/us-work-visa-salary-dataset Formats & how to load Native… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/us-work-visa-salary-dataset.tabulartabular-regression10M<n<100M0 likes134 downloads2mo agoHugging Face10XLearning-SCU /VISA Visual Abstraction: A Plug-and-Play Approach for Text-Visual Retrieval This Hugging Face model repository corresponds to the GitHub project:👉 XLearning-SCU/2025-ICML-VISA Please visit the GitHub repository for full implementation details, code, and additional resources. Usage The processed directory contains intermediate files for datasets used in this project. These files are preprocessed and ready for use in experiments and evaluations. Intermediate File… See the full description on the dataset page: https://huggingface.co/datasets/XLearning-SCU/VISA.text4 likes124 downloads10mo agoHugging Face11globenomad /digital-nomad-visa-data GlobeNomad Visa Dataset Long-stay, remote-work and nomad visa records — one per programme — each sourced to an official government page and carrying the date it was last checked. A country may hold several records. Thailand publishes four: the DTV, the education visa, the retirement route and the Non-B. Group by country_slug, not by slug — slug is the record key. See CHANGELOG.md if you are holding a file from before 2026-08-27, when this was one row per country. Free to use… See the full description on the dataset page: https://huggingface.co/datasets/globenomad/digital-nomad-visa-data.textn<1K0 likes115 downloads15d agoHugging Face12jiyounglee0523 /VisAlign VisAlign: Dataset for Measuring the Alignment between AI and Humans in Visual Perception This is the test set of VisAlign (NeurIPS 2023 Datasets and Benchmarks Track), a dataset for measuring the degree of alignment between AI models and humans in visual perception. It contains 900 images across 8 categories. Ground-truth labels and per-image categories are withheld, and filenames are anonymized IDs — to evaluate your model, submit your predictions to the VisAlign Leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/VisAlign.imageimage-classificationn<1K1 likes98 downloads1mo agoHugging Face13visaadvisor1 /visaadvisor-country-hub-v0-1-0 VisaAdvisor Country Hub v0.1.0 العربية This repository is a discovery mirror of the VisaAdvisor Country Hub v0.1.0 public foundation pre-release. Read the bilingual Country Hub landing page. The canonical, citable release is archived on Zenodo under DOI 10.5281/zenodo.21858354. The full source, schema, governance documents and reproducible build scripts are available in the public GitHub repository. The release preserves the distinction between official claims, editorial… See the full description on the dataset page: https://huggingface.co/datasets/visaadvisor1/visaadvisor-country-hub-v0-1-0.tabularn<1K0 likes80 downloads1mo agoHugging Face14Matutino /visayan-ethnohistory Visayan Ethnohistory and Matutino Lineage Studies Author: Karl Romeo Soland y LacsonORCID iD: 0009-0008-0902-4945Wikidata Item: Q140312355Deutsche Nationalbibliothek (DNB) GND/PND Records: Record ID 1 (PIZ): 1401722016 Record ID 2 (PIZ): 1401575749 Date: June 2026 Dataset: Matutino/visayan-ethnohistory Dieses Dataset enthält grundlegende Forschungsarbeiten von Karl Romeo Soland y Lacson zur Matutino-Linie aus Anilao (heute Barrio Cabutungan, Sara, Iloilo, Panay… See the full description on the dataset page: https://huggingface.co/datasets/Matutino/visayan-ethnohistory.documentn<1K0 likes70 downloads3mo agoHugging Face15benemi /unilink-visa-handbook 1|# UNILINK Visa Handbook Dataset 2| 3|> A neutral, citable corpus of visa & immigration facts across 8 jurisdictions (AU/UK/US/CA/NZ/JP/HK/MY), compiled and structured by **UNILINK Education** (licensed education & migration agent, MARN 1687552 / QEAC G167) from official government sources. 4| 5|[![License: CC BY 4.0](https://img.shields.io/badge/License-CC_BY_4.0-lightgrey.svg)](https://creativecommons.org/licenses/by/4.0/) 6|[![Entries](https://img.shields.io/badge/entries-2… See the full description on the dataset page: https://huggingface.co/datasets/benemi/unilink-visa-handbook.textn<1K0 likes67 downloads4mo agoHugging Face16nvidia /Nemotron-Content-VISafe-v1 Dataset Description: VISafe is a Vietnamese-language AI safety evaluation probe dataset for testing model and guardrail behavior on safety-critical prompts. The current validated build contains 3,212 text probes in Vietnamese across jailbreak, toxicity, misinformation, prompt injection, Vietnam-specific political sensitivity, cybercrime, over-refusal, and privacy categories. The dataset combines translated probes from established English safety benchmarks with Vietnamese-native… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-VISafe-v1.text1K<n<10K2 likes55 downloads3mo agoHugging Face17steven0226 /defectforge-visa-synthetic DefectForge VisA Synthetic Defects Synthetic defect images and generation-time masks for the VisA pcb1 and capsules objects. Each object is generated from only 10 real anomalous training images, while the frozen high-shot test partition is never visible to generation, filtering, or quality reference sets. 繁中摘要:這是 VisA pcb1/capsules 的少樣本工業瑕疵合成資料。每個物件只用 10 張真實瑕疵訓練圖;mask 是生成時使用的標註,不是模型事後預測。資料同時提供 filtered 與 unfiltered 版本,並公開 provenance 與 test SHA-256 blocklist。 What is… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/defectforge-visa-synthetic.imagen<1K0 likes54 downloads2mo agoHugging Face18zli99 /VisAnalog VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images VisAnalog is a diagnostic benchmark for visual concept transfer on natural images. Each example follows an analogy pattern: infer the transformation from pair1_source to pair1_target, transfer that concept to pair2_source, and answer a multiple-choice question about the expected pair2_target. Dataset Structure The uploaded split is test with 617 examples. Main columns: pair1_source: first source… See the full description on the dataset page: https://huggingface.co/datasets/zli99/VisAnalog.imagevisual-question-answeringn<1K0 likes50 downloads4mo agoHugging Face19VISAI-AI /nitibench-statute 📜 NitiBench-Statute: Thai Legal Corpus for RAG Part of the NitiBench Project This dataset contains the complete corpus of legal sections used in the NitiBench benchmark (CCL and Tax subset). It comprises 5,127 legal sections extracted from 35 Thai legislations (primarily focusing on Corporate and Commercial Law). It is designed to be used as a Context Pool (Knowledge Base) for Retrieval-Augmented Generation (RAG) pipelines. Researchers and developers can load this dataset to… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/nitibench-statute.text1K<n<10K0 likes49 downloads10mo agoHugging Face20sagegar /visa-approval-refusal-rates Visa approval and refusal rates: Schengen consulates and US nationalities Three government datasets, normalised across years and made usable. The numbers are not mine — they are the European Commission's and the US State Department's. What is mine is the reconciliation: the EU publishes one spreadsheet per year with country labels that drift between them, and the US publishes PDFs. Maintained at visachances.com, which is built from these files. What's here… See the full description on the dataset page: https://huggingface.co/datasets/sagegar/visa-approval-refusal-rates.tabular10K<n<100K0 likes39 downloads1mo agoHugging Face21passport-visa-photo-studio /document-photo-requirements Verified Document Photo Requirements Dataset A structured reference collection of official-source passport, visa, and national ID photo requirements maintained by Passport Visa Photo Studio. It is not a photo corpus, training dataset, or model artifact. Dataset summary Version: 1.0.0 Release date: 2026-08-14 Latest source review represented: 2026-08-11 Records: 18 (12 passport, 5 visa, 1 national ID) Coverage: 14 countries or regions Formats: CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/passport-visa-photo-studio/document-photo-requirements.tabularn<1K0 likes32 downloads1mo agoHugging Face22foggyforest /ViSA_LlavaOV_700KThis dataset was presented in the paper Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents. Code: https://github.com/HITsz-TMG/ViSA textimage-text-to-text100K<n<1M0 likes29 downloads1y agoHugging Face23foggyforest /ViSA_LlavaOV_80KThis dataset was presented in the paper Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents. Code: https://github.com/HITsz-TMG/ViSA textimage-text-to-text10K<n<100K0 likes27 downloads1y agoHugging Face24JunJiaGuo /VIS-APP-Benchgated anchor_tasks_web — Dataset This dataset is part of the benchmark presented in the paper VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents. Project Page | GitHub Repository | Paper A web-app generation benchmark. Each task is a multi-page UI taken from a public Figma community file. For every task we ship the textual page descriptions, the rendered mockup PNGs, the Figma node structure, the per-page click-annotations, and the distilled data-testid "anchors"… See the full description on the dataset page: https://huggingface.co/datasets/JunJiaGuo/VIS-APP-Bench.imageimage-to-textn<1K1 likes27 downloads3mo agoHugging Face25Whiterocket /passport-visa-photo-specs Passport & Visa Photo Specifications (100 Countries, 248 Document Formats) Machine-readable photo requirements for passports, visas, ID cards, residence permits, and driving licences across 100 countries: physical dimensions (mm), pixel dimensions, DPI, background colour, government portal file-size caps, and a citation to the issuing-authority source for every spec. Files passport_photo_specs.csv - flat table, one row per country+document (248 rows)… See the full description on the dataset page: https://huggingface.co/datasets/Whiterocket/passport-visa-photo-specs.textn<1K1 likes26 downloads3mo agoHugging Face26jiwan-chung /visargs Dataset Card for VisArgs Benchmark Dataset Summary Data from: Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding @article{chung2024selective, title={Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding}, author={Chung, Jiwan and Lee, Sungjae and Kim, Minseo and Han, Seungju and Yousefpour, Ashkan and Hessel, Jack and Yu, Youngjae}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/jiwan-chung/visargs.image1K<n<10K1 likes25 downloads2y agoHugging Face27YanqiDai /VisATB[WWW 2026] (VisATB) Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty Paper Link: https://arxiv.org/abs/2403.04343 Project Link: https://github.com/YanqiDai/VisATB text1M<n<10M1 likes25 downloads8mo agoHugging Face28VISAI-AI /JUSTNLP2025-L-Summ-formatted JUSTNLP20205-L-SUMM Formatted Data This repository provides a filtered and formatted dataset used to train and validate the model prior to submitting. Data Filtering We explore the relationship between the length of a judgment and the length of its summarization, measured in characters and words. When plotted on a log scale, summarization length shows a strong correlation with judgment length. To reduce noise that could affect model performance, we remove samples where… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/JUSTNLP2025-L-Summ-formatted.textsummarization1K<n<10K0 likes22 downloads11mo agoHugging Face29visaadvisor1 /south-america-travel-planning-index-2026 South America Travel Planning Index 2026 Version 1.1 is a source-linked travel-planning dataset covering all 12 sovereign South American countries. It combines editorial trip-length ranges, buffer-day guidance, gateways, route intensity, seasonality, signature experiences, official tourism links, and entry-check links with a separate registry of 44 official-source records. Canonical record Version DOI: https://doi.org/10.5281/zenodo.21765168 Concept DOI for all… See the full description on the dataset page: https://huggingface.co/datasets/visaadvisor1/south-america-travel-planning-index-2026.tabularn<1K0 likes21 downloads1mo agoHugging Face30VISAI-AI /thai-gazette-evidence-retrieval Thai Royal Gazette Evidence Retrieval A retrieval dataset over notices from the Thai Royal Gazette (ราชกิจจานุเบกษา ratchakitcha 'Royal Gazette'). Each question links to the documents that answer it. Each link gives the exact character span of the evidence in the document. The dataset contains questions, documents, and relevance judgments. It contains no generated answers. Layout The dataset uses the MTEB retrieval layout of one repository and three configs.… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/thai-gazette-evidence-retrieval.tabulartext-retrieval10K<n<100K0 likes21 downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.