datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ofac-sdn-vessel-ownership-resolutionCanonical page: https://arcnautical.com/data/ofac-sdn-vessel-ownership-resolution/Permanent record (DOI): https://doi.org/10.5281/zenodo.22808695Publisher: ArcNautical. Public-source compilation (OFAC SDN + GLEIF), CC BY 4.0. This Hugging Face copy mirrors the canonical page; the Zenodo DOI is the citable identifier.
OFAC SDN vessel records (screened subset), snapshot 2026-09-16: flag and public GLEIF ownership-resolution stage
Public-record dataset of 1,527 screened… See the full description on the dataset page: https://huggingface.co/datasets/saltytar01/ofac-sdn-vessel-ownership-resolution.florence2-ofa-captions-500
OFA Florence-2 Dataset (500 Samples)
This dataset was generated using microsoft/Florence-2-large on a subset of COCO 2017 Validation images.
It is pre-formatted for OFA Stage-1 fine-tuning (Headerless TSV, URL-safe base64, max 512x512 resolution).
RagabilityCorpusCurrent version: Dataset_v0.4.tsv (converted to ragability format: v0d4.hjson)
Ragability Corpus
In the following, we introduce WikiContradict (the empirical basis for the Ragability Corpus), describe the Ragability Corpus, and finally explain how the dataset can be extended and how a new one can be created.
Empirical basis
WikiContradict is a benchmark for evaluating LLMs on real-world knowledge conflicts from Wikipedia (see the Hou et. al. 2025 and the dataset for more… See the full description on the dataset page: https://huggingface.co/datasets/ofai/RagabilityCorpus.
