datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
triples-objekts
tripleS Objekts Archive (HD Images, Motion MP4s & Full Metadata)
Complete digital collectible photocards (Objekts) dataset for the K-Pop girl group tripleS (Modhaus / COSMO app).
Dataset Summary
Total Objekts: 11,025 unique digital collectibles
Total Images: 21,340 high-resolution card scans (Front & Back)
Total Motion & Voice Videos: 125 animated card loops & voice message MP4 clips
Members: All 24 members (S1–S24) + Sub-units (AAA, KRE, Assemble24, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/fadhilafif98/triples-objekts.tripmatch-ai-dataset
TripMatch AI Dataset
A reproducible multimodal dataset for the TripMatch AI Final Project. It contains
10,000 synthetic text trip plans with a raw idea generated for every row by the
pretrained Hugging Face model google/flan-t5-small, plus 5,000 real street-view images
retained as extra multimodal work. The two configurations are separate so Dataset
Viewer can load each schema correctly.
Dataset statistics
Configuration
Rows
Main fields
Intended task… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-dataset.trisearch-dataset-64k-v0.0.1
TriSearch-v1 (0.0.1)
Initial public data release for the TriSearch multimodal training stack
(v0.0.1). Curated image–text corpus for joint embedding
spaces (contrastive training + query-style text). Schema is intended to stay
stable; later versions may add fields or tighten quality filters.
Summary
Version
0.0.1 (format v1)
Examples
65,536 image–text records
Train / test
61,440 / 4,096 (test = 1/16 of data)
Image size
1024×1024 RGB JPEG… See the full description on the dataset page: https://huggingface.co/datasets/NuclearManD/trisearch-dataset-64k-v0.0.1.hw1-paper-triage-multimodal
HW1 Multimodal Paper Triage
Purpose
This dataset supports a classroom exercise in assembling and augmenting multimodal data for a personalized research-paper triage system.
Composition and splits
The dataset begins with 100 original multimodal samples. The original samples were split before augmentation using a fixed random seed and stratification by the binary target.
train: 10,070 samples consisting of 70 training originals and 10,000 augmented… See the full description on the dataset page: https://huggingface.co/datasets/ishaanamahajan/hw1-paper-triage-multimodal.KTO_trial
KTO Training Dataset
Processed from kricko/cleaned_auditor using the Auditor model.
Description
Each example contains the original image alongside adversarial heatmaps, feathered masks,
and masked images with detected unsafe regions blacked out.
Features
Column
Type
Description
image
Image
Original input image
prompt
string
Text prompt associated with the image
id
string
Unique identifier
disturbing
int8
Disturbing content score
hate
int8… See the full description on the dataset page: https://huggingface.co/datasets/ShreyashDhoot/KTO_trial.triangle-nc-geotagged
Triangle NC Geotagged Street Imagery
A large, public street-level image corpus for fine-grained visual geolocation research in North Carolina's Triangle region. It was assembled for training and evaluating the Cardinal geolocation model.
Dataset summary
The release contains 1,777,526 unique geotagged images: 1,537,750 train and 239,776 validation images. Sources: KartaView (1,169,521), Mapillary (608,002), Panoramax (3). Images cover ordinary road scenes and were… See the full description on the dataset page: https://huggingface.co/datasets/theminji/triangle-nc-geotagged.I-SPY_2_Breast_Dynamic_Contrast_Enhanced_MRI_Trial_T0_and_T3_DCE-MRI_Datasetkenya-bee-health-qa-image-triples
Kenya Bee Health Training Data
This folder is a starter database for BeeCare Anywhere / Gemma Apiary. It is intentionally small, transparent, and license-aware: use it to prove the Q/A/image-triple pipeline, then expand it with Kenyan field data before trusting model behavior in production.
Important Model Note
google/gemma-2b is a text-to-text, decoder-only model. It cannot directly read pictures. Use these image triples with a vision-capable model path, for… See the full description on the dataset page: https://huggingface.co/datasets/yahelr1/kenya-bee-health-qa-image-triples.VRSBench-satellite-triage-labels
VRSBench Satellite Triage Labels
Triage priority labels for 20,264 satellite image captions from VRSBench. Used to fine-tune LFM2.5-VL-450M for on-board satellite image triage.
Labels were generated via knowledge distillation from a larger LLM, assigning priority, reasoning, and categories to each VRSBench caption.
Files
captions.jsonl — 20,264 captions extracted from VRSBench tasks
labels.jsonl — Triage labels: priority, reasoning, categories
Schema… See the full description on the dataset page: https://huggingface.co/datasets/marcelo-earth/VRSBench-satellite-triage-labels.autotrain-data-trial
AutoTrain Dataset for project: trial
Dataset Description
This dataset has been automatically processed by AutoTrain for project trial.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<32x36 RGBA PIL image>",
"target": 0
},
{
"image": "<32x36 RGBA PIL image>",
"target": 2
}]
Dataset Fields
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Fahad-7864/autotrain-data-trial.cxr14-bpals-trial
NIH-CXR14-BPALS — Label-Quality Audit for NIH ChestX-ray14
By O5I | Patent pending (B-PALS) | Contact: hello@o5i.io
Start by finding what's wrong. Refine from there.
What it does
NIH ChestX-ray14's labels are derived automatically from radiology reports, not verified against the images — so a meaningful share are noisy or wrong (~80% reported accuracy). NIH-CXR14-BPALS independently re-examines each (image, label) pair with a vision-language model and returns a… See the full description on the dataset page: https://huggingface.co/datasets/o5i/cxr14-bpals-trial.TriALS-Report
TriALS-Report: A Multi-Center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT
Study workflow. Non-contrast CT volumes are paired with the triphasic contrast-enhanced report of the same patient; findings are extracted from the report to form the label space, and models are evaluated on disease diagnosis and report generation.
TriALS-Report is a multi-centre benchmark for abdominal disease diagnosis from non-contrast CT (NCCT), where the… See the full description on the dataset page: https://huggingface.co/datasets/marwankefah/TriALS-Report.trident
TRIDENT Challenge Dataset
This repository hosts the dataset release for the TRIDENT challenge:
TRIDENT: Tri-modal Deepfake Perception, Detection, and Hallucination Grand Challenge.
The repository was used for Phase 1 with the public train and public_val splits. For Phase 2, the test set has been added and the submission period is open. Participants must run inference on the test set and submit their predictions through the official competition platform. Ground-truth labels and… See the full description on the dataset page: https://huggingface.co/datasets/j1anglin/trident.Trial_01
