datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
results
TrueVisLies – Results
This dataset contains all raw outputs, extracted fields, semantic similarity scores, and UMAP projections produced in the paper:
True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies
The paper evaluates 16 LLMs, 15 open-weight vision-language models (VLMs), and GPT-5.4 on their ability to (RQ0) detect misleading data visualizations, (RQ1) identify the visualization rhetoric techniques, and… See the full description on the dataset page: https://huggingface.co/datasets/truevislies/results.radread-public-results
RadRead — public results
Rollout-level results for RadRead, a benchmark of frontier models reading 150
radiographs. Every row is one graded model read: 5 saved rollouts per study
per model, scored by a deterministic grader (no judge model).
A read passes only when every required checklist finding, lesion box (the grader's
IoU / centre / containment test), lexical diagnosis check and action-set membership
check match the reference rubric. No partial credit inside a study;… See the full description on the dataset page: https://huggingface.co/datasets/tirandazdylan/radread-public-results.revisable-vlm-memory-results
Revisable VLM memory: artifacts, results, and plans
This dataset repository archives frozen representation artifacts, evaluation
reports, causal-intervention outputs, and preregistered execution plans for
Yunbo-max/revisable-vlm-memory.
It does not redistribute TAP-Vid images or model weights.
Current status (2026-08-18)
The frozen 3B K/V diagnostics previously showed held-out signal for simple
visibility and moving/static variables under their original two-frame… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/revisable-vlm-memory-results.aidm-dogs-vs-cats-results
aidm-dogs-vs-cats-results
The experiment record of a dogs-vs-cats image-classification study, with CIFAR-10 and CIFAR-10-LT transfer and class-imbalance ablations. This repo holds the run registry, the splits, the report tables and figures, and the per-run predicted probabilities. It holds no images and no model weights; the checkpoints are in the companion model repo.
Generated by scripts/90_publish_hf.py on 2026-09-22 19:17 UTC. Every count, fingerprint and metric below was… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/aidm-dogs-vs-cats-results.
