datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openbrush-landscapes
OpenBrush Landscapes
Every landscape painting from OpenBrush-75K — across all artists, movements, and centuries. Largest single-genre subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,612 you actually want.
Why this subset
Every landscape across the parent dataset's full range — Romantic wildernesses, Impressionist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-landscapes.openbrush-impressionist-landscapes
OpenBrush Impressionist Landscapes
Cross-cut subset: Impressionist landscape paintings from OpenBrush-75K. The most-targeted style+genre combination for Impressionist landscape LoRA training.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,308 you actually want.
Why this subset
The intersection of the largest movement… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionist-landscapes.top_landscapedistributional-landscapesReasonAVEdit-Bench-Landscape
ReasonAVEdit-Bench — Landscape Edition
2,313 samples, every one of them landscape, for instruction-guided joint audio-video
editing. Two halves:
Half
N
What it is
Newly sampled
1,200
15 reasoning sub-tracks, selected from measured evidence
Legacy tracks
1,113
Every landscape sample from the hand-curated AniAVEditBench tracks
The benchmark measures not just the final edit but the multimodal reasoning behind it:
which region to change, which sound source to… See the full description on the dataset page: https://huggingface.co/datasets/bigfacing/ReasonAVEdit-Bench-Landscape.finetuning-landscape-paintingThis dataset was compiled for the purpose of finetuning models in the context of benchmarking for art historical research.
Images scraped from Wikimedia Commons via Wikidata; metadata scraped from
Wikidata (CC0). Image licenses vary per file (predominantly public domain,
some CC-BY-SA) see the license_short_name / license_url columns in the
parquet files for the exact terms of each individual image, and the
commons file page for full details.
ai-bias-research-landscape
Dataset Card: AI Bias Research Landscape
Dataset Summary
This dataset contains 692 curated bibliographic records of peer-reviewed and
preprint publications on artificial intelligence (AI) and algorithmic bias,
published between 2012 and 2026. Each record includes publication metadata
(paper title, DOI, authors, author regions, affiliations, publication year,
and research domain), author ORCID identifiers, and OpenAlex-derived
metadata, including OpenAlex IDs… See the full description on the dataset page: https://huggingface.co/datasets/cair-nepal/ai-bias-research-landscape.llm-sensitivity-landscape
LLM Sensitivity Landscape: Semantic Divergence Under Input Perturbation
Systematic analysis of Gemma4 (e2b) semantic divergence under input perturbation using 100 TruthfulQA questions.
Dataset Summary
This dataset measures how much a language model's response changes when:
System prompt changes (skeptical, literal, creative)
Input is randomly perturbed (word swaps)
Same question is asked twice (baseline vs perturbed baseline)
Divergence is measured as 1 -… See the full description on the dataset page: https://huggingface.co/datasets/bjornshomelab/llm-sensitivity-landscape.vc_landscape_datasetclinical-temporal-5node-pressure-buf-lag-cpl-competitive-landscape-v0.1
What this repo does
This dataset tests whether a model can detect manufacturing drift forming over time and predict whether the program crosses into supply disruption lock-in by the final step.
Core quad
pressurebufferlagcoupling
Prediction target
label_cascade_state
Row structure
One row represents a short temporal window (t0–t3) across program months. It includes time-series values for pressure (deviations and schedule stress), buffer capacity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-temporal-5node-pressure-buf-lag-cpl-competitive-landscape-v0.1.biosecurity_landscape
