datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lucid
LUCID — Lunar Captioned Image Dataset
LUCID is a large-scale multimodal dataset for vision-language training on real lunar surface observations. It is introduced as part of the paper "LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration" (Inal et al., 2025, under review).
To the best of our knowledge, LUCID is the first publicly available dataset for multimodal vision-language training on real planetary observations.
Links
📄 Paper: LLaVA-LE: Large… See the full description on the dataset page: https://huggingface.co/datasets/pcvlab/lucid.bankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/lucids112/bankertoolbench.meisho-sft
Meisho SFT
Meisho SFT is a dataset containing 10K curated and fully captioned, high quality images for the fine-tuning of diffusion models.
Usage
from datasets import load_dataset
dataset = load_dataset("lucidityai/meisho-sft", split="train")
Curation
Filtering
This dataset was filtered based via a two-stage process:
AI Detection
We used our Akita-1 AI image detection model to determine which images were AI generated with a 90%… See the full description on the dataset page: https://huggingface.co/datasets/LucidityAI/meisho-sft.LUCID-SOILevante
sol levante - anime super resolution dataset
dataset for training anime upscaling models.
what is this
downloaded the frames from netflix's open content anime Sol Levante
processed it with the LUCID dataset maker: https://github.com/Phhofm/lucid-sisr
license is apache 2.0 (original netflix content is cc-by-4.0)
note on the "weird" noise
if you see some weird noise or grain in the frames, i noticed it too. its not from the processing tool. the… See the full description on the dataset page: https://huggingface.co/datasets/XYETHER/LUCID-SOILevante.
