furry
Datasets
All datasets matching “furry”moviesfurry_dataset_e621_captions_claims_human_reviewedAbout 13,000 human reviewed captions. ~5000 of them were reviewed by me, the rest by independent workers.
I cannot confirm they will have caught all of the mistakes (there will still be some small mistakes). But the accuracy of these captions is higher than what any VLM can produce.
The "claims" are one-liner claims about an image, and given a truth value. Mostly machine-verified, but the ~10000 human-reviewed ones are a very valuable set of sex-related items or items which powerful VLMs had… See the full description on the dataset page: https://huggingface.co/datasets/furproxy/furry_dataset_e621_captions_claims_human_reviewed.furry-e621-sfw-7m-hq
Dataset Card for furry-e621-sfw-7m-hq
Dataset Summary
This is 6.92 M captions of the images from the safe-for-work (SFW) split of e621 ("e926"). It extends to January 2023, before the widespread advent of machine learning images. It includes captions created by LLMs and a custom multilabel classifier along with CogVLM. There are 8 LLM (mistralai/Mistral-7B-v0.1) and 1 CogVLM (THUDM/CogVLM) captions per image.
Most captions are substantially larger than 77 tokens and are… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/furry-e621-sfw-7m-hq.ManiSkill-Memory-dependence
ManiSkill-Memory-Dependence Benchmark
A comprehensive memory dependence robot benchmark across 4 manipulation tasks from different memory dimensions, introduced in the paper "Non-Markovian Long-Horizon Robot Manipulation via Keyframe Chaining".
Dataset Description
This dataset provides a suite of Non-Markovian manipulation tasks built upon the ManiSkill simulator to measure task success rates in scenarios requiring long-horizon memory and state disambiguation.
It is… See the full description on the dataset page: https://huggingface.co/datasets/furry123/ManiSkill-Memory-dependence.aesir-rpg-furry-markdown-flatguard-splitnewloras
