datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mscoco-controlnet-canny-less-colorsfineweb_synth_dense_ocr_colors_with_groundingcolorscolorswap
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
Dataset Description
ColorSwap is a dataset designed to assess and improve the proficiency of multimodal models in matching objects with their colors. The dataset is comprised of 2,000 unique image-caption pairs, grouped into 1,000 examples. Each example includes a caption-image pair, along with a "color-swapped" pair. Crucially, the two captions in an example have the same words, but the color words have… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/colorswap.Colors
Dataset Details
Dataset Description
A dataset that contains the color names and their relations with their respective RGB Values with over 40k rows.
Repository: Generation-of-Colors-using-BiLSTMs
Citation:
@misc{sinha2023generation,
title={Generation Of Colors using Bidirectional Long Short Term Memory Networks},
author={A. Sinha},
year={2023},
eprint={2311.06542},
archivePrefix={arXiv},
primaryClass={cs.CV}
ColorSpectrogram
Google/MusicCapsの音楽をスペクトログラムにしたもの
Google/MusicCapsのスペクトログラム。カラーバージョンも作っておく.
基本情報
sampling_rate: int = 44100
参考資料とメモ
(memo)ぶっちゃけグレースケールもカラーバージョンをtorchvision.transformのグレースケール変換すればいいだけかも?
ダウンロードに使ったコードはこちら
参考:https://www.kaggle.com/code/osanseviero/musiccaps-explorer
仕組み:Kaggleの参考コードでwavファイルをダウンロードする->スペクトログラムつくりながらmetadata.jsonlに{"filename":"spectrogram_*.png", "caption":"This is beautiful music"}
をなどと言ったjson列を書き込み、これをアップロードした… See the full description on the dataset page: https://huggingface.co/datasets/mickylan2367/ColorSpectrogram.recolored-object-colors-devcoco_colors_cleaned
coco_colors_cleaned
The coco_colors__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
117,332
QA turns
570,142
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
109
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/coco_colors_cleaned.geo3k-colorshiftcolor_swatchcss-colorsascii_colors_discussions
Dataset Card for ParisNeo/ascii_colors
Dataset Description
This dataset contains synthetic, structured conversations designed to encapsulate the knowledge within the ascii_colors Python library (specifically around version 0.8.1). The primary goal of this dataset is to facilitate the fine-tuning of Large Language Models (LLMs) to become experts on the ascii_colors library, capable of answering questions and performing tasks related to it without relying on… See the full description on the dataset page: https://huggingface.co/datasets/ParisNeo/ascii_colors_discussions.GenEval_prompts_w_two_colorscolor_squarescolors-normalized
Color Names Normalized
A dataset for normalizing messy, multilingual, free-text color names to a
small, fixed vocabulary. Real-world color attributes — product feeds,
marketplace listings, survey answers — are free text with thousands of
variants. This dataset maps 38,112 real-world color names (English and
French) to a strict 20-color base palette — e.g. "rouge", "dark navy",
"burgundy", "whispering grasslands" all resolve to a canonical base color
— so color data becomes… See the full description on the dataset page: https://huggingface.co/datasets/NacerKr/colors-normalized.coco_colors_hatawColorswikipedia-colorsList of colors from Wikipedia
wikiart_5k_colorsspurious_colorscountry_flag_colorscountry_to_flag_colorsPrueba_colorsveggies_to_colors_v1colors
