vision-text
SmolVLM2-256M-Video-Instruct-vision-W8A16-text-W4A16-G64Text_to_Visiontiny-random-VisionTextDualEncoderModel-vit-bertsd15.realistic_vision.v5_1.text_encoderSmolVLM2-500M-Video-Instruct-vision-BF16-text-W8A16-ASYMSmolVLM2-500M-Video-Instruct-vision-W8A16-text-W4A16-G64-ASYMEElayoutlmv3_jordyvl_rvl_cdip_100_examples_per_class_2023-08-12_text_vision_onlygemma-like-multimodal-speech-vision-text
Vision-DeepResearch-Text-Datauber_text-Vision-QA
Dataset Card for "uber_text_qa"
More Information needed
text-vision-audio-2k-testA 2k sample dataset for testing multimodal (text+vision+audio) format. This is compatible with HF's processor apply_chat_template.
Load in Axolotl via:
datasets:
- path: Nanobit/text-vision-audio-2k-test
type: chat_template
Make sure to download the image and audio via:
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/African_elephant.jpg
wget https://huggingface.co/datasets/Nanobit/text-vision-audio-2k-test/resolve/main/En-us-African_elephant.oga… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-audio-2k-test.text-vision-shieldstral-2k-testA 2k sample dataset for testing the Shieldstral multimodal moderation format. Each sample is a fixed system prompt, a [text, image, text] user message, and a single yes/no answer.
Load in Axolotl via:
datasets:
- path: Nanobit/text-vision-shieldstral-2k-test
type: chat_template
Make sure to download the image via:
wget https://huggingface.co/datasets/Nanobit/text-vision-shieldstral-2k-test/resolve/main/African_elephant.jpg
Image source:… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-shieldstral-2k-test.text-vision-2k-testA 2k sample dataset for testing multimodal (text+vision) format. This is compatible with HF's processor apply_chat_template.
Load in Axolotl via:
datasets:
- path: Nanobit/text-vision-2k-test
type: chat_template
Make sure to download the image via:
wget https://huggingface.co/datasets/Nanobit/text-vision-2k-test/resolve/main/African_elephant.jpg
Image source: https://upload.wikimedia.org/wikipedia/commons/e/ec/African_elephant.jpg
Each sample has the following format and is repeated 2k… See the full description on the dataset page: https://huggingface.co/datasets/axolotl-ai-co/text-vision-2k-test.Text-Before-VisionGitHub: https://github.com/MiliLab/Text-Before-Vision
