CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes535 downloads7mo agoHugging Face02mishl /StreetVision-10K StreetVision-10K Each sample contains: A system prompt instructing the model to act as an OSINT/geospatial expert A user message with a street-level photo and the instruction to determine coordinates An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format Format Each line is a JSON array of ChatML messages: [ {"role": "system", "content": "..."}, {"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.imagevisual-question-answering10K<n<100K1 likes159 downloads6mo agoHugging Face03PJMixers-Images /bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219. The format should be similar to that of liuhaotian/LLaVA-Instruct-150K. Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.imagetext-generation1K<n<10K1 likes66 downloads2y agoHugging Face04ArkaMukherjee /reasoning-10k-v2 🧠 Dataset for Vision-Language Reasoning Reasoning-10K-v2 is a 10,000-sample multimodal reasoning dataset. Our work unifies diverse sources of vision–language reasoning data spanning code, mathematics, geography, tables, and scientific figures. Using a hybrid of synthesis and filtration strategies inspired by prior baselines (LIMO VL, MM MathInstruct, and Multimodal Open R1), we curated high-quality reasoning examples from open and synthetic data. The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/ArkaMukherjee/reasoning-10k-v2.documenttext-generation10K<n<100K2 likes56 downloads11mo agoHugging Face05TurkishCodeMan /recipe-synthetic-images-10k Recipe PDF Dataset A multimodal dataset of 10K+ recipes rendered as PDF images with full metadata. Dataset Description Each sample contains: image: Recipe rendered as a styled PDF page (PNG, ~1654x2339px) name: Recipe title description: Recipe description ingredients: List of ingredients steps: Cooking instructions nutrition: Nutritional values (calories, fat%, sugar%, sodium%, protein%, sat.fat%, carbs%) random_reviews: User reviews minutes: Cooking time tags: Recipe… See the full description on the dataset page: https://huggingface.co/datasets/TurkishCodeMan/recipe-synthetic-images-10k.imageimage-to-text10K<n<100K0 likes51 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.