CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01youssef101 /artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems. The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.imageimage-to-text10K<n<100K2 likes1k downloads3y agoHugging Face02schneewolflabs /ArtemisMix-v1 ArtemisMix-v1 Stage-2 (multimodal instruction fine-tuning) corpus for Artemis, the Schneewolf Labs vision-language flagship that grafts a Qwen3-VL ViT + MLP projector onto the A2 decoder (the A3 lineage). This is the Lite slice (v1): layers L1 + L4 only, 350,000 rows. The planned L2 (multimodal tool/agent) and L3 (custom distill) layers are produced separately and concatenated later. Composition Layer Rows Share Purpose L1 general multimodal instruction… See the full description on the dataset page: https://huggingface.co/datasets/schneewolflabs/ArtemisMix-v1.imagevisual-question-answering100K<n<1M0 likes976 downloads4mo agoHugging Face03artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes530 downloads7mo agoHugging Face04cilabuniba /wikifragments-visual-arts-embeds WikiFragments - Visual Arts Pages with Fragments (WikiFragmentsVA) WikiFragmentsVA is a domain-specific multimodal dataset focused on the visual arts, derived from Wikipedia (en). It consists of textual paragraphs paired with related images (infoboxes and thumbnails), rendered as unified visual fragments. This dataset extends the base WikiFragments project by providing pre-rendered fragment images and multi-vector embeddings obtained via ColQwen2 v1.0, including optimized pooled… See the full description on the dataset page: https://huggingface.co/datasets/cilabuniba/wikifragments-visual-arts-embeds.imagetext-generation1M<n<10M0 likes95 downloads7mo agoHugging Face05Duke-de-Artois /ChemVLM_test_dataUsing this dataset, please kindly cite: @inproceedings{li2025chemvlm, title={Chemvlm: Exploring the power of multimodal large language models in chemistry area}, author={Li, Junxian and Zhang, Di and Wang, Xunzhi and Hao, Zeying and Lei, Jingdi and Tan, Qian and Zhou, Cai and Liu, Wei and Yang, Yaotian and Xiong, Xinrui and others}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={39}, number={1}, pages={415--423}, year={2025} }… See the full description on the dataset page: https://huggingface.co/datasets/Duke-de-Artois/ChemVLM_test_data.imagetext-generation1K<n<10K2 likes90 downloads9mo agoHugging Face06schneewolflabs /ArtemisMix-v1.1 ArtemisMix-v1.1 The as-trained Stage-2 corpus for Artemis1-13B (Schneewolf Labs' Qwen3-VL-ViT-grafted-onto-A2/Mistral vision-language flagship). Extends schneewolflabs/ArtemisMix-v1 with 14,816 non-reasoning creative-writing rows derived from schneewolflabs/Athanorlite-DPO (DPO collapsed to SFT, chosen only, bare assistant content so the A2 chat template renders an empty <think></think> block — i.e. thinking-off mode). The bucket balances v1's reasoning-heavy L4 with a… See the full description on the dataset page: https://huggingface.co/datasets/schneewolflabs/ArtemisMix-v1.1.imagevisual-question-answering100K<n<1M0 likes72 downloads4mo agoHugging Face07AWeirdDev /zh-tw-pts-articles-sm zh-tw-pts-articles-sm 🐣English • 🇹🇼 繁體中文 This dataset contains articles scraped from PNN News. It's a news provider verified by the vast majority. Note: some keys like conclusion may be None. Dataset({ features: ['image', 'title', 'conclusion', 'content', 'timestamp', 'category', 'link'], num_rows: 1400 }) Use The Dataset Use 🤗 Datasets to download, use or modify this dataset. from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-pts-articles-sm.imagetext-generation1K<n<10K7 likes48 downloads3y agoHugging Face08crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes37 downloads6mo agoHugging Face09AWeirdDev /zh-tw-articles-2kHey! Also check out AWeirdDev/zh-tw-pts-articles-sm for a news source verified by the vast majority. zh-tw-articles-2k 🐣English • 🇹🇼 繁體中文 This dataset contains Taiwan news articles scraped from (https://www.storm.mg) on March 2024. Size: 5.0MB (5294263 bytes) Rows: 2000, from 20n20n20n nnn pages: 100 Dataset({ features: ['image', 'title', 'content', 'tag', 'author', 'timestamp', 'link'], num_rows: 2000 }) Use The Dataset Use 🤗 Datasets to download… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-2k.imagetext-generation1K<n<10K3 likes35 downloads2y agoHugging Face10crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes32 downloads1y agoHugging Face11research-artifact-4729 /bankertoolbench BankerToolBench Anonymized submission. Author, institution, and external-link references have been removed from this dataset card for double-blind review. Citation and external links will be restored after the review period. BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for evaluating AI agents. Each task mirrors real junior-banker work — building financial models, preparing pitch decks, writing memos — and produces multi-file deliverables (Excel… See the full description on the dataset page: https://huggingface.co/datasets/research-artifact-4729/bankertoolbench.documenttext-generationn<1K0 likes28 downloads5mo agoHugging Face12sayurio /somewhereinblog-article Somewhereinblog Article Archive Overview This repository contains a large-scale text dataset scraped from m.somewhereinblog.net, the largest and first-ever Bengali community blogging platform. The primary goal of this archive is to preserve a massive collection of purely human-written blog posts, personal stories, socio-political opinions, and community discussions, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/somewhereinblog-article.imagetext-generation10K<n<100K1 likes24 downloads6mo agoHugging Face13AWeirdDev /zh-tw-articles-6k zh-tw-articles-6k This dataset contains Taiwan news articles scraped from (https://www.storm.mg) on March 2024. Size: 10.4MB (15644219 bytes) Rows: 6000 (Max) Dataset({ features: ['image', 'title', 'content', 'tag', 'author', 'timestamp', 'link'], num_rows: 6000 }) Use The Dataset Use 🤗 Datasets to download, use or modify this dataset. from datasets import load_dataset dataset = load_dataset("AWeirdDev/zh-tw-articles-6k") imagetext-generation1K<n<10K0 likes23 downloads3y agoHugging Face14david-sprague /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes17 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.