datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
artelingo-dummyArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. It is an extension of ArtEmis, which is a collection of 80,000 artworks from WikiArt with 450,000 emotion labels and English-only captions. ArtELingo expands this dataset by adding 790,000 annotations in Arabic and Chinese. The purpose of these additional annotations is to evaluate the performance of "cultural-transfer" in AI systems.
The dataset in ArtELingo… See the full description on the dataset page: https://huggingface.co/datasets/youssef101/artelingo-dummy.ArtemisMix-v1
ArtemisMix-v1
Stage-2 (multimodal instruction fine-tuning) corpus for Artemis, the
Schneewolf Labs vision-language flagship that grafts a Qwen3-VL ViT + MLP
projector onto the A2 decoder (the A3 lineage).
This is the Lite slice (v1): layers L1 + L4 only, 350,000 rows.
The planned L2 (multimodal tool/agent) and L3 (custom distill) layers are
produced separately and concatenated later.
Composition
Layer
Rows
Share
Purpose
L1 general multimodal instruction… See the full description on the dataset page: https://huggingface.co/datasets/schneewolflabs/ArtemisMix-v1.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.wikifragments-visual-arts-embeds
WikiFragments - Visual Arts Pages with Fragments (WikiFragmentsVA)
WikiFragmentsVA is a domain-specific multimodal dataset focused on the visual arts, derived from Wikipedia (en). It consists of textual paragraphs paired with related images (infoboxes and thumbnails), rendered as unified visual fragments. This dataset extends the base WikiFragments project by providing pre-rendered fragment images and multi-vector embeddings obtained via ColQwen2 v1.0, including optimized pooled… See the full description on the dataset page: https://huggingface.co/datasets/cilabuniba/wikifragments-visual-arts-embeds.ChemVLM_test_dataUsing this dataset, please kindly cite:
@inproceedings{li2025chemvlm,
title={Chemvlm: Exploring the power of multimodal large language models in chemistry area},
author={Li, Junxian and Zhang, Di and Wang, Xunzhi and Hao, Zeying and Lei, Jingdi and Tan, Qian and Zhou, Cai and Liu, Wei and Yang, Yaotian and Xiong, Xinrui and others},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={39},
number={1},
pages={415--423},
year={2025}
}… See the full description on the dataset page: https://huggingface.co/datasets/Duke-de-Artois/ChemVLM_test_data.ArtemisMix-v1.1
ArtemisMix-v1.1
The as-trained Stage-2 corpus for Artemis1-13B (Schneewolf Labs'
Qwen3-VL-ViT-grafted-onto-A2/Mistral vision-language flagship).
Extends schneewolflabs/ArtemisMix-v1
with 14,816 non-reasoning creative-writing rows derived from
schneewolflabs/Athanorlite-DPO
(DPO collapsed to SFT, chosen only, bare assistant content so the A2
chat template renders an empty <think></think> block — i.e. thinking-off
mode). The bucket balances v1's reasoning-heavy L4 with a… See the full description on the dataset page: https://huggingface.co/datasets/schneewolflabs/ArtemisMix-v1.1.zh-tw-pts-articles-sm
zh-tw-pts-articles-sm
🐣English • 🇹🇼 繁體中文
This dataset contains articles scraped from PNN News.
It's a news provider verified by the vast majority.
Note: some keys like conclusion may be None.
Dataset({
features: ['image', 'title', 'conclusion', 'content', 'timestamp', 'category', 'link'],
num_rows: 1400
})
Use The Dataset
Use 🤗 Datasets to download, use or modify this dataset.
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-pts-articles-sm.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.zh-tw-articles-2kHey! Also check out AWeirdDev/zh-tw-pts-articles-sm for a news source verified by the vast majority.
zh-tw-articles-2k
🐣English • 🇹🇼 繁體中文
This dataset contains Taiwan news articles scraped from (https://www.storm.mg) on March 2024.
Size: 5.0MB (5294263 bytes)
Rows: 2000, from 20n20n20n
nnn pages: 100
Dataset({
features: ['image', 'title', 'content', 'tag', 'author', 'timestamp', 'link'],
num_rows: 2000
})
Use The Dataset
Use 🤗 Datasets to download… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-articles-2k.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.bankertoolbench
BankerToolBench
Anonymized submission. Author, institution, and external-link references
have been removed from this dataset card for double-blind review. Citation
and external links will be restored after the review period.
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel… See the full description on the dataset page: https://huggingface.co/datasets/research-artifact-4729/bankertoolbench.somewhereinblog-article
Somewhereinblog Article Archive
Overview
This repository contains a large-scale text dataset scraped from m.somewhereinblog.net, the largest and first-ever Bengali community blogging platform. The primary goal of this archive is to preserve a massive collection of purely human-written blog posts, personal stories, socio-political opinions, and community discussions, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/somewhereinblog-article.zh-tw-articles-6k
zh-tw-articles-6k
This dataset contains Taiwan news articles scraped from (https://www.storm.mg) on March 2024.
Size: 10.4MB (15644219 bytes)
Rows: 6000 (Max)
Dataset({
features: ['image', 'title', 'content', 'tag', 'author', 'timestamp', 'link'],
num_rows: 6000
})
Use The Dataset
Use 🤗 Datasets to download, use or modify this dataset.
from datasets import load_dataset
dataset = load_dataset("AWeirdDev/zh-tw-articles-6k")
Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.
