CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face02jrzhang /TextVQA_GT_bbox TextVQA validation set with grounding truth bounding box The dataset used in the paper MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs for studying MLLMs' attention patterns. The dataset is sourced from TextVQA and annotated manually with ground-truth bounding boxes. We consider questions with a single area of interest in the image so that 4370 out of 5000 samples are kept. Citation If you find our paper and code useful… See the full description on the dataset page: https://huggingface.co/datasets/jrzhang/TextVQA_GT_bbox.imagequestion-answering1K<n<10K4 likes723 downloads1y agoHugging Face03artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes530 downloads7mo agoHugging Face04yonilev /Text2Receipt Text2Receipt Messy free-text Hebrew income notes -> valid, complete Israeli fiscal documents (receipts & tax invoices). Live demo (Space): yonilev/Text2Receipt Dataset: yonilev/Text2Receipt Dataset Creation A synthetic corpus from a deterministic, rule-based generator plus a bounded LLM-paraphrase layer, so the ground truth is exact by construction. Pipeline Scenario sampling - category, issuer status, document type, client type, year, payment… See the full description on the dataset page: https://huggingface.co/datasets/yonilev/Text2Receipt.imagetext-generation10K<n<100K0 likes103 downloads3mo agoHugging Face05crawlfeeds /Curated-Fox-News-Headlines-and-Full-Text Curated Fox News Headlines and Full Text This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis. 📁 Dataset Format Format: CSV Encoding: UTF-8 Fields: headline: The article title or headline publish_date: Date the article was published (YYYY-MM-DD) content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.imagetext-classification1K<n<10K2 likes89 downloads1y agoHugging Face06REILX /Chinese-Image-Text-Corpus-dataset REILX/Chinese-Image-Text-Corpus-dataset [ English | 中文 ] Introduction The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms. Dataset Structure The dataset is organized into the following categories: Idioms: Traditional Chinese idioms with explanations and… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.imagequestion-answering100K<n<1M0 likes47 downloads2y agoHugging Face07Windsao /eis-text250 EIS-Text250: 1970s U.S. Environmental Impact Statements (text-only) Per-page OCR/extraction text for 250 scanned 1970s U.S. federal Environmental Impact Statements (EIS) from the Northwestern University Library collection — the text-only companion to Windsao/eis-subset50 (which carries full page images for a 50-doc subset). Built to test how current models handle long, dense, historical government text: mean ~300 pages/doc, 1970s typewriter prose, OCR noise from degraded… See the full description on the dataset page: https://huggingface.co/datasets/Windsao/eis-text250.imagetext-generation10K<n<100K0 likes32 downloads2mo agoHugging Face08hkust-gz-w2 /PDD3_text_rendered_v2gatedimagetext-generation0 likes1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.