CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes9.2k downloads7mo agoHugging Face02artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes535 downloads7mo agoHugging Face03MelosY /TextMonkey_Dataimage10K<n<100K4 likes364 downloads2y agoHugging Face04AbdullahRian /Korean.OCR.Img.text.pairimage100K<n<1M1 likes257 downloads1y agoHugging Face05nyuuzyou /cc0-textures Dataset Card for CC0 Textures Dataset Summary This dataset contains 18,785 texture images from cc0-textures.com. It includes textures of wood, metal, concrete, fabric, stone, ceramic, and other materials. The original archives were downloaded, unpacked, and images were compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while keeping good quality. Languages The dataset is monolingual: English (en): Texture titles and tags… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cc0-textures.imageimage-classification10K<n<100K5 likes170 downloads1y agoHugging Face06sheng22213 /speech_text-tts_audioaudio10K<n<100K0 likes134 downloads1y agoHugging Face07cg1177 /qvhighlight_internvideo2_llama_text_featuretextn<1K0 likes103 downloads2y agoHugging Face08alfredplpl /image-text-pairs-ja-cc0-2 はじめに このデータセットは画像生成で日本語を生成したいときに使うデータセットです。 ライセンス CC-0です。著作権を放棄して使いやすくしました。 作り方の概要 gpt-oss-20bを使って、約8万個からなる単語集兼短文集を作りました。 その文章をPillowとPythonでランダム要素を入れながら100万枚と10万枚でレンダリングしました。 フォントはNoto Sans JPなのでライセンス的には問題ないと思います。 image1M<n<10M3 likes100 downloads1y agoHugging Face09koorye /Describable-Textures-Datasetimage1K<n<10K0 likes75 downloads10mo agoHugging Face10nyuuzyou /textureninja Dataset Card for Texture Ninja Dataset Summary This dataset contains 4,540 texture images from texture.ninja. It includes high-resolution textures of brick, concrete, rock, wood, metal, paint, plaster, ground materials, and other surfaces. The original images were downloaded, processed, and compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while maintaining good quality. Languages The dataset is monolingual: English (en):… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/textureninja.imageimage-classification1K<n<10K0 likes38 downloads1y agoHugging Face11rivisia /text-to-image-2M text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset Overview text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/rivisia/text-to-image-2M.imagetext-to-image100K<n<1M0 likes34 downloads9mo agoHugging Face12cg1177 /charade_sta_internvideo2_llama_text_featuretextn<1K0 likes29 downloads2y agoHugging Face13alfredplpl /image-text-pairs-ja-cc0 Japanese Glyph Images with English Captions (CC0) This dataset contains Japanese glyph images rendered with black text on white background. Each .png image has a corresponding .txt file with an English caption: This image is saying "<Japanese>". The background is white. The letter is black. Structure train/ — PNG images and matching TXT captions (same base filename) provenance/assets_registry.csv — Fonts and license info LICENSE.txt — CC0-1.0 license Generation… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/image-text-pairs-ja-cc0.image10K<n<100K4 likes27 downloads1y agoHugging Face14AngelBottomless /TextAtlas5M-CoverBookThis is repacked dataset of https://huggingface.co/datasets/CSU-JPG/TextAtlas5M with "annotations only" / "images only". image100K<n<1M0 likes23 downloads2y agoHugging Face15tekkithorse /mlp-board-yuki-archive-html-textTaken from the yuki archive data. Contains the html text from the /mlp/ board up until a few years ago. It sure would be interesting to have a dataset that's organized by each thread in a spreadsheet separated by row with the posts in the thread separated by column, with the post number in each cell. text100K<n<1M0 likes21 downloads3y agoHugging Face16vindows /qwen2.5-text-to-sql-benchmark-data Qwen2.5 Text-to-SQL Benchmark Data & Results This dataset contains Spider benchmark data and evaluation results for fine-tuned Qwen2.5 models. Contents qwen-spider-data.tar.gz (1.3MB) Spider benchmark training data (~7,000 examples) Spider dev set (1,034 examples) Table schemas (tables.json) Official evaluation scripts: evaluation.py - Official Spider evaluation process_sql.py - SQL parsing utilities README.md - Spider benchmark documentation… See the full description on the dataset page: https://huggingface.co/datasets/vindows/qwen2.5-text-to-sql-benchmark-data.textn<1K0 likes21 downloads9mo agoHugging Face17wsdwJohn1231 /DreamLIP_short_texttext10M<n<100M0 likes17 downloads10mo agoHugging Face18yiyangd /bridge_orig_text_20kimage10K<n<100K0 likes11 downloads8mo agoHugging Face19dddraxxx /text-image-finetune-1k-v2image1K<n<10K0 likes9 downloads8mo agoHugging Face20Kondapally /AIWD16-Textimage10K<n<100K0 likes9 downloads8mo agoHugging Face21pavi1561 /Divehi_text_speech_datasetaudion<1K0 likes7 downloads2y agoHugging Face22Junaid7188 /text-to-image-2M text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset Overview text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/Junaid7188/text-to-image-2M.imagetext-to-image100K<n<1M0 likes6 downloads4mo agoHugging Face23shindeaditya /text-to-image-cleanedimage10K<n<100K0 likes5 downloads2y agoHugging Face24JiaHuang01 /prompt_texttextn<1K0 likes4 downloads8mo agoHugging Face25Chengxiang1122 /mmcl-textvqa textvqa Repo: Chengxiang1122/mmcl-textvqa Visibility: public Included archives archives/train.tar.gz source: /g/data/cp23/ch3329/data/textvqa/TextVQA_0.5.1_train.json archives/val.tar.gz source: /g/data/cp23/ch3329/data/textvqa/TextVQA_0.5.1_val.json Each archive preserves the original source path contents for the corresponding split. textn<1K0 likes4 downloads6mo agoHugging Face26Disty0 /sotediffusion-v1-text_onlygatedWD: SmilingWolf/wd-swinv2-tagger-v3GPU: 1x Intel ARC A770 16 GB NL: vikhyatk/moondream2GPU: 8x NVIDIA H100 80 GB SXM5 BLIP: Salesforce/xgen-mm-phi3-mini-instruct-r-v1GPU: 1x AMD RX 7900 XTX 24 GB TEXT: llava-hf/llava-1.5-7b-hfGPU: 1x AMD RX 7900 XTX 24 GB + 1x Intel ARC A770 16 GB AESTHETICS (Only on WD): shadowlilac/aesthetic-shadow-v2GPU: 1x AMD RX 7900 XTX 24 GB + 1x Intel ARC A770 16 GB QUALITY (Only on WD):… See the full description on the dataset page: https://huggingface.co/datasets/Disty0/sotediffusion-v1-text_only.texttext-to-image10M<n<100M3 likes3 downloads2y agoHugging Face27ToniDO /TeXtract_datasetgated TeXtract_dataset (WebDataset Format) This repository contains approximately 3.2 million pairs of mathematical expression images and their corresponding LaTeX source code, packaged in WebDataset format for large-scale training. The dataset is based on and derived from the original hoang-quoc-trung/fusion-image-to-latex-datasets, transformed for more efficient access. 📂 Dataset Structure Each WebDataset shard (.tar) contains multiple samples. Each sample groups files… See the full description on the dataset page: https://huggingface.co/datasets/ToniDO/TeXtract_dataset.image1M<n<10M0 likes3 downloads1y agoHugging Face28tbd-lab /qwen-image-text-renderingimage10K<n<100K0 likes3 downloads11mo agoHugging Face29trirn /speech-to-textaudio100K<n<1M0 likes2 downloads7mo agoHugging Face30IEMaster /worldedit_text_filteredgatedimage10K<n<100K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.