CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01StarTrail-org /pixelrag-tiles PixelRAG tile corpus Rendered screenshot tiles for PixelRAG, a visual retrieval-augmented-generation system that retrieves over page images instead of parsed text. Each Wikipedia page is rendered to an image and cut into fixed-height tiles; retrieval runs on the tiles directly with a Qwen3-VL embedding model. This repository holds the full tile corpus that the published FAISS indexes and embeddings were built from, so the whole pipeline (tiles → embeddings → index → search) can… See the full description on the dataset page: https://huggingface.co/datasets/StarTrail-org/pixelrag-tiles.textimage-to-textn>1T0 likes2k downloads3mo agoHugging Face02owkin /plism-dataset-tiles PLISM dataset The Pathology Images of Scanners and Mobilephones (PLISM) dataset was created by (Ochi et al., 2024) for the evaluation of AI models’ robustness to inter-institutional domain shifts. All histopathological specimens used in creating the PLISM dataset were sourced from patients who were diagnosed and underwent surgery at the University of Tokyo Hospital between 1955 and 2018. PLISM-wsi consists in a group of consecutive slides digitized under 7 different scanners and… See the full description on the dataset page: https://huggingface.co/datasets/owkin/plism-dataset-tiles.imageimage-feature-extraction1M<n<10M9 likes1.3k downloads2y agoHugging Face03ryankim17920 /nanopath-fairness-tiles nanopath-fairness-tiles Pre-tiled histopathology patches from CPTAC whole-slide images, used as the external out-of-distribution validation set for a study on pretraining-time vs. post-hoc fairness in histopathology foundation models. Contents Per-cohort folders, each slides_full/<slide_id>.parquet (one row per tile: case_id, slide_id, tile_idx, image) + labels.tsv: cohort organ / task slides cptac_lung NSCLC — LUAD vs LSCC subtype 604 cptac_gbm GBM —… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/nanopath-fairness-tiles.textimage-classification1K<n<10K0 likes609 downloads3mo agoHugging Face04F1nnSBK /onlypits-lunar-tilestabular10K<n<100K0 likes467 downloads6d agoHugging Face05tilde-research /popcorn-reportstabular100K<n<1M0 likes379 downloads1mo agoHugging Face06medarc /gtex-10M-balanced-tiles GTEx 10M Balanced Tiles This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed. Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.tabularimage-feature-extraction10M<n<100M0 likes274 downloads3mo agoHugging Face07crafiq /game-asset-tiles Game Asset Tiles This dataset contains 275 samples, each consisting of a tile image, a matching template image and a description. Tiles were generated with Nano Banana 2 and FLUX.2, or sourced from kenney.nl. hex / oct / rect describe the footprint shape, flat-top or pointy-top describe the hex shape orientation, and isometric, oblique, or top-down describe the perspective. Repository layout: tiles/: PNG tile images descriptions/: Short English text prompts coordinates/: JSON point… See the full description on the dataset page: https://huggingface.co/datasets/crafiq/game-asset-tiles.imagen<1K0 likes266 downloads5mo agoHugging Face08tilyupo /trivia_qa Dataset Card for "trivia_qa_passages" More Information needed text100K<n<1M1 likes263 downloads3y agoHugging Face09ai4chems /tilingtopomltabular1K<n<10K0 likes219 downloads29d agoHugging Face10lsr42 /msmarco-passage-tilde Dataset Card for "msmarco-passage-tilde" More Information needed text1M<n<10M0 likes215 downloads4y agoHugging Face11rabbit-hmi /MM-Mind2Web-tilde_test_snapshot_20dist MultiModal-Mind2Web~ (MM-Mind2Web~) rabbit inc. [Leaderboard & Blogpost to be released] Configuration: test split, snapshot with seed 42, 20 distractors Multimodal-Mind2Web is a dataset proposed by Boyuan et al.. It's designed for the development and evaluation of generalist web agents and includes various action trajectories of humans on real websites. We've simplified the raw dump from both Multimodal-Mind2Web and Mind2Web into sequences of observation-action pairs. We've… See the full description on the dataset page: https://huggingface.co/datasets/rabbit-hmi/MM-Mind2Web-tilde_test_snapshot_20dist.texttext-generation1K<n<10K2 likes191 downloads2y agoHugging Face12tilyupo /nq_text Dataset Card for "nq_text" More Information needed text100K<n<1M0 likes190 downloads3y agoHugging Face13wzqacky /mahjong_tilesimage10K<n<100K0 likes180 downloads2mo agoHugging Face14tilyupo /mmlu Dataset Card for "mmlu" More Information needed text1M<n<10M0 likes177 downloads3y agoHugging Face15TildeAI /TildeOpen-MT-Euro-mix TildeOpen-MT-Euro-mix TildeOpen-MT-Euro-mix is a machine-translation datamix for 30 European languages. It was built for the translation instruction-tuning experiments used to validate the TildeOpen LLM (Bergmanis et al., LREC 2026). We publish this data to support the reproducibility of our experiments. Composition Datamix contains 1,890,652 sentence/document pairs across 245 translation directions and 30 languages - 221.44M source-side and 240.56M target-side… See the full description on the dataset page: https://huggingface.co/datasets/TildeAI/TildeOpen-MT-Euro-mix.texttranslation1M<n<10M1 likes164 downloads21d agoHugging Face16tilak1114 /deepfashion DeepFashion Multimodal Dataset This dataset contains annotations for shape, fabric, and texture attributes of clothing images. Metadata id_to_category: { "0": "Sleeve Length", "1": "Lower Clothing Length", "2": "Socks", "3": "Hat", "4": "Glasses", "5": "Neckwear", "6": "Wrist Wearing", "7": "Ring", "8": "Waist Accessories", "9": "Neckline", "10": "Outer Clothing a Cardigan?", "11": "Upper Clothing Covering Navel?" } category_options: {… See the full description on the dataset page: https://huggingface.co/datasets/tilak1114/deepfashion.image10K<n<100K1 likes147 downloads2y agoHugging Face17LynnMass /tilebac-flag-dataset Background Overview This dataset contains ultralow-dose cryoEM montage tile images of the bacteria Pantoea sp. YR343. Segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 have been fine-tuned to predict bacterial inner membranes and outer membranes. This bacterial membrane dataset is a benchmark dataset to challenge current AI workflows in rapid, seamless montage stitching and stitched segmentation in extremely noisy ultralow-dose cryoEM images. Bacterial flagella low-dose… See the full description on the dataset page: https://huggingface.co/datasets/LynnMass/tilebac-flag-dataset.imageimage-segmentation1K<n<10K0 likes136 downloads4mo agoHugging Face18kshitijrajsharma /osm-completeness-tilestabular10M<n<100M0 likes133 downloads19d agoHugging Face19jesbu1 /molmoact2_yam_tiled_rfmtext10K<n<100K0 likes129 downloads24d agoHugging Face20tilikumotp /Global-Ocean-Science-Corpus 🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned) A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes. Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.tabulartext-generation10K<n<100K1 likes116 downloads2d agoHugging Face21thetemirbolatov /TILO.RA_CODER_Dataset TILO.RA CODER Dataset Объединённый русско-английский датасет для обучения и поиска по коду. Формат — пары question / code: вопрос на естественном языке → готовый код-ответ. Датасет собран для локального ассистента TILO.RA CODER — офлайн-помощника по программированию Скачать по ссылке https://github.com/thetemirbolatov/TILO.RA_CODER_Dataset/releases/download/v1.0.0/tilora_knowledge_merged.jsonl Состав Источник Язык Записей English coding… See the full description on the dataset page: https://huggingface.co/datasets/thetemirbolatov/TILO.RA_CODER_Dataset.texttext-generation100K<n<1M1 likes102 downloads11d agoHugging Face22AI4Manufacturing /219-tilesgated Roles Roles: canon repo — annot is the source label, kept machine-parseable as the gold for verification and reward parsing; there is no filled reasoning column and this repo is not itself a training view. Derived repos each state their own regime on their own card. 219-tiles Tile-level hot-crack detection in light-microscope images of a weld surface — the trainable view of 219 (512x512 native-resolution tiles; binary masks). ⚠ Microscopy, NOT production-line… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/219-tiles.image1K<n<10K0 likes97 downloads4d agoHugging Face23casual /til_audioaudio1K<n<10K0 likes77 downloads2y agoHugging Face24tilak1114 /necklines-attr Necklines Attribute Dataset This dataset contains images of different neckline styles categorized by type. Structure category: Main category of the neckline subcategory: Specific type of neckline image_url: Source URL of the image image: The actual image data (PNG or JPEG), resized if dimensions exceed 1000px image10K<n<100K0 likes75 downloads2y agoHugging Face25cornhundred /SpatialData_with_spatial_tilestabular100M<n<1B0 likes74 downloads12d agoHugging Face26tilak1114 /fashion-pediaimage10K<n<100K0 likes72 downloads1y agoHugging Face27Tilakoid /vscode-bug-feature-triage VS Code Bug vs Feature Request Triage Dataset summary 1,993 prepared issue records from public microsoft/vscode issues, reduced to one binary task: classify the issue text as bug or feature-request. The splits are a frozen temporal holdout (80/10/10 by created_at within each class, seed 42) used by the GitHub Triage SLM Fine-Tuning Benchmark to compare fine-tuned small models against their base checkpoints on the same test set. Each record carries cleaned issue… See the full description on the dataset page: https://huggingface.co/datasets/Tilakoid/vscode-bug-feature-triage.texttext-classification1K<n<10K0 likes72 downloads7d agoHugging Face28lzh7522 /til_nlp_test_dataset Dataset Card for "til_nlp_test_dataset" More Information needed text10K<n<100K0 likes50 downloads3y agoHugging Face29TILKI-AI /conversations_seed_storytext100K<n<1M0 likes50 downloads5mo agoHugging Face30tilyupo /marco_cqa Dataset Card for "marco_cqa" More Information needed text100K<n<1M0 likes49 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.