datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pixelrag-tiles
PixelRAG tile corpus
Rendered screenshot tiles for PixelRAG, a visual retrieval-augmented-generation system that retrieves over page images instead of parsed text. Each Wikipedia page is rendered to an image and cut into fixed-height tiles; retrieval runs on the tiles directly with a Qwen3-VL embedding model.
This repository holds the full tile corpus that the published FAISS indexes and embeddings were built from, so the whole pipeline (tiles → embeddings → index → search) can… See the full description on the dataset page: https://huggingface.co/datasets/StarTrail-org/pixelrag-tiles.plism-dataset-tiles
PLISM dataset
The Pathology Images of Scanners and Mobilephones (PLISM) dataset was created by (Ochi et al., 2024) for the evaluation of AI models’ robustness to inter-institutional domain shifts.
All histopathological specimens used in creating the PLISM dataset were sourced from patients who were diagnosed and underwent surgery at the University of Tokyo Hospital between 1955 and 2018.
PLISM-wsi consists in a group of consecutive slides digitized under 7 different scanners and… See the full description on the dataset page: https://huggingface.co/datasets/owkin/plism-dataset-tiles.nanopath-fairness-tiles
nanopath-fairness-tiles
Pre-tiled histopathology patches from CPTAC whole-slide images, used as the
external out-of-distribution validation set for a study on pretraining-time
vs. post-hoc fairness in histopathology foundation models.
Contents
Per-cohort folders, each slides_full/<slide_id>.parquet (one row per tile:
case_id, slide_id, tile_idx, image) + labels.tsv:
cohort
organ / task
slides
cptac_lung
NSCLC — LUAD vs LSCC subtype
604
cptac_gbm
GBM —… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/nanopath-fairness-tiles.onlypits-lunar-tilespopcorn-reportsgtex-10M-balanced-tiles
GTEx 10M Balanced Tiles
This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed.
Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.game-asset-tiles
Game Asset Tiles
This dataset contains 275 samples, each consisting of a tile image, a matching template image and a description.
Tiles were generated with Nano Banana 2 and FLUX.2, or sourced from kenney.nl.
hex / oct / rect describe the footprint shape, flat-top or pointy-top describe the hex shape orientation, and isometric, oblique, or top-down describe the perspective.
Repository layout:
tiles/: PNG tile images
descriptions/: Short English text prompts
coordinates/: JSON point… See the full description on the dataset page: https://huggingface.co/datasets/crafiq/game-asset-tiles.trivia_qa
Dataset Card for "trivia_qa_passages"
More Information needed
tilingtopomlmsmarco-passage-tilde
Dataset Card for "msmarco-passage-tilde"
More Information needed
MM-Mind2Web-tilde_test_snapshot_20dist
MultiModal-Mind2Web~ (MM-Mind2Web~)
rabbit inc.
[Leaderboard & Blogpost to be released]
Configuration: test split, snapshot with seed 42, 20 distractors
Multimodal-Mind2Web is a dataset proposed by Boyuan et al.. It's designed for the development and evaluation of generalist web agents and includes various action trajectories of humans on real websites.
We've simplified the raw dump from both Multimodal-Mind2Web and Mind2Web into sequences of observation-action pairs. We've… See the full description on the dataset page: https://huggingface.co/datasets/rabbit-hmi/MM-Mind2Web-tilde_test_snapshot_20dist.nq_text
Dataset Card for "nq_text"
More Information needed
mahjong_tilesmmlu
Dataset Card for "mmlu"
More Information needed
TildeOpen-MT-Euro-mix
TildeOpen-MT-Euro-mix
TildeOpen-MT-Euro-mix is a machine-translation datamix for 30 European
languages. It was built for the translation instruction-tuning experiments
used to validate the
TildeOpen LLM
(Bergmanis et al., LREC 2026). We publish this
data to support the reproducibility of our experiments.
Composition
Datamix contains 1,890,652 sentence/document pairs across 245 translation directions and 30 languages - 221.44M source-side and 240.56M target-side… See the full description on the dataset page: https://huggingface.co/datasets/TildeAI/TildeOpen-MT-Euro-mix.deepfashion
DeepFashion Multimodal Dataset
This dataset contains annotations for shape, fabric, and texture attributes of clothing images.
Metadata
id_to_category:
{
"0": "Sleeve Length",
"1": "Lower Clothing Length",
"2": "Socks",
"3": "Hat",
"4": "Glasses",
"5": "Neckwear",
"6": "Wrist Wearing",
"7": "Ring",
"8": "Waist Accessories",
"9": "Neckline",
"10": "Outer Clothing a Cardigan?",
"11": "Upper Clothing Covering Navel?"
}
category_options:
{… See the full description on the dataset page: https://huggingface.co/datasets/tilak1114/deepfashion.tilebac-flag-dataset
Background
Overview
This dataset contains ultralow-dose cryoEM montage tile images of the bacteria Pantoea sp. YR343. Segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 have been fine-tuned to predict bacterial inner membranes and outer membranes. This bacterial membrane dataset is a benchmark dataset to challenge current AI workflows in rapid, seamless montage stitching and stitched segmentation in extremely noisy ultralow-dose cryoEM images. Bacterial flagella low-dose… See the full description on the dataset page: https://huggingface.co/datasets/LynnMass/tilebac-flag-dataset.osm-completeness-tilesmolmoact2_yam_tiled_rfmGlobal-Ocean-Science-Corpus
🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned)
A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography
Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes.
Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.TILO.RA_CODER_Dataset
TILO.RA CODER Dataset
Объединённый русско-английский датасет для обучения и поиска по коду.
Формат — пары question / code: вопрос на естественном языке → готовый код-ответ.
Датасет собран для локального ассистента TILO.RA CODER — офлайн-помощника по программированию
Скачать по ссылке
https://github.com/thetemirbolatov/TILO.RA_CODER_Dataset/releases/download/v1.0.0/tilora_knowledge_merged.jsonl
Состав
Источник
Язык
Записей
English coding… See the full description on the dataset page: https://huggingface.co/datasets/thetemirbolatov/TILO.RA_CODER_Dataset.219-tiles
Roles
Roles: canon repo — annot is the source label, kept machine-parseable as the gold for verification and reward parsing; there is no filled reasoning column and this repo is not itself a training view. Derived repos each state their own regime on their own card.
219-tiles
Tile-level hot-crack detection in light-microscope images of a weld surface — the trainable view of 219 (512x512 native-resolution tiles; binary masks). ⚠ Microscopy, NOT production-line… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/219-tiles.til_audionecklines-attr
Necklines Attribute Dataset
This dataset contains images of different neckline styles categorized by type.
Structure
category: Main category of the neckline
subcategory: Specific type of neckline
image_url: Source URL of the image
image: The actual image data (PNG or JPEG), resized if dimensions exceed 1000px
SpatialData_with_spatial_tilesfashion-pediavscode-bug-feature-triage
VS Code Bug vs Feature Request Triage
Dataset summary
1,993 prepared issue records from public microsoft/vscode issues, reduced to one binary task: classify the issue text as bug or feature-request. The splits are a frozen temporal holdout (80/10/10 by created_at within each class, seed 42) used by the GitHub Triage SLM Fine-Tuning Benchmark to compare fine-tuned small models against their base checkpoints on the same test set. Each record carries cleaned issue… See the full description on the dataset page: https://huggingface.co/datasets/Tilakoid/vscode-bug-feature-triage.til_nlp_test_dataset
Dataset Card for "til_nlp_test_dataset"
More Information needed
conversations_seed_storymarco_cqa
Dataset Card for "marco_cqa"
More Information needed
