datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STS-3D-Tooth
STS-3D-Tooth
The 3D Cone-Beam CT (CBCT) subset of the STS (Semi-supervised Teeth
Segmentation) multi-modal dental dataset, as released in
Wang et al., Scientific Data 12, 117 (2025)
and used in the MICCAI 2023/2024 STS Challenges.
The companion 2D panoramic X-ray subset is hosted at
Angelou0516/STS-2D-Tooth.
Dataset Summary
Field
Details
Modality
Cone-Beam CT (CBCT), NIfTI (.nii.gz)
Body Part
Teeth (32 permanent teeth, FDI numbering)
Volumes
371… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/STS-3D-Tooth.rendered-sts17
Dataset Summary
This dataset is rendered to images from STS-17. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load Arabic to Arabic dataset:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts17", name="ar-ar", split="test")
Load French to English dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts17.rendered-stsb
Dataset Summary
This dataset is rendered to images from STS-benchmark. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load English train Dataset:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-stsb", name="en", split="train")
Load Chinese dev Dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-stsb.STS-2D-Tooth
STS-2D-Tooth
The 2D panoramic dental X-ray subset of the STS (Semi-supervised Teeth
Segmentation) multi-modal dataset, as released in
Wang et al., Scientific Data 12, 117 (2025)
and used in the MICCAI 2023 STS Challenge.
Composition
4,000 panoramic X-ray images (PNG, 640x320, 3-channel grayscale-as-RGB) split
across two demographic subsets:
Subset
Total
Labeled
Unlabeled
A-PXI (adult)
3,500
850
2,650
C-PXI (child)
500
50
450
Total
4,000
900
3,100… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/STS-2D-Tooth.rendered-sts13
Dataset Summary
This dataset is rendered to images from STS-13. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts13", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts13.rendered-sts15
Dataset Summary
This dataset is rendered to images from STS-15. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts15", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts15.rendered-sts16
Dataset Summary
This dataset is rendered to images from STS-16. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts16", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts16.rendered-sts14
Dataset Summary
This dataset is rendered to images from STS-14. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts14", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts14.rendered-sts12
Dataset Summary
This dataset is rendered to images from STS-12. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts12", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts12.rendered-sts12
Dataset Summary
This dataset is rendered to images from STS-12. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts12", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts12.rendered-sts16
Dataset Summary
This dataset is rendered to images from STS-16. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts16", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts16.rendered-sts15
Dataset Summary
This dataset is rendered to images from STS-15. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts15", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts15.rendered-sts13
Dataset Summary
This dataset is rendered to images from STS-13. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts13", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts13.pixel_glue_stsb
Dataset Card for "pixel_glue_stsb"
More Information needed
MICCAI-STS26-Challenge-Pre-Taskveshti-controlnet-v4-canny
Dataset Card for "veshti-controlnet-v4-canny"
More Information needed
rendered-sts14
Dataset Summary
This dataset is rendered to images from STS-14. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts14", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts14.veshti-controlnet
Dataset Card for "veshti-controlnet"
More Information needed
pixel_glue_stsb_high_noise
Dataset Card for "pixel_glue_stsb_high_noise"
More Information needed
pixel_glue_stsb_low_noise
Dataset Card for "pixel_glue_stsb_low_noise"
More Information needed
veshti-controlnet-v2-sammed-fingers
Dataset Card for "veshti-controlnet-v2-sammed-fingers"
More Information needed
twitter-stst521-2026.04.04-2040243332639752440-gLjV55Qqg0nWWNVn-part1STS-Challenge-2025sts-data
