datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oxford-iiit-pet
The Oxford-IIIT Pet Dataset
Description
A 37 category pet dataset with roughly 200 images for each class. The images have a large variations in scale, pose and lighting.
This instance of the dataset uses standard label ordering and includes the standard train/test splits. Trimaps and bbox are not included, but there is an image_id field that can be used to reference those annotations from official metadata.
Website: https://www.robots.ox.ac.uk/~vgg/data/pets/… See the full description on the dataset page: https://huggingface.co/datasets/timm/oxford-iiit-pet.mini-imagenet
Dataset Description
A mini version of ImageNet-1k with 100 of 1000 classes present.
Unlike some 'mini' variants this one includes the original images at their original sizes. Many such subsets downsample to 84x84 or other smaller resolutions.
Data Splits
Train
50000 samples from ImageNet-1k train split
Validation
10000 samples from ImageNet-1k train split
Test
5000 samples from ImageNet-1k validation split (all 50 samples per class)… See the full description on the dataset page: https://huggingface.co/datasets/timm/mini-imagenet.imagenet-22k-wds
Dataset Summary
This is a copy of the full ImageNet dataset consisting of all of the original 21841 clases. It also contains labels in a separate field for the '12k' subset described at at (https://github.com/rwightman/imagenet-12k, https://huggingface.co/datasets/timm/imagenet-12k-wds)
This dataset is from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-22k-wds.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.resisc45
Description
RESISC45 dataset is a publicly available benchmark for Remote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This dataset contains 31,500 images, covering 45 scene classes with 700 images in each class.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/resisc45.eurosat-rgb
EuroSat (RGB)
Description
A dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples. This is the RGB version of the dataset with visible bands encoded as JPEG images.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/timm/eurosat-rgb.objectnet-in1k
ObjectNet (ImageNet-1k Overlapping)
A webp (lossless) encoded version of ObjectNet-1.0 at original resolution, containing only the images for the 113 classes that overlap with ImageNet-1k classes.
License / Usage Terms
ObjectNet is free to use for both research and commercial applications. The authors own the source images and allow their use under a license derived from Creative Commons Attribution 4.0 with only two additional clauses.
ObjectNet may never be used to… See the full description on the dataset page: https://huggingface.co/datasets/timm/objectnet-in1k.plant-pathology-2021
Description
Dataset from the Plant Pathology 2021 (FGVC8) Challenge.
'
For Plant Pathology 2021-FGVC8, we have significantly increased the number of foliar disease images and added additional disease categories. This year’s dataset contains approximately 23,000 high-quality RGB images of apple foliar diseases, including a large expert-annotated disease dataset. This dataset reflects real field scenarios by representing non-homogeneous backgrounds of leaf images taken at different… See the full description on the dataset page: https://huggingface.co/datasets/timm/plant-pathology-2021.imagenet-w21-webp-wds
Dataset Summary
This is a copy of the full Winter21 release of ImageNet in webdataset tar format with WEBP encoded images. This release consists of 19167 classes, 2674 fewer classes than the original 21841 class Fall11 release of the full ImageNet.
The classes were removed due to these concerns: https://www.image-net.org/update-sep-17-2019.php
This is the same contents as https://huggingface.co/datasets/timm/imagenet-w21-wds but encoded in webp at ~56% of the size, shard count… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-webp-wds.objectnet
ObjectNet
A webp (lossless) encoded version of ObjectNet-1.0 at original resolution.
License / Usage Terms
ObjectNet is free to use for both research and commercial applications. The authors own the source images and allow their use under a license derived from Creative Commons Attribution 4.0 with only two additional clauses.
ObjectNet may never be used to tune the parameters of any model.
Any individual images from ObjectNet may only be posted to the web including… See the full description on the dataset page: https://huggingface.co/datasets/timm/objectnet.imagenet-w21-p
Dataset Summary
This is a subset of the full Winter21, filtered according to https://github.com/Alibaba-MIIL/ImageNet21K. This instance contains 10450 classes with a train and validation split.
Processing
I performed some processing while sharding this dataset:
Synsets were filtered according to ImageNet-21-P scripts
Images were re-encoded in WEBP
Additional Information
Dataset Curators
Authors of [1] and [2]:
Olga Russakovsky
Jia Deng
Hao Su… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-p.objectnet-720p-in1k
ObjectNet (720P Shortest Edge, ImageNet-1k Overlap)
A webp (lossless) encoded version of ObjectNet-1.0 resized to shortest edge = 720 pixels. Containing only the 113 classes that overlap with ImageNet-1k.
License / Usage Terms
ObjectNet is free to use for both research and commercial applications. The authors own the source images and allow their use under a license derived from Creative Commons Attribution 4.0 with only two additional clauses.
ObjectNet may never be… See the full description on the dataset page: https://huggingface.co/datasets/timm/objectnet-720p-in1k.conus-h3-land-cover
CONUS H3 Land Cover Density (NLCD 2024)
A continental United States land cover dataset aggregated to Uber's H3 hexagonal grid at resolution 10 (~120m cell edge length). Each cell contains continuous per-channel coverage fractions derived from the USGS National Land Cover Database (NLCD) 2024. Unlike NLCD's native discrete classification, this dataset expresses land cover as continuous multi-channel density values — a cell at a forest-wetland boundary is represented as 0.6 forest and… See the full description on the dataset page: https://huggingface.co/datasets/Timmahw/conus-h3-land-cover.imagenet-w21-wds
Dataset Summary
This is a copy of the full Winter21 release of ImageNet in webdataset tar format with JPEG images. This release consists of 19167 classes, 2674 fewer classes than the original 21841 class Fall11 release of the full ImageNet.
The classes were removed due to these concerns: https://www.image-net.org/update-sep-17-2019.php
Data Splits
The full ImageNet dataset has no defined splits. This release follows that and leaves everything in the train split.… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-wds.timms2019-nt24027-emb
ESM-2 layer-wise embeddings — Nt_Timms2019
Timms et al. 2019 Science 365(6448):eaaw4912 의 스크린에서 나온 24,027개 항목에 대해 ESM-2 의 모든 transformer
레이어 per-residue hidden state 를 뽑은 데이터셋입니다. masking 없이 raw forward 1회.
데이터셋 태그: Nt_Timms2019_24027 · 저장 레이아웃: stacked
구성
폴더
모델
레이어
차원
레이어당
합계
ESM2_8M/
facebook/esm2_t6_8M_UR50D
8 (L0–L6, L6_pre)
320
0.40 GB
3.2 GB
ESM2_650M/
facebook/esm2_t33_650M_UR50D
35 (L0–L33, L33_pre)
1280
1.60 GB
56.0 GB
레이어 1개 = 파일 1개입니다.… See the full description on the dataset page: https://huggingface.co/datasets/Limtat99/timms2019-nt24027-emb.animal-clef-2026
AnimalCLEF26 Kaggle Competition Dataset
This is a HuggingFace mirror of the official AnimalCLEF26 competition dataset. All files existing in the original Kaggle dataset are unchanged, this mirror just adds additional Parquet metadata that make the dataset easier to use with HuggingFace.
Loading
from datasets import load_dataset
dataset = load_dataset("BVRA/animal-clef-2026")
print(dataset["train"][0]["image"])
Documentation
For documentation, see… See the full description on the dataset page: https://huggingface.co/datasets/timmhaucke/animal-clef-2026.imagenet-12k-wds
Dataset Summary
This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm.
The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k
NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-12k-wds.timm_inference_benchmarkobjectnet-720p
ObjectNet (720P Shortest Edge)
A webp (lossless) encoded version of ObjectNet-1.0 resized to shortest edge = 720 pixels.
License / Usage Terms
ObjectNet is free to use for both research and commercial applications. The authors own the source images and allow their use under a license derived from Creative Commons Attribution 4.0 with only two additional clauses.
ObjectNet may never be used to tune the parameters of any model.
Any individual images from ObjectNet may only… See the full description on the dataset page: https://huggingface.co/datasets/timm/objectnet-720p.hexo-bootstrap-corpus
Hexo Human Corpus
Encoding-free corpus of decisive human Hex Tac Toe games — hexagonal grid,
six-in-a-row to win (player 1 opens with 1 move, then both players play 2 moves
per turn; the board is theoretically infinite).
Each line is one game as a raw axial move list + outcome. Nothing about any
neural-network encoding is baked in — no planes, no fixed board size, no action
space. Read it with the stdlib json module and build whatever representation
you want.
Files… See the full description on the dataset page: https://huggingface.co/datasets/timmyburn/hexo-bootstrap-corpus.so101_keyboard_gray_cube_20260801_224135This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Timmy/so101_keyboard_gray_cube_20260801_224135.hpo-cultural-pt-curated
hpo-cultural-pt-curated
Dataset curado de sinônimos coloquiais brasileiros para fenótipos HPO. Primeira vez que é disponibilizado aberto.
Conteúdo
cultural_pairs.json — 705 pares (anchor PT coloquial → HPO canonical EN) com register, region, hpo_id
hard_negatives.json — 72 triplets (anchor, positive, negative) para desambiguação
Exemplos
"água na cabeça" → Hydrocephalus (HP:0000238)
"esparro" (BR-NE) → Seizure (HP:0001250)
"cabeção" → Macrocephaly… See the full description on the dataset page: https://huggingface.co/datasets/timmers/hpo-cultural-pt-curated.timm-mobilevit-s-cvnet-in1k-val-top_5so101_keyboard_gray_cube_clean11Capstone-Reposetimm-resnset18-a1-in1k-val-top_5-chunkedExploitVectors
ExploitVectors
tags: Injection, Exploit Methods, Classification
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'ExploitVectors' dataset is a curated collection of textual descriptions detailing various types of cybersecurity exploits. Each entry in the dataset provides a brief explanation of the exploit method, followed by an indication of its potential impact level and classification. This dataset can be useful for… See the full description on the dataset page: https://huggingface.co/datasets/timmaythetoolmann/ExploitVectors.consciousness-assessment-benchmark-v1code-refusal-for-abliteration
code-refusal-for-abliteration
Takes datasets of responses / refusals used for abliteration,
and filters these down to programming-specific tasks for code models to be abliterated.
Sources:
https://github.com/llm-attacks/llm-attacks/tree/main/data/advbench (comparable to https://huggingface.co/datasets/mlabonne/harmful_behaviors )
Also see: https://github.com/AI-secure/RedCode/tree/main/dataset / https://huggingface.co/datasets/monsoon-nlp/redcode-hf for samples using Python… See the full description on the dataset page: https://huggingface.co/datasets/timmaythetoolmann/code-refusal-for-abliteration.timmy-t2-timer-sft
Timmy T2 Timer SFT Dataset
This dataset trains Timmy T2, the Timmy Timer Translator model, to convert natural-language timer requests into Timey's compact action DSL.
Model repo: Satansdeer/timmy-t2
Splits
Split
File
Rows
train
data/train.jsonl
2639
validation
data/validation.jsonl
207
hard_validation
data/hard_validation.jsonl
62
all_public
data/all_public.jsonl
2846
The hidden validation split (16 rows) is intentionally not included in this public… See the full description on the dataset page: https://huggingface.co/datasets/Satansdeer/timmy-t2-timer-sft.
