datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liquidrandom-data
liquidrandom-data
Diverse seed data for ML/LLM training data generation pipelines.
Used by the liquidrandom Python package.
Dataset Summary
This dataset contains 520,080 seed data samples across 24 categories,
generated using a hierarchical taxonomy tree approach with LLM-based quality validation
and fuzzy deduplication. Data is stored as Parquet with zstd compression.
Categories
Category
Samples
File
Coding Tasks
30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.GeoDEOfficial Paper
Number of country classes: 40Total number of images: 61925
Image count per country_ip class
Country
Number of Images
Angola
10
Argentina
3193
Botswana
3
Brazil
16
Bulgaria
1
Cameroon
1
China
1565
Colombia
3703
Egypt
2449
France
59
Ghana
1
Greece
45
Indonesia
5311
Ireland
2
Italy
3933
Japan
6500
Jordan
43
Malaysia
55
Mexico
2723
Moldova
2
Netherlands
18
Nigeria
5729
Philippines
2906
Poland
68
Portugal
139… See the full description on the dataset page: https://huggingface.co/datasets/MLap/GeoDE.NFT-70M_transactions
Dataset Card for "NFT-70M_transactions"
Dataset summary
The NFT-70M_transactions dataset is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea, the leading trading platform in the Web3 ecosystem.
With more than 70M transactions enriched with metadata, this dataset is conceived to support a wide range of tasks, ranging from sequential and transactional data processing/analysis to graph-based… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_transactions.NFT-70M_text
Dataset Card for "NFT-70M_text"
Dataset summary
The NFT-70M_text dataset is a companion for our released NFT-70M_transactions dataset,
which is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea.
As we also reported in the "Data anonymization" section of the dataset card of NFT-70M_transactions,
the textual contents associated with the NFT data were replaced by identifiers to numerical… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_text.NFT-70M_image
Dataset Card for "NFT-70M_image"
Dataset summary
The NFT-70M_image dataset is a companion for our released NFT-70M_transactions dataset,
which is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea.
As we also reported in the "Data anonymization" section of the dataset card of NFT-70M_transactions,
the URLs of NFT images data were replaced by identifiers to numerical vectors that represent an… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_image.citrus-disease-vlm-instruct
Citrus Disease VLM Instruct
An instruction-tuning dataset for training a small vision-language model (VLM) to look at a photo of a citrus leaf, fruit or shoot, name the disease, pest or nutrient deficiency, explain the cause and symptoms, and recommend both biological/organic and chemical management.
Every example pairs one image with a chat conversation in the format used by TRL's SFTTrainer for multimodal models (Qwen-VL, SmolVLM, Idefics, LLaVA and similar).
What… See the full description on the dataset page: https://huggingface.co/datasets/ML-Intern-lab/citrus-disease-vlm-instruct.pompax-classification
Pompax Equipment Classification Dataset
Dataset Description
This dataset contains 2,772 images for industrial equipment nameplate classification. The task is to classify whether an image contains an industrial nameplate ("tabliczka-znamionowa" in Polish) or not.
Dataset Summary
Total Images: 2,772
Task: Binary classification (nameplate vs non-nameplate)
Classes: 3 categories
tabliczka-znamionowa (nameplate): 2,525 images (91.1%)
inne (other/non-nameplate):… See the full description on the dataset page: https://huggingface.co/datasets/kahua-ml/pompax-classification.imagenet-1k-224-less
ImageNet-1k-224 LESS
LESS (Lightweight Embedding-driven Subset Selection) coreset of
mlnomad/imagenet-1k-224.
How it was built
CLIP (google/siglip-base-patch16-224) embeddings were generated for every image in the train split.
For each of the 1,000 ImageNet classes, k-means clustering (k=10) was applied
to the class embeddings.
The image closest to each centroid was selected, yielding at most k images per
class — the most geometrically representative samples.… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/imagenet-1k-224-less.food101-MLOPS
Dataset Card for Food-101
Dataset Summary
This dataset consists of 101 food categories, with 101'000 images. For each class, 250 manually reviewed test images are provided as well as 750 training images. On purpose, the training images were not cleaned, and thus still contain some amount of noise. This comes mostly in the form of intense colors and sometimes wrong labels. All images were rescaled to have a maximum side length of 512 pixels.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/LTM19/food101-MLOPS.sun397
SUN397 dataset
The database contains 397 categories subset from the SUN dataset for Scene Recognition used in the following paper.
The number of images varies across categories, but there are at least 100 images per category, and 108,754 images in total.
All images are in jpg format. The images provided here are for research purposes only.
The file ClassName.txt contains the name list for the 397 categories.
Please cite the following paper if you use this dataset in your research.… See the full description on the dataset page: https://huggingface.co/datasets/pc-ml-dl/sun397.
