datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet-12k-wds
Dataset Summary
This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm.
The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k
NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-12k-wds.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-1k-wds.xet-spec-reference-filesThe files in this repository are intended to provide a reference to content processed using the xet protocol, relative to the original file: Electric_Vehicle_Population_Data_20250917.csv
The original file was exported from https://data.wa.gov/Transportation/Electric-Vehicle-Population-Data/f6w7-q2d2/about_data on September 16, 2025.
The contents are as described:
Electric_Vehicle_Population_Data_20250917.csv - the original file
Electric_Vehicle_Population_Data_20250917.csv.chunks - a… See the full description on the dataset page: https://huggingface.co/datasets/xet-team/xet-spec-reference-files.ready-xet-gotest-public-datasetDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/test-public-dataset.inference-benchmarker-xettest-xet-migration-2xet-scope-probe-0909xet-spec-reference-filesThe files in this repository are intended to provide a reference to content processed using the xet protocol, relative to the original file: Electric_Vehicle_Population_Data_20250917.csv
The original file was exported from https://data.wa.gov/Transportation/Electric-Vehicle-Population-Data/f6w7-q2d2/about_data on September 16, 2025.
The contents are as described:
Electric_Vehicle_Population_Data_20250917.csv - the original file
Electric_Vehicle_Population_Data_20250917.csv.chunks - a… See the full description on the dataset page: https://huggingface.co/datasets/MarcoAlho/xet-spec-reference-files.bdh-cl_xet-probeoutput-XETG00082_C105XetTestxethub-demo-videosxet_testxet-wasm-testtest-xet-sagemakerteste-xetl574-b-xetx
l574-x-xetbroadx
l574-x-xetaltx
lfs_ds_publicl574-x-xetkeyx
twitter-xingyinriqi-2025.09.29-1972483971511627983-K9m49_xETC9159-6-part1my-xet-datasetXeTz7dYlXETAYrl8
