min
Datasets
All datasets matching “min”MINT-1T-HTML
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.witmindeyev2Using these files requires that you have already agreed to the Natural Scenes Dataset's Terms and Conditions: https://cvnlab.slite.page/p/IB6BSeW_7o/Terms-and-Conditions
Webdatasets only contain behavioral information, .tar filename numbering correspondings to the scanning session of the subject.
(Always use "new_test" instead of "test" in the wds folders, "test" refers to using the old NSD data from before they released the full set of scanning sessions.)
behavior numpy files correspond to… See the full description on the dataset page: https://huggingface.co/datasets/pscotti/mindeyev2.minigridSignLanguage_MiniProjectDataset used for training a model to classify Danish Sign Language signs, based on MediaPipe hand landmark data.
The data is not split into training, test and validation sets.
The dataset consist of four classes, 'unknown', 'hello', 'bye' and 'thanks'.
There are 30 datapoints for each class.
Each data point is 30 frames of data stored in individual Numpy files with x, y and z values for each hand landmark.
D4RL
