datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embedded_movies
sample_mflix.embedded_movies
This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast.
In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature.
Overview
This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.LLaVA-CC3M-Pretrain-595K-Embedded
Dataset derived from liuhaotian/LLaVA-CC3M-Pretrain-595K
Dataset details
Dataset type:
LLaVA Visual Instruct CC3M Pretrain 595K is a subset of CC-3M dataset, filtered with a more balanced concept coverage distribution.
Captions are also associated with BLIP synthetic caption for reference.
It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
We aim to build large multimodal towards GPT-4 vision/language capability.
wildfire-korea-embedded-300m
Wildfire Korea Embedded 300m (16-Channel Environmental Tiles)
A geospatial dataset containing 300 m embedded environmental feature tiles for the Korean Peninsula, designed for wildfire spread prediction and reinforcement-learning research.
This dataset provides model-ready 16-channel tensors representing static and quasi-static environmental conditions needed for wildfire simulation and RL inference.
It is paired with the wildfire episode dataset (wildfire-korea-episodes-300m) and… See the full description on the dataset page: https://huggingface.co/datasets/chaseungjoon/wildfire-korea-embedded-300m.embedded_datasetembedded_movies_smallia-loaded-embedded-gpu
Dataset Card for "ia-loaded-embedded-gpu"
More Information needed
ia_embedded
Dataset Card for "ia_embedded"
More Information needed
clevr1000_hf_image_text_embeddedembedded_movies_smallThis dataset was created from the HuggingFace dataset AIatMongoDB/embedded_movies
Why was it needed?
The original dataset is close to 25 GB, for learning and experiments it is an overkill
Data in the dataset needs to be cleaned up e.g., some features are Null that requires extra care
Some of the embeddings are missing
How to use?
Use for sentiment analysis
Text similarity (plot)
Embeddings : ready to use with vector DB & search libraries
dataset_info:
features:
- name:… See the full description on the dataset page: https://huggingface.co/datasets/acloudfan/embedded_movies_small.gte_embedded_moviesThis dataset originates from MongoDB's embedded_movies dataset and contains details on movies
from different genres. Each row represents a single movie with detailed information.
As opposed to the original dataset, this one includes embeddings of the fullplot column using the open source General Text Embeddings model
instead of OpenAI's text-embedding-ada-002 embedding model used in MongoDB Atlas.
Those open source embeddings are also used in Hermes.
images_with_embedded_images_v6embedded-pokemonnih-chest-xray-embeddedimages_with_embedded_images_v4images_with_embedded_images_v5images_with_embedded_images_v3
