datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embedded_movies
sample_mflix.embedded_movies
This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast.
In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature.
Overview
This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.fawkes-training-graph-embedded-260615
Fawkes — WM@Booth Training Graphs (v16 dataset) — PRIVATE
PRIVATE — derived from MIMIC-IV (PhysioNet credentialed, governed by the PhysioNet DUA). Do not redistribute. Credentialed access only.
The exact dataset the WM@Booth Graph-JEPA v16 model (on1onmangoes/fawkes-wmatbooth-graph-jepa-v16-260615) was trained on — 4,000 per-admission clinical knowledge graphs (~3,018 patients).
Each record = one hospital admission
field
what it is
subject_id, hadm_id… See the full description on the dataset page: https://huggingface.co/datasets/wmatbooth/fawkes-training-graph-embedded-260615.
