CoolFace
Datasetpublic

TempoFunk/small

TempoFunk Small 7.8k samples of metadata and encoded latents & prompts of random videos. Data format Video frame latents Numpy arrays 120 frames, 512x512 source size Encoded shape (120, 4, 64, 64) CLIP (openai) encoded prompts Video description (as seen in metadata) Encoded shape (77,768) Video metadata as JSON (description, tags, categories, source URL, etc.)

sourceHugging Faceagpl-3.0updated 3y agoView on Hugging Face
9likes65kdownloads
Dataset Card

TempoFunk Small

7.8k samples of metadata and encoded latents & prompts of random videos.

Data format

  • Video frame latents
  • Numpy arrays
  • 120 frames, 512x512 source size
  • Encoded shape (120, 4, 64, 64)
  • CLIP (openai) encoded prompts
  • Video description (as seen in metadata)
  • Encoded shape (77,768)
  • Video metadata as JSON (description, tags, categories, source URL, etc.)