CoolFace
Modelpublic

DigitalByte/LTX-2.5-Exploded-View-XPLDV

sourceHugging Faceotherupdated 3d agoView on Hugging Face
4likes
Model Card

XPLDV — Exploded-View LoRA for LTX-2.5

Turn a source image into an animated exploded view. XPLDV separates an object's exterior parts to reveal the components underneath, for product showcases, mechanical reveals, architectural breakdowns, and stylized cutaways.

Download LoRA · ComfyUI workflow · Sample image · Workflow guide · Prompt guide

How it works

  1. 1.Load the included workflow and select your installed model files.
  2. 2.Download the sample PC image and upload it in Starting image, or copy it to ComfyUI/input/xpldv_computer.png. The workflow is prefilled with its exact tested prompt.
  3. 3.Generate the video with XPLDV at strength 0.5.

For your own subject, replace the image and fill the BASIC template in the note directly above the positive prompt, or use the COMPLEX template in the reference note. Describe which parts separate, what they reveal, and how the camera moves.

The recommended starting model, based on our tests, is ltx-2.5-22b-distilled-transformer-bf16.safetensors. It can also be used with ltx-2.5-22b-dev-transformer-bf16.safetensors, which it was trained on. The included workflow uses the distilled model. Base models: Lightricks/LTX-2.5.

Featured videos

Six 10-second single-image examples at strength 0.5, plus a six-second shoe example at 0.7 using first and last frame guidance. Each video includes its exact recorded prompt.

<Gallery />

<details> <summary>With and without XPLDV</summary>

Matched comparisons using the same image, prompt, seed, and settings. Comparison videos are silent.

Laptop · Computer · SUV · Watch · House · Fighter jet · Running shoe

The shoe comparison keeps both image guides fixed and uses the recorded strength of 0.7 on the XPLDV side. Compare the movement and component behavior between the guided endpoints. Shoe workflow and guidance.

</details>

Recommended settings

Start with LoRA strength 0.5.

SettingRecommended value
CFG1 in both stages
Duration10 seconds: 241 frames at 24 FPS
First stage960 × 544; Euler ancestral; 8 steps
Second stage1920 × 1088; Euler; 3 steps
Second-stage manual sigmas0.85, 0.7250, 0.4219, 0.0
Image conditioning0.7 in the first stage; 1.0 in the second
Prompt enhancementCan be enabled; off in the recorded examples
Negative promptEmpty
Seed for comparisons42 in both stages

The workflow includes these settings and applies the LoRA to both stages. The manual sigmas above belong in Upscale and re-sampler (3 steps); keep the first-stage schedule unchanged.

Prompt enhancement: Enabling Enhance prompt works with XPLDV, but results may vary. It has not been extensively tested. The recorded examples used enhancement off to preserve the supplied prompts.

<details> <summary>LoRA strength</summary>

These are observed tendencies from the tests, not fixed rules for every subject.

LoRA strengthObserved behavior
0.5Recommended balance of separation, shape, and detail.
0.75Stronger influence; can improve separation, but may simplify detail or repeat parts.

Change one setting at a time while keeping the image, prompt, and seed fixed. </details>

Prompting

Start with "separates into an ordered exploded view." This phrase appears in all 204 separation training captions. Name the parts that move, the components revealed underneath, and the final arrangement. Write one continuous paragraph.

Use the slot reference to fill the placeholders in either template.

“Separates into individual components” is another phrasing we tested. With either opening, describe the movement and the final arrangement.

BASIC template — a short, readable reveal

text
[SUBJECT] separates into an ordered exploded view. [NAMED OUTER PARTS] move outward along controlled paths, revealing [NAMED INTERNAL COMPONENTS] underneath. The parts settle into an organized [LAYERED / RADIAL] arrangement around [EXPOSED CORE] while retaining their original shapes. [CAMERA BEHAVIOR]. [LIGHTING CONSISTENT WITH THE IMAGE] remains consistent throughout the continuous shot.
Example: Desktop PC
text
The desktop computer tower separates into an ordered exploded view. The glass side panel, front panel, top cover and lower power-supply cover move outward along controlled paths, revealing the motherboard, graphics card, liquid-cooling assembly and power supply. The components separate into an organized layered arrangement around the case while retaining their original shapes. The case base stays on the desk. The camera slowly orbits to the right around the computer from the starting three-quarter view. The blue-white computer lighting and surrounding room remain consistent throughout the continuous shot.

Watch the example · Download its starting image

COMPLEX template — a detailed, cinematic reveal

A COMPLEX prompt guides a more detailed, cinematic reveal through successive layers, describing deeper internal components, their appearance, and their positions, including details hidden or difficult to see in the starting image. These descriptions give LTX 2.5 the information needed to construct the intended reveal beyond what the image alone shows.

text
[SUBJECT] separates into an ordered exploded view while [CONTINUING ACTION OR ANCHORED SUPPORT]. [OUTER PART A] detaches and translates [DIRECTION], keeping [ORIENTATION]. [OUTER PART B] moves [DISTINCT PATH], and [OTHER OUTER PARTS] move [PATHS]. These exterior parts move first and travel furthest, leaving clear gaps. They reveal [SPECIFIC COMPONENT AND RECOGNIZABLE FEATURES] at [POSITION RELATIVE TO THE OBJECT], with [OTHER COMPONENTS AND THEIR LOCATIONS]. [CORE COMPONENTS] remain mounted together. Each detached section retains its shape as the arrangement settles and holds steady. [CAMERA BEHAVIOR]. [CONSISTENT LIGHTING] continues through one continuous shot.
Example: Watch
text
A luxury automatic mechanical wristwatch separates into an ordered exploded view while resting stationary on its dark velvet display cushion. The front sapphire crystal and surrounding polished bezel rim detach as a single unit and translate forward and slightly upward, keeping its circular orientation facing the camera. The hour, minute, and slender seconds hands lift cleanly above the dial plane, each translating upward along its own short vertical path with slight lateral spacing between them. The main dial plate carrying its applied hour markers, subdial components, and open-heart cutout details moves forward a shorter distance than the crystal, creating a clear gap between it and the hands above. These exterior parts move first and travel furthest, leaving generous gaps between each successive layer. They reveal the mechanical movement at the center of the stack—its polished bridges, interlocking gears, escapement fork, oscillating balance wheel, ruby jewels, and complication modules all visible in their assembled configuration—with the rear exhibition caseback and its oscillating automatic rotor translating backward and slightly downward as the final layer. The movement's bridges, gears, and escapement remain mounted together as a single mechanical unit. Each detached section retains its original shape as the five layers settle into an organized axial arrangement and hold steady. A slow, steady camera pull-back provides room for the full expanded stack. The soft upper-left key light, cool right-side fill, and subtle rim highlight consistent with the source image continue through one continuous shot.

Watch the example

See the prompt guide for motion, preservation, and optional audio, plus troubleshooting.

Limitations and tips

  • —Hidden construction is imagined by the model. Detailed-looking internals may be inaccurate or misplaced.
  • —Parts can merge, duplicate, bend, or lose their identity. Start with a few major assemblies and give them distinct movement paths.
  • —Background objects, text, and camera framing can change. Leave room around the subject and check the full clip.
  • —Review new images and seeds individually. The settings above are a starting point, and stronger LoRA influence is not always an improvement.

Training

SettingValue
TrainerLightricks LTX trainer
Base transformerLTX-2.5 22B DEV BF16
HardwareOne NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB VRAM
LoRA rank / alpha64 / 64
Training steps4,000
Released weightsStep 3000

Dataset

ItemDetails
CreationSynthetic clips created in-house with Blender EEVEE Next
Objects102 multipart objects
Clips256 total: 204 separation clips and 52 reversed reassembly clips
Resolution960 × 544
Length and frame rate121 frames at 24 FPS
ArrangementsRadial, horizontal, and vertical
AudioNo audio in the training clips; audio was not trained

The photographic example images were used for inference, not training. Image-to-video is the evaluated use. Training includes reassembly, but a complete separate-hold-reassemble sequence within ten seconds is not established.

Training configuration · Dataset summary

License

Distributed under the included LTX-2.x Community License Agreement. See Lightricks/LTX-2.5 for the base model and its terms.