badincite/minimax-h3-soup
MiniMax H3 Soup Reproducibility archive for a local ComfyUI MiniMax H3 Ref2V benchmark on an RTX 3090. What is included Original benchmark workflow graph (source_prompt.json), manifest, and result table. Every one-second MP4 from the original C1-C11 benchmark grid and its Euler repeat sweep. The separate duration experiments are intentionally not included. Labeled C1-C11 visual contact sheets, Ref2VA stock-control sheets, and a static render-time summary chart.… See the full description on the dataset page: https://huggingface.co/datasets/badincite/minimax-h3-soup.
MiniMax H3 Soup
Reproducibility archive for a local ComfyUI MiniMax H3 Ref2V benchmark on an RTX 3090.
What is included
- Original benchmark workflow graph (
source_prompt.json), manifest, and result table. - Every one-second MP4 from the original C1-C11 benchmark grid and its Euler repeat sweep. The separate duration experiments are intentionally not included.
- Labeled C1-C11 visual contact sheets, Ref2VA stock-control sheets, and a static render-time summary chart.
System
- NVIDIA RTX 3090
- 32 GB DDR4 RAM
- CUDA 13
- ComfyUI 0.34.4
- PyTorch 2.9.1+cu130
Layout
Dataset Viewer metadata
The primary-video folder includes a metadata.csv sidecar, so the Dataset Viewer shows a readable case (C1-C11 or Stock), display name, model/LoRA, megapixels, native resolution, steps, ksampler, seed, and measured render time beside every video. The original MP4 filenames are preserved for provenance; use the case and display name columns to browse the experiments.
10-second 0.8 MP run
0p8_10s is a separate duration experiment. Its initial set re-renders one selected 1-second 0.8 MP setting per C1-C11 at 10 seconds, using one fixed two-person Negan/Joe prompt and fixed seed. Jeffrey and Joe RefMods are active; the prior Matthew RefMod is disabled. C7 uses the corrected C7_r1 custom Turbo loader/sampler (simple, 8 steps), and C11 uses Euler at 8 steps. The viewer metadata records sampler, scheduler, steps, and measured total and per-output-second render time. This is the start of the 10-second evaluation, not a final quality ranking or a replacement for the one-second grid.
C3 Euler + simple ablation
This follow-up keeps the original C3 one-second benchmark grid (four native resolutions and 6/8/10/12 steps) but changes the sampler/scheduler from res_multistep + beta to euler + simple. It is included in the unified euler_1s Dataset Viewer config, where scheduler distinguishes it from the original Euler/beta repeat. Two 0.8 MP, 10-second C3 controls (8 and 10 steps) use the same two-person duration prompt as the 0p8_10s set.
Important numbering note: the original on-disk filenames use internal IDs 09_... and 10_... for the user-facing C8 and C9 tests. The metadata column maps them to the intended labels, and C10 is explicitly labelled C10.
Author-recommended one-second correction set
The original grid held res_multistep + beta and 6/8/10/12 steps constant across all combinations. That was useful for a controlled initial screen, but it is not the intended inference contract for several checkpoints. This separate 20-video set preserves the same source graph, prompt, references, fixed seed, and four native resolutions while applying each author’s published recipe exactly:
Use the author_recommended_1s Viewer configuration to browse it. The metadata records the complete recipe and a source note for each row. This set is a correction/validation run, not a replacement for the original fixed-grid comparison.
C1-C11 model / LoRA legend
All primary-grid tests use a one-second duration, fixed seed, and the res_multistep KSampler. The separate Euler grid repeats C1-C11 at the same resolution/step pairs with euler; ksampler in each viewer row makes the two series directly filterable. Only resolution and step count vary within a given case. The Stock control is retained as a separate higher-step reference.
C7 correction: Larry v4 Turbo loader
The archived C7 one-second results were created before the original Larry v4 Turbo LoRA loader requirement was caught. Those legacy rows use the standard ComfyUI LoRA loader and therefore are not a valid quality ranking for the original v4 LoRA. Do not use them to judge C7 against C1-C6/C8-C11.
Use H3_REF2V_RTX3090_C7_ORIGINAL_TURBO_CUSTOM_LOADER.json instead. It uses ComfyUI-MiniMax-H3-Turbo's MiniMax-H3 Turbo LoRA node with low_vram=false (runtime bypass for maximum sharpness), MiniMax-H3 Turbo Sampler, simple scheduler, and 8 steps. A corrected C7 benchmark will be recorded as a separate experiment rather than overwriting the legacy data.
The completed corrected screen is named C7_r1 and is available in the c7_r1_original_turbo_1s Dataset Viewer configuration. The paired comparison below includes only 6 and 8 steps because those are the two step values both the legacy C7 run and C7_r1 share.
Repeatable technical-quality metrics
quality_metrics.csv and the Dataset Viewer columns include a transparent, deterministic technical score for every primary-grid video. The method is h3-transparent-v1: eight evenly-spaced decoded frames, fixed OpenCV algorithms, and fixed (absolute, not rank-based) score mappings.
The exact scoring implementation is included as tools/score_h3_primary_videos.py; it was run with OpenCV 5.0.0 and NumPy 2.4.4. Re-running it against the unchanged MP4s produces the same values.
These metrics do not score identity accuracy, prompt adherence, audio, motion direction, or artistic preference. They are a consistent technical filter for community review—not a replacement for watching the videos.
Visual comparison sheets
Each tile is a mid-frame from the corresponding one-second render. Columns are C1 through C11; rows are 6, 8, 10, and 12 steps. Render time is shown beneath each tile.
0.4 MP
0.6 MP
0.8 MP
0.98 MP
Render-time summary
This chart measures render speed only; it is not a visual-quality ranking.
Euler versus res_multistep
These four paired sheets compare the same C1-C11 model/LoRA, seed, resolution, and step count. Each pair is labelled RM (res_multistep) on the left and EU (euler) on the right; the time below each tile is its measured render time.
0.4 MP
0.6 MP
0.8 MP
0.98 MP
C3 versus C10 FastH3
Paired comparison at matching resolution and step counts. C3 uses the FL2VA Turbo 8-step model without a LoRA; C10 uses Ref2VA with the converted FastH3 LoRA. Prompt, seed, and duration are held constant.
results.jsonl is the append-only execution record. results.csv is the current tabular export. During the original Windows file lock, the live export was used to create the published results.csv.
Reproduction notes
The workflow JSON references local ComfyUI model paths. It does not include model weights, LoRAs, RefMods, input images, or audio assets. Install the matching MiniMax H3 nodes and models first, then update local paths as needed.
This repository is an experimental benchmark, not a claim of general model quality. Visual preference and audio quality were assessed separately; identical seed and prompt settings were used within each controlled sweep.
