CoolFace
Datasetpublic

badincite/minimax-h3-soup

MiniMax H3 Soup Reproducibility archive for a local ComfyUI MiniMax H3 Ref2V benchmark on an RTX 3090. What is included Original benchmark workflow graph (source_prompt.json), manifest, and result table. Every one-second MP4 from the original C1-C11 benchmark grid and its Euler repeat sweep. The separate duration experiments are intentionally not included. Labeled C1-C11 visual contact sheets, Ref2VA stock-control sheets, and a static render-time summary chart.… See the full description on the dataset page: https://huggingface.co/datasets/badincite/minimax-h3-soup.

sourceHugging Faceupdated 13d agoView on Hugging Face
5likes2.2kdownloads
Dataset Card

MiniMax H3 Soup

Reproducibility archive for a local ComfyUI MiniMax H3 Ref2V benchmark on an RTX 3090.

What is included

  • Original benchmark workflow graph (source_prompt.json), manifest, and result table.
  • Every one-second MP4 from the original C1-C11 benchmark grid and its Euler repeat sweep. The separate duration experiments are intentionally not included.
  • Labeled C1-C11 visual contact sheets, Ref2VA stock-control sheets, and a static render-time summary chart.

System

  • NVIDIA RTX 3090
  • 32 GB DDR4 RAM
  • CUDA 13
  • ComfyUI 0.34.4
  • PyTorch 2.9.1+cu130

Layout

PathContents
benchmark/primary_res_multistep_1s/Main C1-C11 model/LoRA screen, using res_multistep at one second.
benchmark/euler_ablation_1s/Original full C1-C11 Euler/beta repeat grid at one second.
benchmark/c7_r1_original_turbo_1s/Corrected C7 original Larry v4 experiment using its custom loader and Turbo Sampler.
videos/primary_res_multistep_1s/Every original-grid one-second MP4, C1 through C11.
videos/euler_1s/All one-second Euler tests. Filter the scheduler column to compare beta and simple.
videos/c7_r1_original_turbo_1s/Corrected C7_r1 original-Turbo videos, 4/6/8 steps.
videos/author_recommended_1s/C1, C2, and C4-C6 rerun at each author's documented sampler, scheduler, shifts, and step count.
videos/0p8_10s/10-second 0.8 MP renders, with sidecar metadata. Additional runs can be added here.
reports/primary_res_multistep_1s/Contact sheets and the render-time chart.
workflows/Cleaned reusable ComfyUI workflow JSON.
workflows/H3_REF2V_0p8_10s_C1_C11_BENCHMARK/Exact submitted API graphs, runner, manifest, prompt, and result records for the 10-second run.

Dataset Viewer metadata

The primary-video folder includes a metadata.csv sidecar, so the Dataset Viewer shows a readable case (C1-C11 or Stock), display name, model/LoRA, megapixels, native resolution, steps, ksampler, seed, and measured render time beside every video. The original MP4 filenames are preserved for provenance; use the case and display name columns to browse the experiments.

10-second 0.8 MP run

0p8_10s is a separate duration experiment. Its initial set re-renders one selected 1-second 0.8 MP setting per C1-C11 at 10 seconds, using one fixed two-person Negan/Joe prompt and fixed seed. Jeffrey and Joe RefMods are active; the prior Matthew RefMod is disabled. C7 uses the corrected C7_r1 custom Turbo loader/sampler (simple, 8 steps), and C11 uses Euler at 8 steps. The viewer metadata records sampler, scheduler, steps, and measured total and per-output-second render time. This is the start of the 10-second evaluation, not a final quality ranking or a replacement for the one-second grid.

[image]

C3 Euler + simple ablation

This follow-up keeps the original C3 one-second benchmark grid (four native resolutions and 6/8/10/12 steps) but changes the sampler/scheduler from res_multistep + beta to euler + simple. It is included in the unified euler_1s Dataset Viewer config, where scheduler distinguishes it from the original Euler/beta repeat. Two 0.8 MP, 10-second C3 controls (8 and 10 steps) use the same two-person duration prompt as the 0p8_10s set.

Important numbering note: the original on-disk filenames use internal IDs 09_... and 10_... for the user-facing C8 and C9 tests. The metadata column maps them to the intended labels, and C10 is explicitly labelled C10.

Author-recommended one-second correction set

The original grid held res_multistep + beta and 6/8/10/12 steps constant across all combinations. That was useful for a controlled initial screen, but it is not the intended inference contract for several checkpoints. This separate 20-video set preserves the same source graph, prompt, references, fixed seed, and four native resolutions while applying each author’s published recipe exactly:

CaseSamplerSchedulerStepsVideo/audio shift
C1res_multistepsimple412 / 3
C2 (baked v1.2)eulersimple46 / 3
C4 (FL2V v1.1)eulersimple46 / 3
C5 (FL2V v1.2)eulersimple46 / 3
C6 (Ref2VA v0.1)eulersimple412 / 3

Use the author_recommended_1s Viewer configuration to browse it. The metadata records the complete recipe and a source note for each row. This set is a correction/validation run, not a replacement for the original fixed-grid comparison.

C1-C11 model / LoRA legend

LabelMeaning
StockRef2VA base with no LoRA, tested separately at 20, 25, and 30 steps
C1minimax_h3_fused_refdelta_r1024_turbo8_mystic07_int8_convrot with no LoRA
C2Minimax_H3_fl2va_Pruned_Lightx2v_turbo_4step_v1.2_768p_INT8 with no LoRA
C3Minimax_H3_FL2VA_PRUNED_Turbo_8step_v1.0_768p_INT8 with no LoRA
C4Ref2VA base + minimax_h3_fl2v_turbo_4step_v1.1_768p LoRA
C5Ref2VA base + minimax_h3_fl2v_turbo_4step_v1.2_768p LoRA
C6Ref2VA base + minimax_h3_ref2v_turbo_4step_v0.1 LoRA
C7 (legacy)Ref2VA base + Larry's original minimax_h3_turbo_v4_step600_ema through the standard ComfyUI LoRA loader. Historical only; not valid for judging the original v4 LoRA.
C7_r1 (corrected)Ref2VA base + the same LoRA through MiniMaxH3TurboLoRA and MiniMaxH3TurboSampler; simple, 4/6/8 steps, and low_vram=false sharp bypass mode.
C8Ref2VA base + minimax_h3_fl2v_turbo_8step_v1.0 LoRA
C9Ref2VA base + minimax_h3_fl2v_turbo_8step_v1.0_768p LoRA
C10Ref2VA base + converted fasth3_6step_converted_dense_datafree LoRA
C11minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot with no LoRA

All primary-grid tests use a one-second duration, fixed seed, and the res_multistep KSampler. The separate Euler grid repeats C1-C11 at the same resolution/step pairs with euler; ksampler in each viewer row makes the two series directly filterable. Only resolution and step count vary within a given case. The Stock control is retained as a separate higher-step reference.

C7 correction: Larry v4 Turbo loader

The archived C7 one-second results were created before the original Larry v4 Turbo LoRA loader requirement was caught. Those legacy rows use the standard ComfyUI LoRA loader and therefore are not a valid quality ranking for the original v4 LoRA. Do not use them to judge C7 against C1-C6/C8-C11.

Use H3_REF2V_RTX3090_C7_ORIGINAL_TURBO_CUSTOM_LOADER.json instead. It uses ComfyUI-MiniMax-H3-Turbo's MiniMax-H3 Turbo LoRA node with low_vram=false (runtime bypass for maximum sharpness), MiniMax-H3 Turbo Sampler, simple scheduler, and 8 steps. A corrected C7 benchmark will be recorded as a separate experiment rather than overwriting the legacy data.

The completed corrected screen is named C7_r1 and is available in the c7_r1_original_turbo_1s Dataset Viewer configuration. The paired comparison below includes only 6 and 8 steps because those are the two step values both the legacy C7 run and C7_r1 share.

[image]

Repeatable technical-quality metrics

quality_metrics.csv and the Dataset Viewer columns include a transparent, deterministic technical score for every primary-grid video. The method is h3-transparent-v1: eight evenly-spaced decoded frames, fixed OpenCV algorithms, and fixed (absolute, not rank-based) score mappings.

The exact scoring implementation is included as tools/score_h3_primary_videos.py; it was run with OpenCV 5.0.0 and NumPy 2.4.4. Re-running it against the unchanged MP4s produces the same values.

FieldMeaning
sharpness_laplacian_variance / sharpness_score_0_100Edge-detail proxy: higher generally means less blur, but can also reward noise or hard edges.
contrast_luma_stddev / contrast_score_0_100Luminance contrast and a fixed score favouring a moderate cinematic range.
clipped_pixel_fraction / exposure_score_0_100Fraction of near-black plus near-white pixels; lower clipping scores higher.
saturation_mean_0_255Raw colour-saturation diagnostic; not part of the composite.
flow_aligned_residual_0_1 / temporal_consistency_score_0_100Frame-to-frame residual after dense optical-flow alignment. High deliberate motion can lower this value.
technical_quality_score_0_10045% sharpness + 25% contrast + 20% exposure + 10% temporal consistency.

These metrics do not score identity accuracy, prompt adherence, audio, motion direction, or artistic preference. They are a consistent technical filter for community review—not a replacement for watching the videos.

Visual comparison sheets

Each tile is a mid-frame from the corresponding one-second render. Columns are C1 through C11; rows are 6, 8, 10, and 12 steps. Render time is shown beneath each tile.

0.4 MP

[image]

0.6 MP

[image]

0.8 MP

[image]

0.98 MP

[image]

Render-time summary

This chart measures render speed only; it is not a visual-quality ranking.

[image]

Euler versus res_multistep

These four paired sheets compare the same C1-C11 model/LoRA, seed, resolution, and step count. Each pair is labelled RM (res_multistep) on the left and EU (euler) on the right; the time below each tile is its measured render time.

0.4 MP

[image]

0.6 MP

[image]

0.8 MP

[image]

0.98 MP

[image]

C3 versus C10 FastH3

Paired comparison at matching resolution and step counts. C3 uses the FL2VA Turbo 8-step model without a LoRA; C10 uses Ref2VA with the converted FastH3 LoRA. Prompt, seed, and duration are held constant.

[image]

results.jsonl is the append-only execution record. results.csv is the current tabular export. During the original Windows file lock, the live export was used to create the published results.csv.

Reproduction notes

The workflow JSON references local ComfyUI model paths. It does not include model weights, LoRAs, RefMods, input images, or audio assets. Install the matching MiniMax H3 nodes and models first, then update local paths as needed.

This repository is an experimental benchmark, not a claim of general model quality. Visual preference and audio quality were assessed separately; identical seed and prompt settings were used within each controlled sweep.