Categories: Image

Pushing MiniMax H3 quality on an RTX 3070 8GB — movie screenshots, voice refs + 0.5MP workflow

Wanted to see how far I could push the quality using what I already have. An RTX 3070 with just 8GB VRAM.

This was done with the standard MiniMax Ref workflow using screenshots from the original movie as character and scene references. I stuck with the standard model rather than Turbo Loras because, at least in my tests, I felt I was losing some of the detail/quality I was trying to preserve.

I also put quite a bit of extra effort into the audio references. For me, getting the voices close makes a huge difference, even a convincing visual starts feeling “AI” very quickly when it has a generic generated voice.

I’m honestly still amazed by what I can get away with on an 8GB VRAM card.

My previous video Penny – Born to Fly video took me about a week to make. This Batman one only took a couple of hours, reference images are still the key in my opinion for great generations.

There was still plenty of rendering, re-rendering, prompt changes and fixing little continuity problems along the way. Definitely not a one click result.

The silly credits were just me having fun and trying to make all the separate renders feel like one little production.

And apologies for the vertical edit, wanted to test it out.

One of the simpler H3 prompts was basically:

Vicki sits at her desk in the same consultation office. Batman crouches extremely low behind a tiny potted plant, with only the two pointed ears of his cowl visible above the leaves. Vicki: "Bruce, I can see your ears." Short pause. Batman, completely deadpan: "Those are leaves." Static camera, same environment and character references, quiet realistic room tone.

Curious what you guys think. Any questions about the workflow, prompting, references or audio are welcome.

submitted by /u/justin_wiggins
[link] [comments]

AI Generated Robotic Content

Share
Published by
AI Generated Robotic Content
Tags: ai images

Recent Posts

Qwen Image 2.1 is an EDITING Beast

Just a few tests with the new Qwen Image 2.1. Although it is not a…

8 hours ago

RAG vs. Fine-Tuning for Domain Adaptation: When to Use Which

In this article, you will learn the mechanical difference between retrieval-augmented generation and fine-tuning, when…

8 hours ago

How to Guide Your Language Flow

We introduce a new method to guide flow matching models. Our approach, which we call…

8 hours ago

From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock

This post is co-written with Mauro Rallo and Patrick van der Plas from HEMA. When…

8 hours ago

Meta Pinky Promises Its Smart Glasses Will Be Private Soon

The company is bringing its Private Processing encryption service to its much-maligned smart glasses.

9 hours ago

When AI disagrees, people change how they perceive the technology—not their original view

When an AI chatbot agrees with our reasoning in resolving a social dilemma, we may…

9 hours ago