| | There are plenty of videos around with characters that look like AI or have plastic skin or talk like a robot. What about emotions, or subtle facial expressions, or natural pauses? The solutions and advice usually offered include complex prompt guides or access to paid platforms. I don’t like that. Paid platforms, I mean. I also don’t like complex prompting. So I have been testing. There is something that H3 users hate in agreement: when characters start talking gibberish. Sometimes, the video is longer than the dialogue, so the characters fill the gap with their own nonsense. Other times, there isn’t supposed to be any dialogue, but the character decides to talk anyway. However, in my opinion, gibberish offers a great advantage: because the characters are not constrained by the words we want them to speak, they have way more freedom in the way they say it, meaning they talk more “naturally,” and with more “expression.” So I started wondering: what if we could just let the characters say whatever they wanted, and then edit the audio and video later, changing the nonsense to something intelligible? I didn’t know if that would work. Well, it does, actually. Not only that, but I used one of such clips as a reference, and then downscaled it by 0.25 (for performance reasons), and then used a blur node with a value of “2” to give the model a bit more freedom (and to avoid H3 making it too sharp, which I didn’t want). For consistency, you can also give it an image reference, but I didn’t do that in this case as it had a negative impact on the “acting.” On the final video, you will notice that, in one of the first clips and the last one, the character looks “different”: different clothing, different hair style, different environment. I did this to give the illusion that the footage was taken at a different time. But all of the clips use the same original clip as the reference. You can, however, do it differently, as the voice reference can be given separately. To achieve this using the same reference, you simply prompt for it, use a lower denoise value, or increase the blur. The actual prompt includes the camera style, but you don’t need to do this. The only requirement is to include the dialogue. That’s all. I use the official format, as this reduces the chance of getting more gibberish, but not for every clip, so it’s certainly not necessary. There’s no reason to mention the character if you are using a single one. The original clip was generated at 544×384 and 12 seconds. (For some reason, I get better random stuff at 12 seconds than at 10 seconds or 15 seconds.) The subsequent clips were generated at the same resolution. Then, after editing on Vegas Pro, I rendered the video at 256×180 to help with the lo-fi aesthetic. Finally, I upscaled it to 1920×1350 using Handbrake to avoid Reddit’s compression. So the final video is very low res (technically, ten times smaller than 1080p). Everyone is so focused on making them at 2K, but I just love a good analog/digital horror video. By the way, mine is not “horror” or anything. I simply like the aesthetic. I’m also not claiming it’s any good. I hope it is, but it’s not my place to say. This is the workflow, in case someone wants to test the method I used, though my explanation above should be enough. I hope this helps someone to generate more believable characters. submitted by /u/Etsu_Riot |
In this article, you will learn how to benchmark a deterministic 3-Tiered Graph-RAG system against…
Diffusion-based models decompose sampling into many small Gaussian denoising steps, an assumption that breaks down…
When an AI agent runs, it often needs to buy something to finish a task:…
In recent decades, Ireland has grown into a vibrant hub for global technology. As modernization…
Documents obtained by Democracy Forward show that ICE looked into feeding voter roll data into…
Most of us now carry in our pockets devices capable of running small AI chatbots.…