AI Topic

AI Image & Video News

Image generation, video AI, computer vision. Curated and summarized from dozens of sources by AIBriefs. RSS

AnalysisVisual AI1 source

MiniMax H3 video model demoed in standard T2V workflow

A Reddit r/StableDiffusion post shows output from MiniMax H3 generated with the standard text-to-video workflow. No benchmark scores, pricing, or release details were included in the post.

AnalysisVisual AI1 source

TaoMate-H3 streams synchronized audio-video on MiniMax H3

TaoMate-H3 is a low-latency streaming audio-video runtime built on MiniMax H3 that generates synchronized audio and video in small chunks, supporting continuous long-form generation at 480p, 768p, and 1080p. It was developed by the Alibaba TaoLive AIGC Team.

AnalysisVisual AI1 source

ChatGPT Images 2.5 vs Nano Banana 2: head-to-head review

Decrypt ran OpenAI's ChatGPT Images 2.5, launched September 8, against Google's Nano Banana 2 across six categories; Nano Banana 2 won three. OpenAI claims up to 50% lower latency than Images 2.0, with GPT-Image-2.5 Flare and Sunburst now in the API.

How-ToVisual AI2 sources

MiniMax RefMod workflows create reusable identities without training

A ComfyUI tutorial and workflow pack builds reusable "refmods" for image, video, or audio from reference material, with no model training required. The tutorial covers preparing training images, and the workflows are shared via a Google Drive folder.

AnalysisVisual AI1 source

Reddit users probe motion-context degradation in video diffusion

A r/StableDiffusion thread examines motion-context degradation, a quality issue in video generation workflows. The poster notes the H3-director node claims a refine pass can fix it, but calls that refine a black box when used with low-level motion-context nodes.

How-ToVisual AI1 source

Reddit user asks how to train character LoRAs for Krea2

A r/StableDiffusion post asks for the most accurate way to train character LoRAs on Krea2, covering face and body type, and requests Hugging Face resources. The poster says existing guides are 1-2 months old and the options are overwhelming.

AnalysisVisual AI1 source

SMACK! LoRA Beta 2 adds gunshots and blood squibs to MiniMax H3

SMACK! Beta 2 is a LoRA for MiniMax H3 (Ref2V) that adds impacts, gunshots, and blood squibs, with no trigger word required. Beta 1 covered fists, weapons, car hits, and falls; the older version was removed from Civitai for gore and now lives on Hugging Face.

AnalysisVisual AI1 source

Users note new image generator adds unprompted slogans

Reddit users report the new image generator inserts unrequested uplifting text onto objects like coffee mugs and background posters. One test prompt of stormtroopers fighting flamethrower ninjas still produced the added slogans.

AnalysisVisual AI1 source

Relight node brings 3D lighting studio to MiniMax H3 in ComfyUI

A community-built ComfyUI node adds a relighting studio for MiniMax H3, letting users place up to three lights on a 3D dome around an image. Each light's type, intensity, and color are configurable, along with background and atmosphere settings.

AnalysisVisual AI2 sources

Reddit user tests long-form AI video generation with H3

An r/StableDiffusion user generated a 1344x768 long-form video at int8/32 steps, taking about 7 hours and hitting 192/192GB of RAM during decoding. No image anchor was used, so the character shifts between invisible seams.

LaunchVisual AI2 sources

MiniMax H3 camera control node lands in ComfyUI

A community node adds a visual camera planner for MiniMax H3 inside ComfyUI: drag the camera around a 3D sphere, place keyframes on a timeline, and the node drives the shot. Posted to r/ComfyUI and r/StableDiffusion, with the model hosted on Civitai.

AnalysisVisual AI1 source

Creator makes TV spot in 4 hours with MiniMax H3

Reddit user aurelm says a 10-second-clip TV job took about 4 hours total, including roughly 3 hours of 720p rendering on an RTX 5090, with a 4K upscale done separately. Prompting was handled locally with Gemma, avoiding non-local models.

AnalysisVisual AI1 source

Reddit user praises H3 text-to-video output

A r/StableDiffusion post shows text-to-video results from H3 generated from a simple prompt, with the poster calling the output "really over the top." No model details, benchmarks, or release information are included.

LaunchDevelopers1 source

TwelveLabs Marengo Embed 3.0 hits Amazon Bedrock Knowledge Bases

TwelveLabs Marengo Embed 3.0 is now generally available as an embedding model in Amazon Bedrock Knowledge Bases, enabling video and image search by meaning. AWS says media, sports analytics, education, security, and retail teams need to find specific moments in video, which remains largely unsearchable today.

AnalysisVisual AI2 sources

Creators share MiniMax H3 video tests on r/StableDiffusion

Two community posts demo MiniMax H3 generations: a Truman Show-style Kirby short built from 30 workflow files and edited in KDEnlive, and a Batman: The Animated Series Harley Quinn clip run on a 4070 Ti Super with 16GB VRAM and 64GB RAM.

AnalysisAI Models1 source

Dreamina Seedance 2.0 720p enters top 5 on Text-to-Video Arena

ByteDance Seed's Dreamina Seedance 2.0 720p debuted at #5 on Artificial Analysis's Text-to-Video Arena with a score of 1267.8. Google's Gemini Omni Flash leads the snapshot at 1324.9, ahead of Alibaba-ATH's HappyHorse-1.1 at 1261.4.

How-ToVisual AI1 source

Reddit users compare speed-up options for Minimax H3 video workflows

A thread on r/StableDiffusion asks which acceleration stack is the gold standard for Minimax H3 reference-to-video workflows on a single RTX 3090 Ti, weighing Comfy Kitchen, Triton, EasyCache and ComfyUI Spectrum. The poster also asks which turbo LoRA works best for the ref2va workflow.

AnalysisVisual AI2 sources

Community tests pit Qwen-Image-Edit-2511 against SenseNova and LLaDA editors

Two r/StableDiffusion comparisons benchmark Qwen-Image-Edit-2511 against SenseNova-U1.5-Lite and LLaDA-Image-Turbo on multi-reference fusion and editing. Testers call Qwen's texture and lighting quality impressive and say it stays on top, though one notes a partial style-transfer test disappointed.

AnalysisVisual AI1 source

Reddit users ask whether Krea 2 Edit model is in development

A r/StableDiffusion thread asks if Krea has confirmed a Krea 2 Edit model, noting Krea 2 is already impressive and an edit variant would be interesting. No confirmation or official indication is cited in the post.

EventVisual AI1 source

Google Arts & Culture brings AI art experiments to Southbank Centre

The inaugural Creative Intelligence weekend at London's Southbank Centre features four Google Arts & Culture interactive experiments, including Yinka Ilori's "Dreaming with Flamingos" built on Google DeepMind's Lyria music model. "Splash Canvas" uses Gemini and fluid dynamics, while "Learn Everything" turns photos of objects into visual metaphors.

AnalysisVisual AI1 source

Recraft V4.1 Utility climbs to #19 on Text-to-Image Arena

Recraft V4.1 Utility rose from #26 to #19 on Artificial Analysis's Text-to-Image Arena, scoring 1026.2, up from 1017.9. The leaderboard is topped by MAI-Image-2.6 (Microsoft AI) at 1145.0, ahead of Reve 2.1 (1126.8) and Nano Banana 2 (1121.8).

AnalysisVisual AI1 source

Matt Wolfe builds AI slop detector

Matt Wolfe attempts to build a tool that detects AI-generated videos from YouTube, TikTok, Instagram, or X links. He finds the task harder than expected.

How-ToVisual AI1 source

Workaround reduces plastic skin in Minimax H3 outputs

A Reddit user shares a node-based tweak that adds detail to Minimax H3 generations, including skin, at no extra computational cost. The method involves inserting a node between SamplerCustomAdvanced and the VAE Decode to reduce contrast.