H3-metal – Native MiniMax-H3 inference for Apple Silicon

292 points - today at 1:22 AM

Source

Comments

Meleagris today at 3:12 AM
I've been using MiniMax H3 on my M5 Pro 64GB MacBook Pro through ComfyUI. It works extremely well.

I had to modify the default ComfyUI workflows to use a GGUF quant (city96's ComfyUI-GGUF custom node, UnetLoaderGGUF in place of the stock loader) [0].

I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest.

The main issue is speed, a ~9-second 480x864 clip at 20 steps takes me a bit over an hour. So this will be cool to try for the speed up alone.

There's a lot of great information and workflows available to follow on the r/StableDiffusion subreddit.

[0] https://huggingface.co/Abiray/MiniMax-H3-GGUF/tree/main/unet

antirez today at 6:14 AM
In the AMA Minimax said that H3 could support sparse attention, that would be a huge speedup! I wonder if there are any news on that. H3 is very cool. EDIT: testing a --sparse-attention optional mode based on what they said in the Reddit post.
linzhangrun today at 6:39 AM
On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.

Put Codex to work on deploying it now, hoping the speed can improve quite a lot :-) Thanks anyway

aaqaishtyaq today at 10:32 AM
Anyone tried it with M4 Pro, 48GB of memory?
diddid today at 4:06 AM
This is where the DGX spark makes up a bit of the ground it loses on llm work, diffusion and cuda go together like peanut butter and jelly.
iamyatin today at 10:06 AM
Noob question to all, is there any open source coding model that I can run on Mac mini 16gb?
TechSquidTV today at 2:42 AM
This still requires 128Gb of memory, right? Me and my lowly 96Gb, like a commoner; missing out on the fun.
c0rruptbytes today at 3:22 AM
wow antirez does not sleep
myshapeprotocol today at 8:53 AM
Native inference optimization for Apple Silicon is such a game-changer for local-first workflows. Incredible performance work.
abhinai today at 2:55 AM
How similar are Jeff Dean and Salvatore Sanfilippo?
tipiirai today at 4:38 AM
I'd love to know what the alternatives are and how this is better
luciana1u today at 10:13 AM
neat — now I just need a machine with the memory bandwidth to render the three-second clip of my cat before the cat itself forgets what happened
yieldcrv today at 7:28 AM
Alright I’ve been afraid to ask but have been having trouble finding

What are some adult entertainment workflows in comfyui, I need best loras, best prompts to start with

and the communities, are they on telegram or something?

songhonglei1985 today at 4:49 AM
[dead]