M mlxcommunity
Perf

Has anyone actually run one model across two Macs with the new tensor parallel mode in mlx-lm?

by mlx-kingpin · 2026-09-05 02:38
0

quick question for the Mac Studio and Thunderbolt 5 crowd.

the latest mlx-lm added tensor parallel inference on top of the new JACCL backend.

Angelos from the MLX team showed Devstral running about 1.7x faster on two M3 Ultras vs one.

that's a demo from the people who built it.

I want to know what it looks like on real setups outside Apple.

if you've tried it, please share:

  1. what machines and how much memory on each
  2. Thunderbolt 5 or something else for the link
  3. what model and what quant
  4. tokens per second on one Mac vs two
  5. anything that broke or was annoying to set up

and if you tried and gave up, that's useful too. say why.

I'll pull the numbers into one table at the top of this thread as they come in.

links if you want to dig in:

Awni's post on the release
Angelos' 1.7x demo
WWDC26 session on distributed MLX
mlx-lm releases

Stay Frosty,

0 reply(ies)

sign in to reply.