How We Fixed a Silent Killer in Our Mac Studio LLM Cluster
We run a small but powerful distributed inference cluster: five Apple Mac Studio machines (each with 96 GB unified memory), connected over Thunderbolt 5 and running open-weight LLMs via MLX. Everything was…