AIREITER

Muse Glimmer MLX Setup on Mac: SGLang Backend Guide

Last Updated: 2026-08-11 00:54:14

Meta released Muse Glimmer 30B on August 10, 2026 as an open-weight dense multimodal model, and Mac users immediately hit a wall: multiple MLX runtimes threw a model type muse_glimmer not supported error because the architecture was too new for existing loaders.

SGLang's MLX backend is one path that works. It requires building from source, locking to Python 3.11, and setting a single environment flag. Once running, it serves an OpenAI-compatible API that coding agents and chat front-ends can hit directly. This guide assembles the steps and fixes from SGLang's roadmap issue #19137 into one walkthrough.

What You Need Before Starting

Muse Glimmer 30B is a 30-billion-parameter dense multimodal model. At 4-bit MLX quantization, the weights alone occupy roughly 16-18 GB; with KV cache for a 32K-token context window, you need about 18-20 GB of working memory. SGLang's roadmap also caps memory using Metal's recommended-max-working-set size (PR #21539), so the practical ceiling is lower than your total unified memory.

Mac ConfigurationCan Run Muse Glimmer Q4?Recommended Max Context
16 GB (M1/M2/M3 base)No - OOM before model loads-
32 GB (M2/M3/M4 Pro)Yes, tight8K-16K tokens
48 GB (M3/M4 Pro)Comfortable32K tokens
64 GB+ (M3/M4 Max)Comfortable64K+ tokens
128 GB+ (M3/M4 Ultra)Headroom for Q8128K+ tokens

As a Reddit user in r/opencodeCLI noted when the weights dropped:

"Macs with 32 GB or more unified memory should be viable for higher quantizations."

You also need macOS 13.5 or later (for Metal support), Xcode Command Line Tools, and Homebrew. SGLang's MLX backend has been verified only on Python 3.11 - other versions are known to break, as the roadmap explicitly warns.

Step 1: Install Dependencies (Python 3.11, uv, MLX)

SGLang's Mac installation starts with two Homebrew packages and a Python 3.11 virtual environment managed by uv.

  1. Install Homebrew dependencies:
brew install ffmpeg uv

ffmpeg handles audio/multimodal processing pipelines; uv is the fast Python package manager SGLang's roadmap recommends for creating the virtual environment.

  1. Clone the SGLang repository:
git clone https://github.com/sgl-project/sglang.git
cd sglang
  1. Create and activate a Python 3.11 environment:
uv venv -p 3.11 my-venv
source my-venv/bin/activate
python -m pip install --upgrade pip

Do not use Python 3.12 or 3.13. The roadmap issue documents that Triton stub imports break on Python 3.12+ (fixed in PR #21551 but not fully verified), and MLX's compilation chain is only tested against 3.11.

  1. Install MLX runtime packages at their latest versions:
pip install mlx mlx-lm mlx-vlm --upgrade

The roadmap specifically warns that outdated mlx or mlx-lm causes noisy profiling traces and architecture-detection failures. PR #22162 added these as explicit SGLang dependencies. mlx-vlm is needed for multimodal models like Muse Glimmer; without it, you will hit the model type muse_glimmer not supported error at launch.

Step 2: Build SGLang with MLX Backend from Source

SGLang does not ship MLX support in the standard pip install sglang package. You must build from source with the Apple MPS extras.

  1. Swap the pyproject.toml:
cp python/pyproject.toml python/pyproject.toml.bak
cp python/pyproject_other.toml python/pyproject.toml

The pyproject_other.toml file removes CUDA-only dependencies that fail to build on macOS, replacing them with MPS-compatible alternatives.

  1. Install SGLang in editable mode with MPS extras:
uv pip install -e "python[all_mps]"

This compiles the Metal kernel stubs and installs the Apple Silicon runtime path. The build takes several minutes depending on your Mac; the sgl-kernel Metal builds (PR #23449) are the slowest part.

  1. Verify the install:
python -c "import sglang; print(sglang.__version__)"

If this imports without a Triton error, the MPS path is correctly configured.

Step 3: Download the Muse Glimmer MLX Model

The MLX Community has published a 4-bit quantized build of Muse Glimmer on Hugging Face:

huggingface-cli download mlx-community/Muse-Glimmer-30B-4bit

If huggingface-cli is not installed, add it first:

pip install huggingface-hub

The download is approximately 16-17 GB. By default, huggingface-cli download stores the model under ~/.cache/huggingface/hub/; SGLang can resolve the Hugging Face repository ID directly in --model-path, or you can point to the local cache directory.

Memory math at a glance:

ComponentApproximate Memory (Q4)
Model weights (4-bit)~16-17 GB
KV cache (32K context, F16)~1.5-2 GB
Runtime + overhead~1-2 GB
Total working set~18-21 GB

This means a 32 GB Mac can load the model but has limited headroom for large context windows. If you see the server launch but then crash on the first long prompt, reduce --context-length to 8192 or 16384.

Step 4: Launch the SGLang Server

With dependencies installed and the model downloaded, the launch command is a single line, but the environment flag is critical:

SGLANG_USE_MLX=1 python -m sglang.launch_server \
  --model-path mlx-community/Muse-Glimmer-30B-4bit \
  --port 30000 \
  --context-length 32768

What each part does:

  • SGLANG_USE_MLX=1 enables the native MLX execution backend instead of falling back to PyTorch MPS or CPU. Without this flag, the server starts but runs at a fraction of the speed.
  • --model-path points to the MLX-format 4-bit model. SGLang's PR #25191 added automatic detection of MLX-format quantization_config, so it should recognize the format without additional flags.
  • --context-length caps the maximum context window. Lower this value if you experience memory pressure. Muse Glimmer supports up to 262K tokens theoretically per community testing and Meta's release notes, but on a Mac with unified memory, practical limits are much lower.

Advanced: SGLang also supports on-the-fly quantization from BF16 weights using --quantization mlx_q4 or mlx_q8 (PR #24907). This takes longer to start than loading a pre-built 4-bit model, so use it only if you need to control the quantization process.

Step 5: Verify with an OpenAI-Compatible API Call

Once the server prints Server is ready, test it with a curl request to the OpenAI-compatible endpoint:

curl http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "muse-glimmer",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Explain how GQA reduces KV cache size in one sentence."}
    ],
    "max_tokens": 200
  }'

A successful response returns a JSON object with the completion. On an M5 Pro with the 4-bit model, expect approximately 17.6 tokens/second for single-user decode based on benchmark data from the SGLang roadmap.

Important: Set max_tokens generously (200+). Muse Glimmer uses a reasoning-first design where chain-of-thought tokens can consume a large portion of the output budget. If the model seems to produce empty or truncated responses, the most common cause is a max_tokens value that's too low - the reasoning eats the entire budget before the answer appears.

MLX Tuning Variables Explained

SGLang exposes three MLX-specific environment variables documented in the official environment variables reference. All three default to off or conservative values.

VariableDefaultWhat It Does
SGLANG_MLX_USE_CUSTOM_ROPEfalseUses a custom Metal RoPE kernel with fused KV-cache storage (PR #22868). Enable for a potential prefill speedup on long contexts.
SGLANG_MLX_FUSE_SWIGLUfalseFuses the SwiGLU activation in a single Metal kernel. Muse Glimmer uses SwiGLU activations across its 52 layers, so this can reduce kernel-launch overhead during decode.
SGLANG_MLX_CLEAR_CACHE_STEPS256Clears the MLX internal cache every N decode steps to prevent memory fragmentation. Set to 0 to disable clearing entirely (only do this if you have abundant memory).

Example with tuning enabled:

SGLANG_USE_MLX=1 \
SGLANG_MLX_USE_CUSTOM_ROPE=true \
SGLANG_MLX_FUSE_SWIGLU=true \
SGLANG_MLX_CLEAR_CACHE_STEPS=128 \
python -m sglang.launch_server \
  --model-path mlx-community/Muse-Glimmer-30B-4bit \
  --port 30000 \
  --context-length 32768

These are experimental features on the roadmap. If enabling either kernel-fusion flag causes a crash, disable it and file a report - the MLX backend is still under active development.

Common Errors and Fixes

"Model type muse_glimmer not supported"

This is the most common error on day one. It means your MLX runtime (mlx-lm or mlx-vlm) does not recognize the muse_glimmer architecture type. Fix:

pip install mlx-lm mlx-vlm --upgrade

If the error persists, check whether your SGLang checkout includes the Qwen3 dense MLX support PR (#25754), which added architecture rewrites for dense transformer models. You may need to git pull the latest main branch to get the required architecture support.

Python 3.12 Triton stub crash

SGLang's setup imports Triton stubs that are incompatible with Python 3.12+. The fix is to recreate your virtual environment with Python 3.11:

deactivate
rm -rf my-venv
uv venv -p 3.11 my-venv
source my-venv/bin/activate
uv pip install -e "python[all_mps]"

PR #21551 patched the Triton import path, but Python 3.11 remains the only fully verified version.

Server starts but runs on CPU

If token generation is extremely slow (under 2 tokens/second), SGLang likely fell back to CPU because SGLANG_USE_MLX=1 was not exported. Verify:

echo $SGLANG_USE_MLX

If it returns empty, export it before launching the server, or prefix the launch command with the variable inline.

MLX memory crash or system reboot

Running out of Metal-recommended working-set size causes either a server crash or, in severe cases, a full macOS reboot. The roadmap added working-set capping in PR #21539 to mitigate this, but large context windows can still exceed the limit. Fix:

  • Reduce --context-length to 8192 or lower
  • Set SGLANG_MLX_CLEAR_CACHE_STEPS=64 to clear cache more frequently
  • Use the 4-bit model instead of on-the-fly quantization from BF16 weights
  • Close other GPU-intensive applications (especially Safari with hardware acceleration)

Tool-calling loops or empty results

Community threads on r/LocalLLaMA report that Muse Glimmer's tool calling is inconsistent across quantizations, with users testing both MLX and GGUF variants flagging tool-calling loops. The issue is not MLX-specific - it appears across runtimes. Set max_tokens to 500+ when using function calling, test with single-call workflows first, and consider Qwen 3.6 27B if reliable tool calling is your primary need.

FAQ

Does SGLang's MLX backend support speculative decoding for Muse Glimmer?

Not yet. SGLang's roadmap lists EAGLE speculative decoding as planned but unimplemented for the MLX backend. On Mac, you are limited to standard autoregressive decoding at roughly 17.6 tokens/second (M5 Pro, Q4) per benchmark data in the roadmap discussion.

Should I use MLX or GGUF for Muse Glimmer on Mac?

MLX is the native Apple Silicon path: it uses Metal directly and benefits from unified memory without explicit CPU-to-GPU copies. GGUF via llama.cpp is the fallback if your MLX runtime lacks architecture support for muse_glimmer. The MLX community 4-bit build and the Unsloth GGUF build (available on Hugging Face) are the two main options. MLX generally produces faster decode speeds once it works; GGUF has broader tool compatibility (LM Studio, Ollama).

How does SGLang MLX compare to mlx-lm or Ollama for serving?

SGLang gives you an OpenAI-compatible API server with radix caching and the tuning variables described above. mlx-lm is simpler: it loads and generates text with fewer configuration options but no server abstraction. A Reddit user in r/LocalLLM reported an Ollama muse-glimmer:30b-mlx tag with its own API layer. If you need a drop-in API for coding agents like OpenCode CLI, SGLang or Ollama are the practical choices; for quick one-off generation, mlx-lm is sufficient.