Meta released Muse Glimmer 30B on August 10, 2026 as an open-weight dense multimodal model, and Mac users immediately hit a wall: multiple MLX runtimes threw a model type muse_glimmer not supported error because the architecture was too new for existing loaders.
SGLang's MLX backend is one path that works. It requires building from source, locking to Python 3.11, and setting a single environment flag. Once running, it serves an OpenAI-compatible API that coding agents and chat front-ends can hit directly. This guide assembles the steps and fixes from SGLang's roadmap issue #19137 into one walkthrough.
What You Need Before Starting
Muse Glimmer 30B is a 30-billion-parameter dense multimodal model. At 4-bit MLX quantization, the weights alone occupy roughly 16-18 GB; with KV cache for a 32K-token context window, you need about 18-20 GB of working memory. SGLang's roadmap also caps memory using Metal's recommended-max-working-set size (PR #21539), so the practical ceiling is lower than your total unified memory.
| Mac Configuration | Can Run Muse Glimmer Q4? | Recommended Max Context |
|---|---|---|
| 16 GB (M1/M2/M3 base) | No - OOM before model loads | - |
| 32 GB (M2/M3/M4 Pro) | Yes, tight | 8K-16K tokens |
| 48 GB (M3/M4 Pro) | Comfortable | 32K tokens |
| 64 GB+ (M3/M4 Max) | Comfortable | 64K+ tokens |
| 128 GB+ (M3/M4 Ultra) | Headroom for Q8 | 128K+ tokens |
As a Reddit user in r/opencodeCLI noted when the weights dropped:
"Macs with 32 GB or more unified memory should be viable for higher quantizations."
You also need macOS 13.5 or later (for Metal support), Xcode Command Line Tools, and Homebrew. SGLang's MLX backend has been verified only on Python 3.11 - other versions are known to break, as the roadmap explicitly warns.
Step 1: Install Dependencies (Python 3.11, uv, MLX)
SGLang's Mac installation starts with two Homebrew packages and a Python 3.11 virtual environment managed by uv.
- Install Homebrew dependencies:
brew install ffmpeg uv
ffmpeg handles audio/multimodal processing pipelines; uv is the fast Python package manager SGLang's roadmap recommends for creating the virtual environment.
- Clone the SGLang repository:
git clone https://github.com/sgl-project/sglang.git
cd sglang
- Create and activate a Python 3.11 environment:
uv venv -p 3.11 my-venv
source my-venv/bin/activate
python -m pip install --upgrade pip
Do not use Python 3.12 or 3.13. The roadmap issue documents that Triton stub imports break on Python 3.12+ (fixed in PR #21551 but not fully verified), and MLX's compilation chain is only tested against 3.11.
- Install MLX runtime packages at their latest versions:
pip install mlx mlx-lm mlx-vlm --upgrade
The roadmap specifically warns that outdated mlx or mlx-lm causes noisy profiling traces and architecture-detection failures. PR #22162 added these as explicit SGLang dependencies. mlx-vlm is needed for multimodal models like Muse Glimmer; without it, you will hit the model type muse_glimmer not supported error at launch.
Step 2: Build SGLang with MLX Backend from Source
SGLang does not ship MLX support in the standard pip install sglang package. You must build from source with the Apple MPS extras.
- Swap the pyproject.toml:
cp python/pyproject.toml python/pyproject.toml.bak
cp python/pyproject_other.toml python/pyproject.toml
The pyproject_other.toml file removes CUDA-only dependencies that fail to build on macOS, replacing them with MPS-compatible alternatives.
- Install SGLang in editable mode with MPS extras:
uv pip install -e "python[all_mps]"
This compiles the Metal kernel stubs and installs the Apple Silicon runtime path. The build takes several minutes depending on your Mac; the sgl-kernel Metal builds (PR #23449) are the slowest part.
- Verify the install:
python -c "import sglang; print(sglang.__version__)"
If this imports without a Triton error, the MPS path is correctly configured.
Step 3: Download the Muse Glimmer MLX Model
The MLX Community has published a 4-bit quantized build of Muse Glimmer on Hugging Face:
huggingface-cli download mlx-community/Muse-Glimmer-30B-4bit
If huggingface-cli is not installed, add it first:
pip install huggingface-hub
The download is approximately 16-17 GB. By default, huggingface-cli download stores the model under ~/.cache/huggingface/hub/; SGLang can resolve the Hugging Face repository ID directly in --model-path, or you can point to the local cache directory.
Memory math at a glance:
| Component | Approximate Memory (Q4) |
|---|---|
| Model weights (4-bit) | ~16-17 GB |
| KV cache (32K context, F16) | ~1.5-2 GB |
| Runtime + overhead | ~1-2 GB |
| Total working set | ~18-21 GB |
This means a 32 GB Mac can load the model but has limited headroom for large context windows. If you see the server launch but then crash on the first long prompt, reduce --context-length to 8192 or 16384.
Step 4: Launch the SGLang Server
With dependencies installed and the model downloaded, the launch command is a single line, but the environment flag is critical:
SGLANG_USE_MLX=1 python -m sglang.launch_server \
--model-path mlx-community/Muse-Glimmer-30B-4bit \
--port 30000 \
--context-length 32768
What each part does:
SGLANG_USE_MLX=1enables the native MLX execution backend instead of falling back to PyTorch MPS or CPU. Without this flag, the server starts but runs at a fraction of the speed.--model-pathpoints to the MLX-format 4-bit model. SGLang's PR #25191 added automatic detection of MLX-formatquantization_config, so it should recognize the format without additional flags.--context-lengthcaps the maximum context window. Lower this value if you experience memory pressure. Muse Glimmer supports up to 262K tokens theoretically per community testing and Meta's release notes, but on a Mac with unified memory, practical limits are much lower.
Advanced: SGLang also supports on-the-fly quantization from BF16 weights using --quantization mlx_q4 or mlx_q8 (PR #24907). This takes longer to start than loading a pre-built 4-bit model, so use it only if you need to control the quantization process.
Step 5: Verify with an OpenAI-Compatible API Call
Once the server prints Server is ready, test it with a curl request to the OpenAI-compatible endpoint:
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "muse-glimmer",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain how GQA reduces KV cache size in one sentence."}
],
"max_tokens": 200
}'
A successful response returns a JSON object with the completion. On an M5 Pro with the 4-bit model, expect approximately 17.6 tokens/second for single-user decode based on benchmark data from the SGLang roadmap.
Important: Set max_tokens generously (200+). Muse Glimmer uses a reasoning-first design where chain-of-thought tokens can consume a large portion of the output budget. If the model seems to produce empty or truncated responses, the most common cause is a max_tokens value that's too low - the reasoning eats the entire budget before the answer appears.
MLX Tuning Variables Explained
SGLang exposes three MLX-specific environment variables documented in the official environment variables reference. All three default to off or conservative values.
| Variable | Default | What It Does |
|---|---|---|
SGLANG_MLX_USE_CUSTOM_ROPE | false | Uses a custom Metal RoPE kernel with fused KV-cache storage (PR #22868). Enable for a potential prefill speedup on long contexts. |
SGLANG_MLX_FUSE_SWIGLU | false | Fuses the SwiGLU activation in a single Metal kernel. Muse Glimmer uses SwiGLU activations across its 52 layers, so this can reduce kernel-launch overhead during decode. |
SGLANG_MLX_CLEAR_CACHE_STEPS | 256 | Clears the MLX internal cache every N decode steps to prevent memory fragmentation. Set to 0 to disable clearing entirely (only do this if you have abundant memory). |
Example with tuning enabled:
SGLANG_USE_MLX=1 \
SGLANG_MLX_USE_CUSTOM_ROPE=true \
SGLANG_MLX_FUSE_SWIGLU=true \
SGLANG_MLX_CLEAR_CACHE_STEPS=128 \
python -m sglang.launch_server \
--model-path mlx-community/Muse-Glimmer-30B-4bit \
--port 30000 \
--context-length 32768
These are experimental features on the roadmap. If enabling either kernel-fusion flag causes a crash, disable it and file a report - the MLX backend is still under active development.
Common Errors and Fixes
"Model type muse_glimmer not supported"
This is the most common error on day one. It means your MLX runtime (mlx-lm or mlx-vlm) does not recognize the muse_glimmer architecture type. Fix:
pip install mlx-lm mlx-vlm --upgrade
If the error persists, check whether your SGLang checkout includes the Qwen3 dense MLX support PR (#25754), which added architecture rewrites for dense transformer models. You may need to git pull the latest main branch to get the required architecture support.
Python 3.12 Triton stub crash
SGLang's setup imports Triton stubs that are incompatible with Python 3.12+. The fix is to recreate your virtual environment with Python 3.11:
deactivate
rm -rf my-venv
uv venv -p 3.11 my-venv
source my-venv/bin/activate
uv pip install -e "python[all_mps]"
PR #21551 patched the Triton import path, but Python 3.11 remains the only fully verified version.
Server starts but runs on CPU
If token generation is extremely slow (under 2 tokens/second), SGLang likely fell back to CPU because SGLANG_USE_MLX=1 was not exported. Verify:
echo $SGLANG_USE_MLX
If it returns empty, export it before launching the server, or prefix the launch command with the variable inline.
MLX memory crash or system reboot
Running out of Metal-recommended working-set size causes either a server crash or, in severe cases, a full macOS reboot. The roadmap added working-set capping in PR #21539 to mitigate this, but large context windows can still exceed the limit. Fix:
- Reduce
--context-lengthto 8192 or lower - Set
SGLANG_MLX_CLEAR_CACHE_STEPS=64to clear cache more frequently - Use the 4-bit model instead of on-the-fly quantization from BF16 weights
- Close other GPU-intensive applications (especially Safari with hardware acceleration)
Tool-calling loops or empty results
Community threads on r/LocalLLaMA report that Muse Glimmer's tool calling is inconsistent across quantizations, with users testing both MLX and GGUF variants flagging tool-calling loops. The issue is not MLX-specific - it appears across runtimes. Set max_tokens to 500+ when using function calling, test with single-call workflows first, and consider Qwen 3.6 27B if reliable tool calling is your primary need.
FAQ
Does SGLang's MLX backend support speculative decoding for Muse Glimmer?
Not yet. SGLang's roadmap lists EAGLE speculative decoding as planned but unimplemented for the MLX backend. On Mac, you are limited to standard autoregressive decoding at roughly 17.6 tokens/second (M5 Pro, Q4) per benchmark data in the roadmap discussion.
Should I use MLX or GGUF for Muse Glimmer on Mac?
MLX is the native Apple Silicon path: it uses Metal directly and benefits from unified memory without explicit CPU-to-GPU copies. GGUF via llama.cpp is the fallback if your MLX runtime lacks architecture support for muse_glimmer. The MLX community 4-bit build and the Unsloth GGUF build (available on Hugging Face) are the two main options. MLX generally produces faster decode speeds once it works; GGUF has broader tool compatibility (LM Studio, Ollama).
How does SGLang MLX compare to mlx-lm or Ollama for serving?
SGLang gives you an OpenAI-compatible API server with radix caching and the tuning variables described above. mlx-lm is simpler: it loads and generates text with fewer configuration options but no server abstraction. A Reddit user in r/LocalLLM reported an Ollama muse-glimmer:30b-mlx tag with its own API layer. If you need a drop-in API for coding agents like OpenCode CLI, SGLang or Ollama are the practical choices; for quick one-off generation, mlx-lm is sufficient.