Building a Local AI Workstation on a Mac M5 with 32GB RAM: Complete Guide to LLMs, Speech, Vision, Multimodal AI, and Deployment (2026)
Building a Local AI Workstation on a Mac M5 with 32GB RAM
Introduction
Apple Silicon has made local AI practical. A Mac M5 with 32GB unified memory can comfortably run language, coding, speech, vision, multimodal, and lightweight video-analysis workloads.
This guide uses the MLX ecosystem as the main deployment approach. MLX is Apple's machine learning framework optimized for Apple Silicon, making it a strong choice for local LLM inference, multimodal models, and efficient use of unified memory and Metal GPU acceleration.
Hardware Recommendations
- M5 32GB: ideal for 14B-class models
- M5 64GB: suitable for 27B–32B models
- M5 Max 64GB+: best for multitasking and larger models
Environment Setup
Install Homebrew
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
Install Dependencies
brew install git cmake python ffmpeg wget aria2
Create an MLX Python Environment
mkdir -p ~/AI/mlx
cd ~/AI/mlx
python3 -m venv mlx-env
source mlx-env/bin/activate
pip install -U pip
Install MLX Packages
pip install mlx
pip install mlx-lm
pip install mlx-vlm
Verify GPU Acceleration
python -c "import mlx.core as mx; print(mx.default_device())"
If the output includes gpu, MLX is using Apple Metal GPU acceleration.
Deploying LLMs with MLX-LM
MLX-LM is the core tool for running local language models on Apple Silicon. It supports models such as Qwen, DeepSeek, Gemma, Llama, and Mistral.
Run Qwen3-14B
mlx_lm.generate \
--model mlx-community/Qwen3-14B-4bit \
--prompt "Explain the advantages of MLX for local LLM deployment on Apple Silicon."
Run DeepSeek-R1 Distill 14B
mlx_lm.generate \
--model mlx-community/DeepSeek-R1-Distill-Qwen-14B-4bit \
--prompt "Explain the basic architecture of a RAG system."
Run Gemma 3
mlx_lm.generate \
--model mlx-community/gemma-3-27b-it-4bit \
--prompt "Write a Python FastAPI example."
Start an OpenAI-Compatible API Server
python -m mlx_lm.server \
--model mlx-community/Qwen3-14B-4bit \
--host 127.0.0.1 \
--port 8080
API endpoint:
http://127.0.0.1:8080/v1/chat/completions
This API can be connected to:
- Cherry Studio
- LobeChat
- Dify
- AnythingLLM
- Open WebUI
- Custom Python / Node.js applications
API Request Example
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mlx-community/Qwen3-14B-4bit",
"messages": [
{"role": "user", "content": "Explain Apple MLX in three sentences."}
],
"temperature": 0.7
}'
Optional Open WebUI Integration
If you prefer a ChatGPT-like web interface, you can connect Open WebUI to the MLX-LM server.
docker run -d \
-p 3000:8080 \
-v open-webui:/app/backend/data \
--name open-webui \
--restart always \
ghcr.io/open-webui/open-webui:main
Open:
http://localhost:3000
Configure OpenAI Compatible API in Open WebUI:
Base URL: http://host.docker.internal:8080/v1
API Key: any value
Recommended Models
General Chat
- Qwen3-14B-4bit
- Gemma 3 12B / 27B 4bit
Reasoning
- DeepSeek-R1-Distill-Qwen-14B-4bit
Coding
- Qwen2.5-Coder-14B-4bit
- DeepSeek-R1-Distill-Qwen-14B-4bit
Speech Recognition
MLX Whisper
MLX Whisper is a good fit for local speech recognition on Apple Silicon.
pip install mlx-whisper
Transcribe audio:
mlx_whisper audio.wav
Specify model and language:
mlx_whisper audio.wav \
--model large-v3-turbo \
--language zh
It is suitable for meeting transcription, subtitle generation, and voice-to-text workflows.
SenseVoice
SenseVoice performs well for Chinese speech recognition and mixed Chinese-English scenarios.
mkdir -p ~/AI/sensevoice
cd ~/AI/sensevoice
python3 -m venv venv
source venv/bin/activate
pip install -U pip
pip install funasr modelscope torch torchaudio
Test script:
from funasr import AutoModel
model = AutoModel(
model="iic/SenseVoiceSmall",
trust_remote_code=True
)
result = model.generate(
input="audio.wav",
language="zh"
)
print(result)
Text-to-Speech
CosyVoice2
CosyVoice2 is suitable for Chinese speech synthesis, voice cloning, and low-latency voice applications.
mkdir -p ~/AI/cosyvoice
cd ~/AI/cosyvoice
git clone https://github.com/FunAudioLLM/CosyVoice.git
cd CosyVoice
python3 -m venv venv
source venv/bin/activate
pip install -U pip
pip install -r requirements.txt
pip install modelscope
Download the model:
python -c "from modelscope import snapshot_download; snapshot_download('iic/CosyVoice2-0.5B', local_dir='pretrained_models/CosyVoice2-0.5B')"
Start the Web Demo:
python webui.py \
--port 50000 \
--model_dir pretrained_models/CosyVoice2-0.5B
Open:
http://localhost:50000
Vision and Multimodal Models
MLX-VLM can run vision-language models on Apple Silicon for image understanding, OCR, screenshot analysis, document analysis, and chart interpretation.
Install
pip install mlx-vlm
Qwen2.5-VL-7B
python -m mlx_vlm.generate \
--model mlx-community/Qwen2.5-VL-7B-Instruct-4bit \
--image image.jpg \
--prompt "Analyze the main content of this image."
Gemma 3 Vision
python -m mlx_vlm.generate \
--model mlx-community/gemma-3-12b-it-4bit \
--image image.jpg \
--prompt "Describe this image."
Pixtral
python -m mlx_vlm.generate \
--model mlx-community/pixtral-12b-4bit \
--image report.png \
--prompt "Analyze the chart and summarize the key insights."
Recommended models:
- Qwen2.5-VL-7B
- Gemma 3
- Pixtral
- MiniCPM-V
Image Generation
For image generation, ComfyUI with FLUX is still the more mature choice. MLX is better suited for LLM and VLM inference, while ComfyUI currently has a stronger image-generation ecosystem.
ComfyUI
mkdir -p ~/AI
cd ~/AI
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
python3 -m venv venv
source venv/bin/activate
pip install -U pip
pip install -r requirements.txt
Start ComfyUI:
python main.py
Open:
http://127.0.0.1:8188
FLUX.1-dev
Use Q4/Q5 quantized versions on a 32GB Mac.
Recommended model folders:
ComfyUI/models/checkpoints
ComfyUI/models/clip
ComfyUI/models/vae
ComfyUI/models/unet
Video Understanding
Local video understanding usually does not feed the entire video directly into a model. A practical workflow is:
Video → Frame Extraction → MLX-VLM Analysis → MLX-LM Summary
Extract Frames
mkdir -p frames
ffmpeg -i video.mp4 \
-vf "fps=1/5" \
frames/frame_%04d.jpg
This extracts one frame every five seconds.
You can then analyze key frames with Qwen2.5-VL, Gemma 3, or Pixtral, and summarize the results with Qwen3 or DeepSeek.
Directory Structure
~/AI
├── mlx
│ └── mlx-env
├── models
├── ComfyUI
├── CosyVoice
├── sensevoice
├── datasets
├── outputs
└── scripts
Performance Tips
- Use 4bit models on a 32GB Mac
- Use 14B models as the primary daily driver
- 27B models can be tested in 4bit, but are not ideal for multitasking
- Avoid loading multiple large models at the same time
- Keep context windows reasonable
- Prefer pre-converted
mlx-communitymodels - Use
mlx_lm.serveras a unified OpenAI-compatible API layer - Store large files, video frames, and generated outputs under
~/AI/outputs
Recommended Stack
Chat
Qwen3-14B-4bit
Reasoning
DeepSeek-R1-Distill-Qwen-14B-4bit
Coding
Qwen2.5-Coder-14B-4bit
ASR
MLX Whisper + SenseVoice
TTS
CosyVoice2
Vision
Gemma 3 + Pixtral + Qwen2.5-VL-7B
Image Generation
FLUX.1-dev + ComfyUI
Video Understanding
ffmpeg + MLX-VLM + MLX-LM
Conclusion
The goal of a 32GB Mac is not to run the largest possible model. The real value comes from building a balanced AI ecosystem that covers language, coding, speech, vision, and multimodal workflows.
With MLX-LM, MLX-VLM, MLX Whisper, ComfyUI, SenseVoice, and CosyVoice2, a modern Apple Silicon Mac can become a capable local AI workstation. Compared with generic deployment tools, MLX is closer to the hardware characteristics of Apple Silicon and can better leverage unified memory and Metal GPU acceleration.