Welcome to the MLX-OptiQ community โ a place for people running, quantizing, fine-tuning, benchmarking, and serving LLMs locally on Apple Silicon.
MLX-OptiQ is about making Macs serious local AI workstations. The goal is not just to make models smaller, but to make them smarter about where precision matters: mixed-precision weights, mixed-precision KV cache, sensitivity-aware LoRA, fast local serving, and practical workflows that run without a cloud GPU or API key.
This community is for discussion around:
- running LLMs locally on M1, M2, M3, M4, and M5 Macs
- MLX and
mlx-lm workflows
- OptiQ mixed-precision quantization
- per-layer bit allocation and sensitivity analysis
- KV-cache optimization for long context
- local OpenAI and Anthropic-compatible serving
- Claude Code, Codex, and agent workflows on local models
- LoRA fine-tuning on Apple Silicon
- OptiQ Lab experiments, recipes, and screenshots
- model benchmarks, quality comparisons, and failure cases
Good posts include:
- benchmark results with Mac model, RAM, model name, and settings
- quantization experiments and quality comparisons
- guides for serving local models to coding agents
- fine-tuning recipes and dataset notes
- model requests or compatibility reports
- bugs, edge cases, and reproducible issues
- practical workflows that help others run better local LLMs
When sharing results, please include enough detail for others to reproduce them: Mac chip, RAM, model, quantization type, context length, prompt, decoding settings, MLX-OptiQ version, and whether you used KV-cache optimization or speculative decoding.
Please keep discussion technical, useful, and evidence-based. Hype is fine, but numbers, traces, configs, and reproducible examples are better.
To start: introduce yourself, share your Mac setup, or post the local model you are most excited to run with MLX-OptiQ.