EXL3 inference on Apple Silicon.
I wrote an EXL3 inference engine for Apple Silicon using MLX and custom Metal kernels. It loads compressed models and runs them without expanding the weights into dense tensors. I work on the model loader, GPU kernels and benchmarks.
MLXL3 Desktop is the SwiftUI app I built around it. Download a model, chat locally and track memory use. Conversations are saved, MCP tools are optional, and the engine is bundled with the app.
MLXL3 · EXL3 3.10 bpw4.02 GB
MLX · 8-bit baseline9.04 GB
Peak allocation for LFM2.5-8B-A1B on a 10-core M5. The paired 12-run warm benchmark reports 65.8 generated tokens/s for MLXL3 and 58.9 for MLX 8-bit. Different weight precisions; results apply to this model and setup.