laya-mlx
Native MLX runtime for fast, local decision models on Apple Silicon
Visit laya-mlx ↗
link: rel="ugc noopener"
laya-mlx is a Python package providing a native MLX runtime for Laya typed decision models. It enables low-latency inference on Apple Silicon hardware, delivering decisions in 7–14 milliseconds on M3 Max without requiring text generation, PyTorch, or external cloud APIs.
Designed for on-device machine learning workflows, laya-mlx integrates with the Laya framework to support structured decision-making tasks with minimal dependencies and maximum performance.
Discussion