Repository index
2 indexed repositories
3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar