Lemonade Server version 11 transforms a Ryzen AI Max+ 395 mini PC into a fully local, multi-modal AI orchestration platform. A single chat request to the bundled RPG-HaloTales-V1 pipeline — combining Qwen3.6-27B, Flux, Whisper, and Kokoro — produced text, an image, and narrated audio without any cloud involvement. The release also adds a Hybrid Router that dispatches prompts to the most appropriate model using LLM-based or rule-based classification, though its default privacy gate fails open and should be reconfigured. An image-to-3D pipeline using Trellis generates glTF files in under two minutes, though results are convincing only from the front. Benchmark data shows the iGPU is 3.4x faster than the NPU for token generation, and the machine can hold six concurrent backends consuming 117 GB of its 128 GB unified memory. Beyond the announced features, Lemonade 11 quietly includes image editing, upscaling, music generation, sound effects, and retrieval components.

6m read timeFrom xda-developers.com
Post cover image
Table of contents
One request, and it wrote, drew, and narrated a tavernThe router that picks your model has a privacy gate, and it fails openIt makes 3D models but not the kind you're picturingThe NPU is the slowest way to run a model on this machine
79 Impressions