A coding agent that fits on one GPU: Qwen3.8-27B hits 131 tokens per second on RTX cards
NVIDIA announced day-zero RTX support for Qwen3.8-27B, a 27-billion-parameter open model. It is sized to fit a single GPU and reaches 131 tokens per second on a GeForce RTX 5090.