checking both engines
Left · original

Qwen3-8B

standard decoding
The original model’s response will unfold here.
Output— tok/s
First token
Elapsed
Right · accelerated

A Woolly Qwen3-8B

Woolly accelerated decoding
Woolly
The Woolly response will appear here.
Output— tok/s
First token
Elapsed
Ready for the same prompt in both lanes
Try