The discussion
Most commenters land on the model being good enough to replace API models for coding, because several report real speeds above the claimed 100 tokens/s on a 4090 and one says they stopped using Claude. The sharpest objection is the 2-bit quantization: one commenter cites a paper showing 4-bit usually preserves performance while 2-bit often causes broad degradation, and another notes that published numbers show a 4-bit 27B beating Flash Next at 3 bits. Against that, defenders argue the IQ3 imatrix quants cut unimportant weights surgically and that the Coder variant’s 91% of full-model SWE-bench Verified is the number that matters. A secondary argument is over what surprisingly well means: speed is one thing, benchmarked accuracy another, and almost nobody in the thread posts both.