I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.”
I'm getting somewhat conflicted. There is a Twitter account I follow that has great takes and lots of deeply personal posts... BUT he obviously uses an LLM in his writing pipeline somewhere. Too many AIisms scattered about to ignore.
I'm not thrilled with it, but he is obviously using it to improve his writing overall- to communicate some great ideas that are personal and germane. I've decided that being too inflexible serves no one. If it is true slop, I'll not revisit the writer in the future- if they are using AI to polish writing that at its core is a unique voice, I'll accept it and learn to live with it...
> “It's worth being precise about where the benefit comes from, because it isn't raw speed.“
What's funny it's that is as if AI "learned" to speak english but not really. People simply don't speak using those strange constructs: those sentences sound a bit like if a "Karen" was trying to make a point.
What's scary, to me, as a dev using AI, is that those LLMs do the same thing with code: it looks like proper code, but it really ain't so once you dig a bit.
It's verbose and doesn't add anything: it's just infinite verbiage / sloppy-pasta.
Crazy thing though it's that it's 2026 and apparently devs can't be bothered to copy/paste their sloppy-pasta LLMish into a de-sloppifier before publishing blog posts.
What's the tell, other than it butchering the context (bad writing)? The article as a whole doesn't seem horribly LLMy. The overall topic is pretty dry, and that sentence is honkingly bad, but not all bad writing is LLM-originated -- sometimes insufficient proofreading is at fault.
Nice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their unquantised weights, even though KV quantisation can degrade quality more than weight quantisation.
However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Secondly, the evaluation suite they use to claim that FP8 KV quantisation is indistinguishable is noticeably lacking coding benchmarks; in long-running tasks, minor tool call errors compound over time.
vLLM's study also concluded that "FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss". Their benchmarks include LiveCodeBench 6.
> None of this would matter if it changed the model's answers
If they want to assert that the answers don’t change, then perhaps they should calculate the statistical distance between the token probability outputs or something to that effect. I doubt the results would indicate that the answers don’t change by any reasonable interpretation.
So they quantize models, only tell about it in the blog post (instead of a warning on the model page), and even in the blog post pretend there's no difference by benchmarking on small context tasks many of which are saturated. Coding agents will probably be severely negatively affected by KV quantization.
I'd say serving quantized models without saying so on the "store" page is fraud.
Yeah, I've noticed them recently on OR and whitelisted. Then I backed-off pretty quickly after seeing the cache hit rates. It was also rather revealing to see how some provider hit rates differ when you are using them directly vs via OR.
22 comments
[ 0.27 ms ] story [ 36.8 ms ] threadI love AI, but I really hate reading it.
I'm not thrilled with it, but he is obviously using it to improve his writing overall- to communicate some great ideas that are personal and germane. I've decided that being too inflexible serves no one. If it is true slop, I'll not revisit the writer in the future- if they are using AI to polish writing that at its core is a unique voice, I'll accept it and learn to live with it...
What's funny it's that is as if AI "learned" to speak english but not really. People simply don't speak using those strange constructs: those sentences sound a bit like if a "Karen" was trying to make a point.
What's scary, to me, as a dev using AI, is that those LLMs do the same thing with code: it looks like proper code, but it really ain't so once you dig a bit.
It's verbose and doesn't add anything: it's just infinite verbiage / sloppy-pasta.
Crazy thing though it's that it's 2026 and apparently devs can't be bothered to copy/paste their sloppy-pasta LLMish into a de-sloppifier before publishing blog posts.
However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Secondly, the evaluation suite they use to claim that FP8 KV quantisation is indistinguishable is noticeably lacking coding benchmarks; in long-running tasks, minor tool call errors compound over time.
https://vllm-project.github.io/2026/04/22/fp8-kvcache.html
> None of this would matter if it changed the model's answers
If they want to assert that the answers don’t change, then perhaps they should calculate the statistical distance between the token probability outputs or something to that effect. I doubt the results would indicate that the answers don’t change by any reasonable interpretation.
Maybe the results are still good enough.
Why… I wanted to see if it’s worth it to use cloudflare’s endpoint but I can’t even see the pricing
I'd say serving quantized models without saying so on the "store" page is fraud.
What is the typical job title and/or skillset for this?
We let all traffic get MITM'd, now we're letting our AI conversation get tracked. Cloudflare reeks like a US Honeypot.
One vCPU means nothing, which chip? which SIMD? what RAM speed? what storage? what latency?
Quantization is the next layer of lies
Society is moving towards offloading intelligence to clouid overlords, since nothing runs on your computer anymore, you can't trust anything
And since datacenters occupy a physical space, they are building a monopoly defacto
Unless transparency becomes mandatory, we are headed towards the biggest self sabotage mankind has ever witnessed