Kimi K3 Now Available via Telnyx Inference API (telnyx.com)

132 points by fionaattelnyx ↗ HN
Moonshot AI released open weights for Kimi K3 today and it's live on Telnyx Inference. The architecture and Moonshot's own benchmarks are in their technical blog. This post is about running it on Telnyx.

What we are adding: K3 is now available via the Telnyx Inference API, hosted on GPUs that we own and operate.

This matters due to the size of this model. A 2.8T model needs dedicated infra to serve well. Because we own and operate the GPUs, we can contorl throughput. with no inter-provider hops and no cloud tenant, latency is minimized. We have GPUs in each of the US, EU, APAC, and MENA. Inference runs in the region you pick, with zero data retention. We do not store prompts or completions after the response returns. Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.

We have not benchmarked K3 ourselves yet. Moonshot's own numbers and early third-party evaluations put it at frontier level for coding and agentic work, trailing only Claude Fable 5 and GPT 5.6 Sol on aggregate. Full breakdown in their blog.

Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens. Prompt caching enabled by default. Served via an OpenAI-compatible endpoint so you can test easily.

Model ID: [MODEL_ID] API: https://api.telnyx.com/v2/ai/chat/completions Docs: https://developers.telnyx.com/docs/inference Technical blog (Moonshot): https://www.kimi.com/blog/kimi-k3

43 comments

[ 0.15 ms ] story [ 42.6 ms ] thread
(comment deleted)
Very cool. What are your throughput and latency like?
Telnyx was cool until they started demanding KYC. I would use them for burner phone numbers until they started saying they needed my government ID. Fuck that.

OVH same thing. Tried to buy a VPS from them some years back and they said no VPS unless I provided ID. Would not refund me. Tried to dispute but my bank just gave me a credit instead.

We've been fighting this battle with the FCC. They tried fining us (and then dropped it), but they continue to push the issue via NPRMs. We generally think KYC is ineffective due to ID mules. It's obviously a privacy issue as well.
what quantization? FP4?

    > Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.

is not compatible with

    > Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens.
since you asserted something false and bizarre, how about telling us what is your markup?
10% cheaper than official! Let the inference pricing wars begin!
I'm really curious how far and how fast prices will drop (if at all)!
any of these providers are HIPAA compliant?
Amazon Bedrock, but they don't have this model yet
At Telnyx, we are willing to enter into a BAA. Just hit up our sales team.
[dead]
We'll be pushing benchmarks over the next few days, will update here when we have them
I think telnyx is a good product, with the only stain to its name being the supply chain attack on their python library.

But I don't feel like providing inference is a professional move, it feels like out of scope for a telephony IaaS company. Feels like a FOMO moment where a reputation of years is crashed in a couple of weekends of being drawn into a fad.

And the fact that it's a chinese model doesn't quite help? I guess it's on brand with the 'cheap' pay as you go brand telnyx might already be associated to.

But more so it reads like Telnyx is trying to 'jump' into the trend of the 'open weights' discussion to compete with closed source incumbents. But we are at the tail end of the boom, anti ai sentiment is ever growing, customers now despise AI, especially in support channels, which is presumably the hook that Telnyx would have into 'AI'(LLMs). At this stage any company or individual that tries to join into the buzzword fueled cycle will pay the full fixed cost reputational price, but only reap the leftover hay from when the sun shone.

AI(LLM) on support channels is essentially a decapitalization of a company/brand, the company has a reputation that customers value, and might be worth good money in the market, and by implementing AI (LLMs) on support, a lot of costs can be cut, while the brand loses value, not sustainable. And by Telnyx (or any B2B company)asking their clients to participate in this decapitalization move, they essentially gamble their reputation as well, albeit with better odds as shovel sellers.

Large profitable corporations with very poor customer support have existed long before AI, so why should it be different? Using AI to provide poor customer service is just an implementation detail.
Companies changing their scope or opening side products is not at all unusual. Microsoft is Windows, what's this Xbox thing? Google is search, why do they have email? Y Combinator is a startup accelerator, why'd they make their own Reddit?
Yeah, the pypi thing sucked, and we've taken the necessary measures to prevent something like that from happening again.

We're moving past "telephony" and focusing on building more agentic primitives at the telecom edge.

This isn't a recent jump. We've been on this path for some time.

Disagree with you on the customer support side and re: decapitalization in general. People will want to talk to competent bots. We have people literally prompting our AI agents (which are capable of doing troubleshooting) in our shared Slack channels. You simply cannot beat the speed with which they can provide a quality response.

We remain committed to delivering the highest quality product at the lowest possible price point - across all of our product lines.

Jevon's paradox depends on how cheap the tokens get as the price of tokens get driven to zero as intelligence gets better and cheaper.
Well this is a very large and expensive model. Try something like Qwen-7B if you want ultra-cheap.
In my first interaction ("hi there kimi k3!"), Kimi K3 identified twice out of three times as Claude:

> Hi there! Quick note — I'm actually Claude, made by Anthropic, not Kimi. But no worries!

https://imgur.com/a/jqpc2Jc

and

> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.

https://imgur.com/a/AKxeysH

Models don't have an inherent identity. It should be obvious by now that every models trains on public AI chat session transcripts. I've seen Claude say it's Qwen.
This happens with other models too - Gemini often identifies as ChatGPT for me, confusing many a debugging attempt
How many times do people need to point out that every model has this behavior until this stops being posted?
I remember when Claude identified as ChatGPT a long time ago. It proves nothing else than that there is a lot of training material on the internet with Claude as the AI.
Could be on purpose to disguise as a US made model
(comment deleted)
The meme/trope of China copying everything really keeps playing into itself
Models don't know who they are. Stop asking this question to the model as a source of truth.
I have used Telnyx for, on the opposite end of cool-ness, their fax API. Curious if their AI pricing and quality are actually competitive or if this is just a rapid pivot to try to ride the AI wave.
That is not going to be cheap for long if K3 prices does the same as GLM5.2 prices. If nothing else open weight models are great to get providers to compeete on price.
Is it also offered via a ZDR + BAA / HIPAA eligible? Would love to use it in prod
Is this an ad?

Why is this provider specifically on the front page?

Also why is this not available over openrouter?

Token Type Price per 1M tokens Cached Input $0.27 Input $2.70 Output $13.50

Looks like they are first vendor to undercut in price.