From what I'm seeing elsewhere, context size up to 128k should be possible on this hardware. It really matters for agentic workloads to push that context size headroom up. Anthropic are spoiling us with models that do…
It gave the card an aura of mystique (pun intended) and I remember really wanting it. I guess that means they succeeded in making a statement with it.
Nvidia is also pushing for local inference IMHO. They want open models and competition in the model layer, not two big labs controlling all of it.
This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped…
By leveraging the experience Broadcom has in this area. Still remains to be seen how that goes when they want to scale production.
Most of the generative art is hosted on IPFS.
Wait, it's allowed again? Completely missed that. Been avoiding to use it and trying to find workarounds, not great.
The Cerebras hardware is not locked to specific models / model families. Taalas is the company that's etching models into their silicon, locking it to that model forever.
My default assumption (with no knowledge about this particular case) is that all gear in YouTube videos is sponsored. Too easy to get jealous of all the fancy setups. But I don't want to become a YouTuber so that…
OpenAI's upcoming mega IPO
> Time for me to go research the early history of electrification. The Stepchange podcast has an amazing episode on The Grid [1], walking us through the arc of history of how it became the utility it is today. [1]:…
> posting this sentence was part of the deal to get the compute This 100%
> but given that you can actually run the models yourself on AWS Bedrock That's not exactly how it works. Anthropic are hosting their models in AWS Bedrock as a managed service. Customers call those LLMs just like…
There is a surprising amount of code needed in each of the inference frameworks (LM Studio, llama.cpp, etc) to support each new model release. For example to format the input in the right way using a chat template, to…
I think we lost that terminology war. Open source models mean open weight. There are only a couple examples of fully open source models with open data and code, and the labs are not incentivized to go that far.
Yes, retooling gas stations is the way to go. Already happening in Norway where stations now show the price of kWh in addition to gas and diesel prominently on signs by the road. Charging is just a different kind of…
America certainly did not invent electric cars. Depending on which electric car you consider the first real one, the inventor was either French, British or German [1]. [1]:…
Unsloth is providing the best and most reliable libraries for finetuning LLMs. We've used it for production use-cases where I work, definitely solid.
Spend time building a test harness and evaluations of whether the solution meets the requirements. Then you don't need to look at the code because those other pieces will bring the necessary guarantees and trust.
Because we all prefer it over Gemini and Codex. Anthropic knows that and needs to get as much out of it as possible while they can. Not saying the others will catch up soon. But at some point other models will be as…
Still do. Great for workloads where it's okay to bundle a bunch of requests and wait some hours (up to 24h, usually done faster) for all of them to complete.
Thank you, was confused there for a second XD
You’re ignoring Jevons paradox. Everyone, both people and companies, will be making exponentially more software with these tools. Software that both needs to get created, debugged and updated to realize the intention of…
Apple has a solid hardware business and massive profits from their App Store tax, they are not dependent on ad business in the way Google is. Very different incentives.
Sounds neat but what kind of range limits would that impose on each trip? Switching from one means of transportation to another, even if both are buses, increases the total travel time significantly. Not to mention all…
From what I'm seeing elsewhere, context size up to 128k should be possible on this hardware. It really matters for agentic workloads to push that context size headroom up. Anthropic are spoiling us with models that do…
It gave the card an aura of mystique (pun intended) and I remember really wanting it. I guess that means they succeeded in making a statement with it.
Nvidia is also pushing for local inference IMHO. They want open models and competition in the model layer, not two big labs controlling all of it.
This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped…
By leveraging the experience Broadcom has in this area. Still remains to be seen how that goes when they want to scale production.
Most of the generative art is hosted on IPFS.
Wait, it's allowed again? Completely missed that. Been avoiding to use it and trying to find workarounds, not great.
The Cerebras hardware is not locked to specific models / model families. Taalas is the company that's etching models into their silicon, locking it to that model forever.
My default assumption (with no knowledge about this particular case) is that all gear in YouTube videos is sponsored. Too easy to get jealous of all the fancy setups. But I don't want to become a YouTuber so that…
OpenAI's upcoming mega IPO
> Time for me to go research the early history of electrification. The Stepchange podcast has an amazing episode on The Grid [1], walking us through the arc of history of how it became the utility it is today. [1]:…
> posting this sentence was part of the deal to get the compute This 100%
> but given that you can actually run the models yourself on AWS Bedrock That's not exactly how it works. Anthropic are hosting their models in AWS Bedrock as a managed service. Customers call those LLMs just like…
There is a surprising amount of code needed in each of the inference frameworks (LM Studio, llama.cpp, etc) to support each new model release. For example to format the input in the right way using a chat template, to…
I think we lost that terminology war. Open source models mean open weight. There are only a couple examples of fully open source models with open data and code, and the labs are not incentivized to go that far.
Yes, retooling gas stations is the way to go. Already happening in Norway where stations now show the price of kWh in addition to gas and diesel prominently on signs by the road. Charging is just a different kind of…
America certainly did not invent electric cars. Depending on which electric car you consider the first real one, the inventor was either French, British or German [1]. [1]:…
Unsloth is providing the best and most reliable libraries for finetuning LLMs. We've used it for production use-cases where I work, definitely solid.
Spend time building a test harness and evaluations of whether the solution meets the requirements. Then you don't need to look at the code because those other pieces will bring the necessary guarantees and trust.
Because we all prefer it over Gemini and Codex. Anthropic knows that and needs to get as much out of it as possible while they can. Not saying the others will catch up soon. But at some point other models will be as…
Still do. Great for workloads where it's okay to bundle a bunch of requests and wait some hours (up to 24h, usually done faster) for all of them to complete.
Thank you, was confused there for a second XD
You’re ignoring Jevons paradox. Everyone, both people and companies, will be making exponentially more software with these tools. Software that both needs to get created, debugged and updated to realize the intention of…
Apple has a solid hardware business and massive profits from their App Store tax, they are not dependent on ad business in the way Google is. Very different incentives.
Sounds neat but what kind of range limits would that impose on each trip? Switching from one means of transportation to another, even if both are buses, increases the total travel time significantly. Not to mention all…