DeepSeek's Long-Term Effect

2 points by piotrgrudzien ↗ HN
Unlike other sensational tech news, I believe DeepSeek has shifted mindsets in the mainstream and will have a long-lasting effect on the broader industry. Below is my 7-point summary:

1/ "Training AI models can be cheap" Details aside, the idea that training AI models doesn't have to be that expensive has hit the mainstream. Raising huge funding for training large models will become much more difficult.

2/ More scrutiny of large LLM vendors' energy spend Environmental footprint of Generative AI models has been a topic of discussions for some time now. The news of an unknown startup demonstrating that great results can be achieved an order of magnitude cheaper can only enliven them.

3/ More scrutiny of national initiatives to train large models The need for huge investments into national / non-commercial initiatives will be put into question.

4/ "Open source is on the right side of history" DeepSeek's success might, in the eyes of many, be the final blow to closed-source as the preferred AI strategy. Sam Altman himself stated that regarding Open Source they might have been "on the wrong side of history" (https://www.vice.com/en/article/openai-ceo-sam-altman-says-theyve-been-on-the-wrong-side-of-history/).

5/ Push towards revenue-generating applications More difficulty in seeking funding will put pressure on big AI apps to move to revenue-generating use cases like e-commerce, just-in-time research or government contracts (https://www.youtube.com/watch?v=YkCDVn3_wiw, https://www.perplexity.ai/shopping, https://openai.com/global-affairs/introducing-chatgpt-gov/).

6/ More fragmentation in the AI market AI app market leaders will struggle to keep up the pace of technological advancements and launching new features. That opens up room for new players to launch specialized solutions and chip away at the ChatGPT consumer market. It has already been marked by Perplexity's entrance, more up-and-comers will follow.

7/ Leading models' knowledge cutoff becomes a real problem Competition is pushing AI apps to constantly launch new features and there is less focus on solving fundamental problems like the knowledge cutoff. For use cases such as e-commerce or just-in-time research being trained on data up until October, 2023 is a big issue. Broad availability of the web search feature is supposed to be the solution but it is silently ignores the original reason why we loved AI chat apps in the first place - their breadth of knowledge and understanding. To use top web search results as the sole basis for AI app's answers deprives the user of that depth and feels like we're back to using Google search.

Original post: https://www.linkedin.com/feed/update/urn:li:activity:7293161576402939904/

4 comments

[ 4.4 ms ] story [ 24.3 ms ] thread
I'd say the biggest conclusion to be drawn from all what has happened is "there is no moat".

Which, in turn, means that there is no first-mover advantage. Which, in turn, invalidates a ton of the ridiculous valuations by VCs like YC.

Or that moats exist but in the old-fashioned way: brand, user experience, great understanding of users, highly optimised sales tactics. In short: no free lunch :)
No technical moat, no money moat. Anybody can get to SOTA AI fastly, with reasonable money (no U$S 500B datacenters, no 5 to 10B cash in the bank required), just like anybody could develop the next TikTok or Instagram.

The most important thing is that new hardware is most probably demonstrated now NOT to be able to add any significant moat for future frontier models. If you own or have access to relatively old GPU datacenter level hardware (like the ones offered in many public clouds and/or private offerings -think old crypto farms pivoting - everywhere), you should be good provided to develop frontier models in a snap (2-3 months), from the scratch.

There's quite new thing, distillitation of models that can be done using R1, and now you could train a relatively week model -not yet frontier - but you could upgrade the thing right to the SOTA level - right now o3 probably - just distilling it with a R1 reasoner. This changes the game because you already have lots of advanced publicly downloadable models, plus whatever you can train, then you can now just begin to distill stuff seriously improving intelligence in those, it is so new that we have yet to see how it develops in the next weeks, months.