5 comments

[ 2.8 ms ] story [ 24.2 ms ] thread
Fascinating how Jeremy openly and honestly says that a technique he developed and everyone uses is now not the best way to go.

I have become obsessed with smaller LLMs lately. Every time I use mistral-7B I wonder what exactly the developers did - anyone have links to good papers by the Mistral developers?

I think the link to their paper was posted on HN. Try algolia to find it.
thanks for submitting this! unfortunately HN doesnt seem to like podcasts :/