Actually there is even a straight connection: Step-Fun DeepResearch trained on SYNTH (the open Baguettotron dataset).
Hi. Yes this is wholly correct. On the second points: * Well I'm very much involved in making open more models, pretrained the first model on free and open data without copyrigh issues, released the first version fo…
So to clarify: the important product that people will ultimately want is the model. Obviously you need to design an infra/UI around it but that's not the core product. The really important distinction is between…
Hi. So quickly: * RL is Reinforcement Learning. Already used for a while as part of RLHF but now we have started to find a very nice combo of reasoning+RL on verifiable tasks. Core idea is that models are not just good…
Hi, author here. An important background is the imminent rise of actual LLM agents I discuss in the next post: https://vintagedata.org/blog/posts/designing-llm-agents So answering to a few comments: *The shift is coming…
Ah it's completely volontary on my part: I want to keep the historical spelling as much a possible. That's why I used the google books OCR which does a better work at it than Gallica. That's still a bit erased in the…
Either that or appending archaic expressions in the prompts (a bit like the prompt extension of Midjourney)
Well you're not going to believe it but I do have a FoucaultGPT just being trained (indirectly: Foucault is just part of my extended French historical corpus). As a sample: Prompt : Écrit un livre de Michel Foucault sur…
I published the completely dataset here: https://huggingface.co/datasets/Pclanglais/MonadGPT While I don't think Saint-Simon is included, a French colleague did a few try with it that turned out better than ChatGPT. I'm…
In a way you could do so by prompting Monad with artificial intelligence stuff. I did a try lately on the latest OpenAI events and it went on like this: "In this sad and tragical storye, you shall heare how Sam Altman,…
Yes I needed that for the conversational/instructional capacities. I've made a lot of tests with base models and it would not listen to instruction very well…
Yes you're perfectly right. I've currently tried to maintain some kind of uneasy balance between good conversational capacities (so that it really is a "chatGPT") and cultural reset, which means it may revert from its…
Yes. I think we may have enough for "full finetuning" and erasing to a large extent the previous knowledge. But that's still very far off for pretraining. "RomeGPT" is next on my list of Monad successors and to give you…
Yes it happens once in a while. It's still a small model (7B) and I've done very weird things with it. If I were historically reconditioned in the 17th century mindset, I would also likely have strange lapses of…
Already in the work. Just had a meeting today with two latinists about it.
Certainly. In fact I see we already follow each other on Twitter :D And yes totally. The other massive impact could be in source analysis. I have started using Mistral-Hermes for text annotation and it is both…
Hi! Model creator here. I happen to be a cultural historian and that's a main use case that I see. It's not complicated to learn about past events but having a general idea of the culture of the time (and its alieness…
Hi. TheBloke has quantized the model: https://huggingface.co/TheBloke/MonadGPT-GGUF You may be able to run the Q3 or Q4 variant. Although in my experience, the quality of quantization takes a hit on "weirder" data…
Actually there is even a straight connection: Step-Fun DeepResearch trained on SYNTH (the open Baguettotron dataset).
Hi. Yes this is wholly correct. On the second points: * Well I'm very much involved in making open more models, pretrained the first model on free and open data without copyrigh issues, released the first version fo…
So to clarify: the important product that people will ultimately want is the model. Obviously you need to design an infra/UI around it but that's not the core product. The really important distinction is between…
Hi. So quickly: * RL is Reinforcement Learning. Already used for a while as part of RLHF but now we have started to find a very nice combo of reasoning+RL on verifiable tasks. Core idea is that models are not just good…
Hi, author here. An important background is the imminent rise of actual LLM agents I discuss in the next post: https://vintagedata.org/blog/posts/designing-llm-agents So answering to a few comments: *The shift is coming…
Ah it's completely volontary on my part: I want to keep the historical spelling as much a possible. That's why I used the google books OCR which does a better work at it than Gallica. That's still a bit erased in the…
Either that or appending archaic expressions in the prompts (a bit like the prompt extension of Midjourney)
Well you're not going to believe it but I do have a FoucaultGPT just being trained (indirectly: Foucault is just part of my extended French historical corpus). As a sample: Prompt : Écrit un livre de Michel Foucault sur…
I published the completely dataset here: https://huggingface.co/datasets/Pclanglais/MonadGPT While I don't think Saint-Simon is included, a French colleague did a few try with it that turned out better than ChatGPT. I'm…
In a way you could do so by prompting Monad with artificial intelligence stuff. I did a try lately on the latest OpenAI events and it went on like this: "In this sad and tragical storye, you shall heare how Sam Altman,…
Yes I needed that for the conversational/instructional capacities. I've made a lot of tests with base models and it would not listen to instruction very well…
Yes you're perfectly right. I've currently tried to maintain some kind of uneasy balance between good conversational capacities (so that it really is a "chatGPT") and cultural reset, which means it may revert from its…
Yes. I think we may have enough for "full finetuning" and erasing to a large extent the previous knowledge. But that's still very far off for pretraining. "RomeGPT" is next on my list of Monad successors and to give you…
Yes it happens once in a while. It's still a small model (7B) and I've done very weird things with it. If I were historically reconditioned in the 17th century mindset, I would also likely have strange lapses of…
Already in the work. Just had a meeting today with two latinists about it.
Certainly. In fact I see we already follow each other on Twitter :D And yes totally. The other massive impact could be in source analysis. I have started using Mistral-Hermes for text annotation and it is both…
Hi! Model creator here. I happen to be a cultural historian and that's a main use case that I see. It's not complicated to learn about past events but having a general idea of the culture of the time (and its alieness…
Hi. TheBloke has quantized the model: https://huggingface.co/TheBloke/MonadGPT-GGUF You may be able to run the Q3 or Q4 variant. Although in my experience, the quality of quantization takes a hit on "weirder" data…