we use mailgun's email thread parser. which is not perfect but good enough, and keep only the latest email (we ignore text from older emails in thread). figuring out the output language is where a lot of the domain…
the memory blows up with the length of encoder sequence. for that reason we truncate the email at ~300 tokens, which is for the vast majority of cases enough to capture the relevant info. other than that we don't get…
marcos here (one of the authors). i know the word "breakthrough" in the title is a "little" ambitious, but i really think we've done something interesting ... we'd like to publish so this is a way to collect…
Marcos here, ds at x.ai ... happy to answer that question in person. If you are around NYC, feel free to pass by our offices for a coffee/chat :-)
we use mailgun's email thread parser. which is not perfect but good enough, and keep only the latest email (we ignore text from older emails in thread). figuring out the output language is where a lot of the domain…
the memory blows up with the length of encoder sequence. for that reason we truncate the email at ~300 tokens, which is for the vast majority of cases enough to capture the relevant info. other than that we don't get…
marcos here (one of the authors). i know the word "breakthrough" in the title is a "little" ambitious, but i really think we've done something interesting ... we'd like to publish so this is a way to collect…
Marcos here, ds at x.ai ... happy to answer that question in person. If you are around NYC, feel free to pass by our offices for a coffee/chat :-)