2 comments

[ 4.3 ms ] story [ 20.0 ms ] thread
What's the round trip latency on this? ask question -> response. Do you parse the question word by word and feed into llm or wait for the whole question before feeling into llm?