2 comments

[ 3.3 ms ] story [ 12.5 ms ] thread
Conjecture: Does GLM only handle the prefill stage, compress its output hidden vectors (trained?), and then send them to Qwen for decoding?