Conjecture: Does GLM only handle the prefill stage, compress its output hidden vectors (trained?), and then send them to Qwen for decoding?
Conjecture: Does GLM only handle the prefill stage, compress its output hidden vectors (trained?), and then send them to Qwen for decoding?