2 comments

[ 0.23 ms ] story [ 10.4 ms ] thread
RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention