1 comment

[ 4.3 ms ] story [ 16.2 ms ] thread
Hey HN! Just sharing some work we did to make gpt-oss finetuning use O(N) and not O(N^2) VRAM via Flex Attention + some bug fixes :)