1 comment

[ 4.2 ms ] story [ 11.7 ms ] thread
Try it out and see how fast inference can be for agentic workflows. Works on CUDA and ROCm. Feedback appreciated!