1.8-3.3x faster Embedding finetuning now in Unsloth (unsloth.ai) 3 points by electroglyph 7mo ago ↗ HN
[–] storystarling 7mo ago ↗ Do the memory savings carry over to inference or is this strictly optimizing the backward pass? I'm running embedding pipelines via Celery and being able to squeeze this into lower VRAM would help the margins quite a bit.
3 comments
[ 5.8 ms ] story [ 23.3 ms ] thread