kisjovan

↗ HN profile [ 33.9 ms ] full profile
Karma
0
Created
()
Submissions
0
  1. Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with…