OpenAI Models Dominate Structured Code Edit Benchmark (blog.mentat.ai) 5 points by biobootloader 2y ago ↗ HN
[–] granawkins 2y ago ↗ I saw the same thing with Mailogy.I only tested variants of gpt-3.5 and -4 but got ~50% invalid syntax errors with 3.5, and virtually none with 4.
1 comment
[ 4.2 ms ] story [ 87.2 ms ] threadI only tested variants of gpt-3.5 and -4 but got ~50% invalid syntax errors with 3.5, and virtually none with 4.