IMO this is because responding within a json response with all those other fields shifts the distribution to text that's staler. If you told it to roleplay a pixel knight fighting to the death and provided it responses…
I mean can't you just directly test some variations of the query against many models at once and pick the cheapest model that hits your accuracy goal? If it's a one time run then ofc this is all pointless, but for…
[dead]
IMO this is because responding within a json response with all those other fields shifts the distribution to text that's staler. If you told it to roleplay a pixel knight fighting to the death and provided it responses…
I mean can't you just directly test some variations of the query against many models at once and pick the cheapest model that hits your accuracy goal? If it's a one time run then ofc this is all pointless, but for…
[dead]