Show HN: How do you know your LLM giving accurate info? (github.com)

1 points by bhanuhai2 ↗ HN
I spend a good amount of time chasing the LLM for evidence that its answer is actually accurate. To be honest, I can usually get there, but it works great for fresh work and gets messy on months-old work. And it's not just me; a lot of people have raised the same thing.

So for the last 6 months I've been trying to fix it. Today I can say it's working for me — it flags me in advance. Obviously not perfect yet, and that's where I need help: I need people to test it and tell me what else needs fixing.

PS: This isn't a memory layer; it flags stale answers before you act on them. Give it a try, and if you think I'm solving a real problem, please drop a star.

1 comment

[ 0.30 ms ] story [ 5.8 ms ] thread
I never tested this with science folks, would love if someone share some use case