13 comments

[ 0.95 ms ] story [ 3.6 ms ] thread
The fact that even the abstract is very clearly AI-generated does not fill me with confidence in the rigour of the research.
I had to use an AI to undertand the abstract made by AI…
Interesting. My initial thought was “this is so horribly written it must be human”
"Verify Presence, Not Absence" is very LLM-coded all on its own.
The entire paper is almost certainly Claude-generated, given the clear Claude-speak throughout. It seems like they didn’t even try to have Claude write a polished paper: it reads like Claude writing up a lengthy report, down to the needless sectioning with idiosyncratic title language. There appears to be no disclosure, and the authors contribution section appears to falsely claim that a particular author wrote the text.

I know arXiv has taken some measures to combat spam like this, but it seems like they’ll need to do more. There’s just very little barrier now to creating giant slop papers like this and then dumping them anywhere that won’t reject them. It is an insult to everyone’s time, and I can’t imagine they expect people to actually read this. If the expectation is that everyone will use an LLM to interpret it, then maybe they should have at least had a few more rounds of tightening and polishing the paper, even via LLM, to save the redundant token use.

May be AI written. Used my AI to get the gist of it. One thing that I believe is LLMs need to be considered a pure play tool and can be guided to point out misses as well. The reason is simple. To find whats missing, one needs to ask the right questions on how to judge and when that is there.
Not sure about the paper but the results make sense, we see it in PDF extraction too. Fields that aren't in the document are being made up 11 to 40% of the time depending on the API.

The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.

> Fields that aren't in the document are being made up 11 to 40% of the time depending on the API.

Do you have an example?

I'm in healthcare and ambient documentation is obviously a huge thing now but I don't have any experience with it. We have anywhere from 5-10 companies reach out a week trying to sell us on their product and the demos are mostly okay (though you can tell they're rely on the happy path through a lot of it), but we haven't actually pulled the trigger on anything. Thanks in advance.

"the floored note before its clean twin"

People using AI like this should be run out of their jobs.