1 comment

[ 0.21 ms ] story [ 14.1 ms ] thread
VLM can already process both the document images and the query to produce an answer directly. Do we still need the intermediate OCR step?