The model finds the structure. It is not trusted with the numbers.
Paste a menu — or photograph one — and Counter turns it into categories, items, descriptions and options. That part works well: layout is exactly the sort of thing a language model is good at.
Prices are a different problem. Menus print them in a right-hand column, far from the item, and every vision model I tested on this stack either skipped them or invented plausible round numbers. So Counter does not take the model's word for it: every price is checked against the transcription, and anything that is not literally printed on the menu comes back marked price not read for you to type in.
Being told “I could not read this one” is worth more than a confident wrong number — especially when the number is what somebody pays.