Interfaze

logo

Beta

pricing

help

docs

blog

sign in

All leaderboards

Object detection (NL prompts)

RefCOCO

Visual grounding: given a free-form natural-language description, the model must return the exact bounding box of the object referenced — not classify it, but locate it.

Acc@0.5 — fraction of predicted bounding boxes whose IoU with the ground truth is ≥ 0.5. Higher is better. Hover a bar to reveal the exact score.

Model rankings

Scores

Every model evaluated on RefCOCO, ranked highest to lowest.

#ModelScore
1Interfaze82.1%
2Gemini-3.5-Flash80.9%
3Claude-Sonnet-4.675.5%
4Gemini-3-Flash75.2%
5Claude-Sonnet-569.2%
6GPT-5.4-Mini67.0%
7Grok-4.325.0%

Interfaze

logo

Product

Playground

OCR

Models

Leaderboards

Pricing