PitVQA: Surgical Vision-Language Model

Multi-task VLM for pituitary surgery understanding

This demo showcases verified examples that produce valid outputs. Select an example type to load a pre-tested surgical frame.

Load Verified Example

Example Type

Stage 1: Locate instruments and anatomy

Target

Models: pitvqa-qwen2vl-unified-v2 | Dataset: pitvqa-comprehensive-spatial

Output formats: <point x='45' y='68'>target</point> | <box x1='20' y1='30' x2='60' y2='70'>target</box>