Multi-task VLM for pituitary surgery understanding
This demo showcases verified examples that produce valid outputs. Select an example type to load a pre-tested surgical frame.
Stage 1: Locate instruments and anatomy
Stage 2: Detect surgical objects
Stage 1: Identify surgical phase
Stage 4: Ask any question
Models: pitvqa-qwen2vl-unified-v2 | Dataset: pitvqa-comprehensive-spatial
Output formats: <point x='45' y='68'>target</point> | <box x1='20' y1='30' x2='60' y2='70'>target</box>
<point x='45' y='68'>target</point>
<box x1='20' y1='30' x2='60' y2='70'>target</box>