Skip to main content
Arka supports vision tasks: photo description, blueprint analysis, and screen capture.

Describe images

Two-layer analysis: OCR extracts exact text; vision describes layout and colors.
Backends (auto-selected): Gemini, Ollama (llava), vLLM.

Drawings and blueprints

Analyze floor plans, elevations, MEP schematics, and scanned contracts with Gemini vision:

Screen capture

10-second countdown, full-display capture, vision describe:

Describe videos

describe_video samples frames with ffmpeg and sends those frames through the configured vision backend. Use it for gameplay recordings, UI animation checks, demos, and people-location questions.
For people-focused prompts, Arka asks the vision model to identify visible people and approximate their screen positions. For general describe/analyze prompts, it describes subjects, actions, setting, text, framing, and visual issues.