Agent skill · creative production · aaaaqwq

vision-analyze

Image analysis using multimodal vision models. Use when user needs to: (1) Describe what's in an image, (2) Extract text from images (OCR), (3) Analyze visual content, (4) Compare images, (5) Answer questions about images. Supports JPG, PNG, GIF, WebP formats.

Why this skill is useful

Provides domain-specific commands for image analysis and extraction that the AI wouldn't reliably generate on its own.

What it needs

About 1k tokens when loaded. Last updated 2026-08-06. 83 stars on the source repository.

What this skill does

Vision Analyze Analyze images using the built-in vision capabilities of multimodal AI models. Quick Start Analyze an Image Describe what's in an image: Extract Text (OCR) Extract text from images: Analyze Multiple Images Compare or analyze multiple images: Usage Patterns Visual Q&A Ask specific questions about image content: Content Moderation Check image content: Data Extraction Extract structured data from visual content: Visual Comparison Compare images: Tips Be specific: The more specific your prompt, the better the results Multiple images: You can analyze up to 20 images at once Supported formats: JPG, PNG, GIF, WebP Size limits: Large images are automatically resized When to Use Reading text from screenshots, documents, or photos Describing visual content for accessibility Analyzing charts, graphs, or diagrams Comparing visual changes Extracting data from forms or receipts Understanding UI elements or error messages

How to use it

Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:

@skills aaaaqwq/image-vision

View the source on GitHub

Browse the @skills marketplace