Vision models are now good enough to describe most images accurately, which makes AI-generated alt text realistic for sites with thousands of images that will never get written by hand. It is not a blind substitute for judgment. Here's where it holds up and where it doesn't.
| Case | Why |
|---|---|
| Product photos | Concrete objects, consistent framing — a vision model describes color, type and key features reliably. |
| Stock and editorial photography | Generic real-world scenes are exactly what these models are trained on. |
| Screenshots of UI | Text and layout in the image are usually read correctly, which covers most "here's how the dashboard looks" cases. |
| Bulk backlogs | A site with 5,000 untouched images gets 5,000 reasonable alt texts today, instead of zero forever because no one will ever do it by hand. |
| Case | Why it's risky unchecked |
|---|---|
| Charts and graphs | A model can describe the shape of a chart but usually can't read exact data values reliably — these need a text summary written by someone who has the data. |
| Brand, people and context | A model doesn't know that the person in the photo is your CEO, or that the building is your new office — it will describe generically ("a person standing in an office") instead of meaningfully. |
| Decorative images | Deciding alt="" vs. a real description is a judgment
call about the image's role on the page, not something visible in the pixels alone. |
| Functional images (buttons, linked icons) | The correct alt text describes the action, not the image — "Download the report" not "a downward arrow icon" — and a model only sees the icon, not what it links to. |
Generate alt text with AI for volume, review the categories above (charts, people, decorative, functional) by hand, and ship the rest as-is. That's the approach behind both of our own tools: