AI-Generated Alt Text: What It Gets Right, and What It Doesn't

Vision models are now good enough to describe most images accurately, which makes AI-generated alt text realistic for sites with thousands of images that will never get written by hand. It is not a blind substitute for judgment. Here's where it holds up and where it doesn't.

Where it works well

CaseWhy
Product photosConcrete objects, consistent framing — a vision model describes color, type and key features reliably.
Stock and editorial photographyGeneric real-world scenes are exactly what these models are trained on.
Screenshots of UIText and layout in the image are usually read correctly, which covers most "here's how the dashboard looks" cases.
Bulk backlogsA site with 5,000 untouched images gets 5,000 reasonable alt texts today, instead of zero forever because no one will ever do it by hand.

Where it still needs a human check

CaseWhy it's risky unchecked
Charts and graphsA model can describe the shape of a chart but usually can't read exact data values reliably — these need a text summary written by someone who has the data.
Brand, people and contextA model doesn't know that the person in the photo is your CEO, or that the building is your new office — it will describe generically ("a person standing in an office") instead of meaningfully.
Decorative imagesDeciding alt="" vs. a real description is a judgment call about the image's role on the page, not something visible in the pixels alone.
Functional images (buttons, linked icons)The correct alt text describes the action, not the image — "Download the report" not "a downward arrow icon" — and a model only sees the icon, not what it links to.

The practical middle ground

Generate alt text with AI for volume, review the categories above (charts, people, decorative, functional) by hand, and ship the rest as-is. That's the approach behind both of our own tools:

← See all products