
Understanding How MLLMs Describe Artworks Using Token Activation Maps
We ask what multimodal LLMs are actually looking at when they describe a painting. Using Token Activation Maps, we produce a heatmap for every generated token that isolates the visual evidence specific to that token, then analyse five categories across periods and genres: common objects, style descriptors, metadata, iconographic tokens and affective expressions. Grounding varies sharply by token type — models attribute artists far more reliably than they predict titles, where hallucination concentrates.






