12 publications

Research

Two threads: getting vision–language models to work in a domain as demanding as art, and working out how they arrive at what they say.

2026

Figure from “Understanding How MLLMs Describe Artworks Using Token Activation Maps”

2026PRESTIGE Workshop at the International Conference on Pattern Recognition (ICPR)

Understanding How MLLMs Describe Artworks Using Token Activation Maps

Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi, Eva Cetinić, Gennaro Vessio, Giovanna Castellano

We ask what multimodal LLMs are actually looking at when they describe a painting. Using Token Activation Maps, we produce a heatmap for every generated token that isolates the visual evidence specific to that token, then analyse five categories across periods and genres: common objects, style descriptors, metadata, iconographic tokens and affective expressions. Grounding varies sharply by token type — models attribute artists far more reliably than they predict titles, where hallucination concentrates.

Figure from “Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment”

2026arXiv preprint

Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment

Ivan Rinaldi, Matteo Mendula, Nicola Fanelli, Florence Levé, Matteo Testi, Giovanna Castellano, Gennaro Vessio

Image-conditioned music systems are trained on natural photographs, which leaves them unable to capture the semantic, stylistic and cultural content of artworks. We introduce ArtSound, a dataset of 105,884 artwork–music pairs with dual-modality captions, and ArtToMus, a framework that maps digitized artworks directly to music by projecting visual embeddings into the conditioning space of a latent diffusion model — no image-to-text step in between.

2025

Figure from “ArtSeek: Deep Artwork Understanding via Multimodal In-Context Reasoning and Late Interaction Retrieval”

2025arXiv preprint

ArtSeek: Deep Artwork Understanding via Multimodal In-Context Reasoning and Late Interaction Retrieval

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

A multimodal pipeline for captioning and visual question answering in the art domain. ArtSeek extends ColQwen2 for multimodal retrieval, introduces WikiFragments — a dataset of multimodal fragments mined from Wikipedia — and a multitask classification framework, unlocking agentic retrieval-augmented generation for multimodal LLMs through in-context learning.

Figure from “Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual Prompts”

202528th European Conference on Artificial Intelligence (ECAI)

Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual Prompts

Pasquale De Marinis, Nicola Fanelli, Raffaele Scaringi, Emanuele Colonna, Giuseppe Fiameni, Gennaro Vessio, Giovanna Castellano

A neural architecture for few-shot semantic segmentation that supports multi-class segmentation from points, boxes or masks as prompts, relaxing several of the constraints that normally govern how support sets have to be built.

Figure from “I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting”

2025IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

Best Paper Honorable Mention — IEEE CIS Italy Chapter, IJCNN 2025

We inpaint several regions of a painting at once, each driven by its own text prompt, and train a multimodal LLM to propose those prompts automatically — so the model suggests what could plausibly fill each gap rather than requiring a human to describe every mask.

2025International Conference on Image Analysis and Processing (ICIAP)

Unveiling Visual Features in Artwork Classification: Towards Explainable Vision Transformers in the Arts

Raffaele Scaringi, Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

An investigation into which visual features vision transformers actually rely on when classifying artworks, and how to surface them in a form an art historian can interrogate.

2025IEEE 35th International Workshop on Machine Learning for Signal Processing (MLSP)

Multimodal Artwork Topic Modeling via Fine-Tuned CLIP and Knowledge-Driven Prompts

Raffaele Scaringi, Giacomo Stea, Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

Topic modelling over artwork collections that combines a fine-tuned CLIP with knowledge-driven prompts, so the discovered topics line up with art-historical vocabulary instead of arbitrary clusters.

2025Ital-IAabstract

Generative AI Across Modalities: Insights from Our Research on Domain-Aware Content Generation

Giovanna Castellano, Emanuele Colonna, Nicola Fanelli, Luca Laraspata, Ivan Rinaldi, Andrea Gabriele Valerio, Gennaro Vessio

An overview of the lab’s work on domain-aware generative models spanning images, music and text.

2024

Figure from “Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation”

2024AI4VA Workshop at the European Conference on Computer Vision (ECCV)

Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation

Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano, Gennaro Vessio

We extend the AudioLDM2 architecture to generate music directly from artworks, trained on a dataset of image–music pairings collected with ImageBind.

Figure from “Converso: Improving LLM Chatbot Interfaces and Task Execution via Conversational Forms”

2024LUHME Workshop at the European Conference on Artificial Intelligence (ECAI)

Converso: Improving LLM Chatbot Interfaces and Task Execution via Conversational Forms

Gianfranco Demarco, Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

A fully containerized architecture for building LLM chatbots, with conversational forms that measurably improve how reliably they collect structured data from a user.

2023

Figure from “Exploring the Synergy Between Vision-Language Pretraining and ChatGPT for Artwork Captioning: A Preliminary Study”

2023FAPER Workshop at the International Conference on Image Analysis and Processing (ICIAP)

Exploring the Synergy Between Vision-Language Pretraining and ChatGPT for Artwork Captioning: A Preliminary Study

Giovanna Castellano, Nicola Fanelli, Raffaele Scaringi, Gennaro Vessio

Caption generation for digitized artworks from a noisy dataset of LLM-written descriptions. We introduce CLIPScore weighting, which weighs each caption by its estimated quality during training and recovers much of the performance lost to the noise.

2023Lecture Notes in Computer Scienceabstract

Exploring New Frontiers at the Intersection of AI and Art

Giovanna Castellano, Nicola Fanelli, Raffaele Scaringi, Gennaro Vessio

A short position piece on where computer vision and art history usefully meet.