Rectified flow

The background is a flow matching field. Flow matching trains a velocity field to carry a simple distribution — here a Gaussian cloud of samples — onto a data distribution, which in this case is the photograph.

xt=(1t)x0+tx1,dxtdt=x1x0x_t=(1-t)\,x_0+t\,x_1,\qquad \frac{\mathrm{d}x_t}{\mathrm{d}t}=x_1-x_0

Rectified flow is the variant that couples each noise sample to its data sample along a straight line. The velocity is then constant, so the paths you see are exactly straight — and a model that learns them can be integrated in a handful of steps, rather than the many small ones a diffusion sampler's curved trajectories need.

Liu et al., 2022

Portrait of Nicola Fanelli

Zürich, Switzerland

Nicola Fanelli

I work on multimodal models for text, image and video. My PhD: getting vision–language models to hold up in a domain as demanding as art — and working out how they get there.

AI Engineer

Studio Jadu

Building multimodal text–image–video models for Studio Jadu's animation pipeline: tooling that extends what a small team can produce, with artists and storytellers directing the work at every step.

PhD candidate, Computer Science & Mathematics

University of Bari Aldo Moro

Vision–language models and multimodal AI at CILab, supervised by Prof. Giovanna Castellano and Prof. Gennaro Vessio, funded by Italy's NRRP Mission 4. Two threads run through it: getting these models to hold up in a domain as demanding as art, and working out how they arrive at what they say. Thesis submitted September 2026; defence expected December 2026.

About me

Figure from “Understanding How MLLMs Describe Artworks Using Token Activation Maps”

2026ICPR Workshops

Understanding How MLLMs Describe Artworks Using Token Activation Maps

Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi, Eva Cetinić, Gennaro Vessio, Giovanna Castellano

Figure from “ArtSeek: Deep Artwork Understanding via Multimodal In-Context Reasoning and Late Interaction Retrieval”

2025arXiv

ArtSeek: Deep Artwork Understanding via Multimodal In-Context Reasoning and Late Interaction Retrieval

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

Figure from “Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual Prompts”

2025ECAI

Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual Prompts

Pasquale De Marinis, Nicola Fanelli, Raffaele Scaringi, Emanuele Colonna, Giuseppe Fiameni, Gennaro Vessio, Giovanna Castellano

Figure from “I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting”

2025WACV

I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

Best Paper Honorable Mention — IEEE CIS Italy Chapter, IJCNN 2025

All 12 publications