Abstract pixel-art graphic with purple, white, and red circular and square dots forming a bold curved digital pattern on a black background.
COLM
2024

Can MLLMs Perform Multimodal In-Context Learning for Text-to-Image Generation?

Yuchen Zeng, Wonjun Kang, Yicong Chen, Hyung Il Koo, Kangwook Lee
text-to-image
in-context-learning

Abstract

The evolution from Large Language Models (LLMs) to Multimodal Large Language Models (MLLMs) has spurred research into extending In-Context Learning (ICL) to its multimodal counterpart. Existing such studies have primarily concentrated on image-to-text ICL. However, the Text-to-Image ICL (T2I-ICL), with its unique characteristics and potential applications, remains underexplored. To address this gap, we formally define the task of T2I-ICL and present CoBSAT, the first T2I-ICL benchmark dataset, encompassing ten tasks. Utilizing our dataset to benchmark six state-of-the-art MLLMs, we uncover considerable difficulties MLLMs encounter in solving T2I-ICL. We identify the primary challenges as the inherent complexity of multimodality and image generation. To overcome these challenges, we explore strategies like fine-tuning and Chain-of-Thought prompting, demonstrating notable improvements.

​
Read paper
View Github

Related Publications

Scaling Test-Time Compute for DLLMs via Parallel Search

COLM
2026
diffusion-llm
search
parallel-decoding
​
View Job

Characterizing High Bandwidth Flash for LLM Serving

arXiv
2026
kv-cache
long-context
​
View Job

AsyncOPD: How Stale Can On-Policy Distillation Be?

NeurIPS
2026
reasoning
distillation
asynchronous-execution
​
View Job