Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Abstract
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- InstanceControl: Controllable Complex Image Generation without Instance Labeling (2026)
- Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation (2026)
- Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping (2026)
- UniVerse: A Unified Modulation Framework for Segmentation-Free,Disentangled Multi-Concept Personalization (2026)
- LCG: Long-Context Consistent Image Generation with Sparse Relational Attention (2026)
- BindEdit: Taming Attention Leakage for Precise Multi-Object Image Editing (2026)
- Decoupled Guidance: Disentangling Subject and Context Pathways in Text-to-Image Personalization (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.19344 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper