Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
Abstract
Atelier improves artist-grounded image generation by translating vague artistic intent into explicit control states that separate scene content from style, reducing reliance on stereotypical shortcuts.
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.
Community
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text-to-Image Generation
๐ arXiv: https://arxiv.org/abs/2608.06751
TL;DR: Naming an artist in a prompt isn't style control. T2I models tend to fall back on canonical shortcuts โ the recurring motifs, stock palettes, and over-represented period signatures they associate with that name โ and quietly overwrite the scene the user actually asked for. This paper introduces Atelier, which plans an explicit control state before generation instead of hoping the backend infers intent.
What Atelier does
- Decomposes an underspecified request ("a quiet subway platform in the style of Van Gogh") into an explicit control state: scene anchors, preserve vs. transform decisions, style-regime hypotheses, role-bound artist evidence, and anti-shortcut constraints.
- Grounds that state with artist-level knowledge plus local patch references โ and unlike generic RAG prompting, binds each patch to a specific scene role rather than dumping it in as an undifferentiated style exemplar.
- Compiles backend-aware generation plans, then iteratively refines candidates with global + local authenticity critics.
ArtIntentBench
A companion benchmark that trades artist breadth for supervision depth โ two deliberately contrastive artists (Van Gogh and Qi Baishi, different media and cultural traditions) across four tasks: artwork re-rendering, period-controlled generation (Paris / Arles / Saint-Rรฉmy / Auvers), historically unseen subjects, and Qi Baishi re-rendering, plus shortcut auditing and human preference eval. The argument for going deep: a wide benchmark of many artist names with shallow labels would end up measuring generic style association, not artist-grounded control.
Results
Across both open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure better, and substantially cuts shortcut substitution relative to prompt-engineered, retrieval-augmented, and general-purpose agent baselines.
Why it's interesting: the framing that artist-grounded generation is bottlenecked upstream of the diffusion model โ in inferring explicit, evidence-grounded controls โ is a nice counterpoint to the "just scale the generator" default. The shortcut-auditing setup also seems reusable well beyond these two artists.
Curious how the control state holds up on artists with less catalogued, more diffuse oeuvres, and whether the period-regime hypotheses transfer to movements rather than individuals.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling (2026)
- GMO-E$^2$DIT: Grounded Multi-Operation Editing for E-Commerce Images (2026)
- RAVA: Retrieval-Augmented Viewpoint Alignment for Subject-Driven Image Generation (2026)
- Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis (2026)
- FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining (2026)
- ArtMine: Discovering and Formalizing Artistic Processes (2026)
- PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.06751 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper