Agentic Visual Generation: From Generative Models to Agentic ControlExplained for Beginners
Yinming Huang, Shuyuan Tu, Xi Yan +7 more
Abstract
Visual generation is evolving from generative models used through a single invocation into agentic control processes that can plan, select tools, inspect intermediate synthesized outputs, revise failures, and reuse prior experience. In most existing systems, the controller is an LLM or VLM, while visual generation models serve as tools or executors. However, existing work lacks a consistent criterion for determining when a generation system becomes agentic. Planning depth, tool use, multi-role collaboration, and reinforcement learning are often treated as evidence of agenticity, even though none of them necessarily determines which generation decisions the controller can make. We organize the field according to what the controller can directly control in the generation process. At L1 Conditioning Control, the controller prepares the input to a predetermined generator but does not control which visual operation is executed. At L2 Execution Control, it selects and invokes actual generation, editing, rendering, or other content-modifying operations. At L3 Outcome-Adaptive Control, it observes an intermediate outcome and uses that observation to change a subsequent operation within the current task. At L4 Experience-Adaptive Control, it retains experience from completed tasks and uses that experience to change decisions on future tasks. L0 Fixed Support separately denotes generators, editors, evaluators, reward models, benchmarks, and fixed pipelines without a deployed controller that makes generation-level decisions. These levels describe a progressively broader decision-making scope rather than model size, system complexity, output quality, tool or role count, or training method. Applying this framework across image, video, editing, 3D, world, slide, and user-interface generation reveals how controller capabilities have evolved and how their mechanisms are distributed across levels.
1. The Problem
When you ask a modern image generator to create a picture, you are interacting with a generative model—a statistical engine trained on billions of images to translate text into pixels. For years, this was a one-shot affair: you type a prompt, you get a result. But recently, the frontier of AI has shifted. We are moving from "generative models" to "agentic" systems. The promise is that an AI won't just give you a static image; it will actively manage the creative process, deciding which tool to use, checking its own work, and iterating until the result matches your intent.
However, the research community has hit a snag. As this paper argues, there is currently no consistent "ruler" to measure how agentic a system actually is. Researchers often throw around terms like "planning," "tool use," and "reinforcement learning" as if they are the same thing. But a system can use a tool without being an agent, or plan a course of action without being able to change its mind when things go wrong.
The paper identifies a core frustration: existing work treats evidence of agenticity (like having a planner) as the defining feature, even though these elements don't necessarily dictate what the controller can actually do during generation. The authors propose that to truly understand these systems, we need to look at decision-making scope. Not how big the model is, but what levers the controller actually pulls.
2. How It Works (The Technical Mechanics)
The paper’s central contribution is a framework that organizes visual generation systems into five distinct levels of controller capability. Think of this not as a ranking of "better" or "worse," but as a map of what the system is actually allowed to control during the creative process.
L0: Fixed Support This is the "traditional" generative AI. Imagine a calculator. You punch in numbers, and it gives you an answer. The calculator doesn't decide how to solve the problem; it just executes the operation you request. In AI terms, L0 denotes generators, editors, and evaluators that exist within a fixed pipeline. There is no deployed controller making generation-level decisions on the fly. You type a prompt, the model generates, and that’s it.
L1: Conditioning Control Here, the controller prepares the input, but the generator runs the show. Think of this like a photographer directing a model. You (the controller) decide on the lighting, the angle, and the subject (the "conditioning"), but the camera (the generator) decides exactly how the image is rendered. The controller sets the stage but doesn't dictate the specific visual operations happening inside the model.
L2: Execution Control This is where things get "agentic" in the way most people imagine it. At L2, the controller can select and invoke actual generation, editing, rendering, or other content-modifying operations. Using our analogy, this is like a director on a film set. The director doesn't just set the lighting; they yell "Cut!" they choose which camera angle to shoot, and they decide when to switch from wide shot to close-up. The controller is actively choosing which tool to use and when to use it.
L3: Outcome-Adaptive Control This level introduces feedback loops. At L3, the controller observes an intermediate outcome and uses that observation to change a subsequent operation within the current task. Imagine you are drawing a portrait, and you sketch the eyes, but they look a bit too wide. You don't start the whole drawing over; you just go back and adjust the eyes. The system sees the "intermediate result" (the eyes) and adapts the next step (the face shape) based on what it saw. This is the difference between a fixed plan and a reactive one.
L4: Experience-Adaptive Control This is the highest level of autonomy described in the paper. At L4, the controller retains experience from completed tasks and uses that experience to change decisions on future tasks. This is essentially "learning how to do the task" over time. If the system realizes that "using a specific editing tool first usually leads to better results," it will start doing that automatically on new, unrelated prompts. It transfers knowledge across tasks, much like a human artist who realizes that starting with a sketch layer always improves their final paintings.
The paper emphasizes that these levels describe decision-making scope, not model size. A tiny model could theoretically operate at L4 if it has the right control logic, while a massive model might be stuck at L0 if it's just responding to fixed prompts.
3. Key Results & Benchmarks
The paper doesn't just propose a theory; it applies this framework across a wide range of visual modalities—image, video, editing, 3D, world models, slides, and user interfaces. By doing so, it reveals how controller capabilities have evolved and how mechanisms are distributed across these levels.
While the paper is primarily conceptual and organizational, its "results" are the insights gained from mapping existing systems onto these levels. For example, the framework helps explain why some recent "agentic" systems feel more capable than others: they might be using L2 tools (selecting operations) but lack the adaptive feedback of L3 (correcting course mid-task).
The paper implies that as we move up the levels, we should expect to see systems that are more robust to user errors and better at complex, multi-step creative tasks. The "benchmark" here is the framework itself—a way to compare the "agency" of different systems apples-to-apples.
4. Why It Matters (Key Takeaways)
- A Common Language for "Agency": The most immediate takeaway is that the field now has a structured way to talk about agentic AI. Instead of debating whether a system "is" or "isn't" agentic, researchers can now say, "This system operates at L2, so it can select tools, but it can't adapt its plan when it makes a mistake." This clarity is crucial for developers and users alike.
- The Shift from Prompting to Orchestration: For product managers and users, this framework signals a shift. We are moving away from the era of "crafting the perfect prompt" (which works great for L0/L1 systems) toward an era of "orchestrating" a process. In L2 and beyond, the AI becomes a partner that manages the workflow, choosing tools and iterating, rather than just a fancy paintbrush.
- Limitations to Watch: Systems higher up the ladder (L3 and L4) are likely more complex to build and potentially harder to predict. An L4 system that "learns from experience" might develop quirks or biases based on its specific history of tasks. Users should be aware that these systems might behave unpredictably if they haven't been deployed in the specific context the system "learned" from.
- What to Watch Next: Keep an eye on the transition from L3 to L4. The ability to retain experience across tasks is the holy grail for making AI assistants feel truly "intelligent" and personalized. The next big leap will likely be in how these systems manage long-term memory without losing the ability to handle new, unique creative requests.
In summary, "Agentic Visual Generation" provides a much-needed map of the landscape. It moves the discussion from vague claims of "intelligence" to a concrete taxonomy of control, helping us understand exactly how much autonomy an AI has—and how much control we, as users, should expect to retain.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →