Procedural Graphs: Self-Evolving Execution Structures for LLM AgentsExplained for Beginners
Yuxing Lu, Yicheng Chen, Shanchan Wu +1 more
Abstract
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order, and under which conditions. As trajectories lengthen, agents can lose track of their objectives, invoke tools out of order, and repeat unproductive actions. We introduce the Procedural Graph: just as a knowledge graph organizes factual knowledge into (entity, relation, entity) triplets for what-is questions, a Procedural Graph organizes procedural knowledge into (procedure, relation, procedure) triplets for what-to-do questions. At each decision step, the framework localizes the agent's active node, and a guidance model translates the surrounding subgraph into step-level situational guidance that biases the solver's next action without dictating it. The graph is self-evolving: an LLM refiner contrasts failed trajectories with successful ones and edits the graph's topology and attributes, committing edits that preserve or improve held-out validation performance while retaining rejected ones to discourage repetition. Starting from a minimal skeleton, the loop builds graphs that match or surpass hand-designed ones. It can also repair a flawed expert prior. Across multiple datasets, task types, and LLMs, the Procedural Graph delivers consistent gains over memory-based baselines, and self-evolution further improves performance without manual engineering.
1. The Problem
Imagine you are trying to assemble a piece of furniture using only a verbal description of the steps, without a diagram or a parts list. You might start by attaching a leg, then realize you need a washer you didn't know existed, and eventually end up with wobbly furniture because you tightened the bolts in the wrong order. For large language models (LLMs) acting as agents—think of them as the "person" trying to assemble the furniture—the situation is surprisingly similar.
Current LLM agents typically operate on an "accumulating history" model. They read a prompt, take an action (like clicking a button or calling a API), and then append that action to their memory. Their "brain" is essentially a ever-growing transcript of and full combined taking and.. and in short-term memory. As the agent works toward a goal over a long horizon—say, booking a flight or debugging a piece of code—the accumulated history balloons in size. The model has to mentally parse this growing transcript to figure out where it left off and what to do next.
This approach has three major flaws. First, the agent can lose track of its original objective as the context window fills up with irrelevant tool outputs. Second, it can invoke tools out of order or repeat actions that previously failed, essentially taking a wrong turn and then walking back down that same wrong road. Third, there is an implicit, invisible map of "procedural knowledge." The model knows what to do to some extent, but the map of what to do next, and in what order, is implicit in the weights of the model rather than explicit in the world. This makes the agent prone to "getting lost" or repeating unproductive loops, especially as tasks become longer and more complex.
The paper identifies this gap: current LLM agents have the raw materials for knowledge (the tools and facts), but they lack an explicit "map" of the
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →