Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive OptimizationExplained for Beginners
Sihan Ge, Yichen Lin, Chenyu Zhou +3 more
Abstract
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descriptions, yet realistic operations research (OR) requests are often incomplete: missing objectives, constraints, or business rules can change the resulting mathematical program. Existing evaluations largely assume a complete specification and therefore overlook whether an agent knows when clarification is needed before modeling. We introduce OR-Clarify, a benchmark for pre-formulation clarification. Each task presents a partial public problem description, withholds structured hidden slots, and evaluates agents through bounded interaction with a simulated user. The benchmark supports both openended and choice-based clarification, and measures slot recovery, stopping behavior, silent assumptions, and interaction cost. We further propose Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop. In our choice-based experiments, InterOPT substantially outperforms all baselines in exact slot recovery; in the open-ended setting, it remains competitive with strong prior methods. Together, OR-Clarify and InterOPT reframe OR assistance as a selective completeness decision: clarify when needed, stop when ready, and quantify what remains missing.
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
1. The Problem
Imagine asking a smart colleague to plan a delivery route. They pull out a notebook and start sketching paths, only to realize later that you never said whether the trucks need to return to the warehouse. The route they drew is wrong not because they’re bad at routing, but because a single, unasked binary choice—open or closed tour—fundamentally changes the mathematical constraints of the problem.
This is the core issue that the paper “Ask Before You Optimize” addresses. Large language models (LLMs) are increasingly being used to translate natural language into optimization models—think supply chain logistics, production scheduling, or resource allocation. In theory, this lowers the barrier to Operations Research (OR), letting non-experts get mathematical models without hiring a specialist. But in practice, real-world business requests are messy. A user might describe capacities and demands without specifying the objective, or state a routing rule without clarifying whether time windows are hard constraints or soft preferences.
Most existing evaluations of LLM-based optimization agents assume the problem description is already complete. They measure whether the agent can solve a given model, but they overlook a critical earlier question: Does the agent know when it needs to ask for clarification before it starts modeling?
If an agent silently assumes a default—like “every route returns to the origin”—it might generate a mathematically valid model that solves the wrong problem. The paper identifies this “premature formulation” as the central failure mode. It argues that we need to evaluate agents not just on their ability to solve models, but on their ability to recognize missing information and ask the right questions before the model is built.
2. How It Works (The Technical Mechanics)
To make this concrete, the authors introduce two main contributions: a new benchmark called OR-Clarify, and a new framework called InterOPT.
The OR-Clarify Benchmark
Think of OR-Clarify as a controlled training ground for clarification. The benchmark takes fully specified OR problems and deliberately withholds key details (the "hidden slots"). For example, in a vehicle-routing task, the agent might see the locations and the demand, but the hidden fact—whether the route is open-ended or must return to the start—is kept secret.
The benchmark supports two interaction modes:
- Open/Free-form: The agent asks natural-language questions.
- Choice-based: The agent asks questions with multiple-choice options (e.g., "Is the tour open, closed, or does it depend on demand?").
It measures several things:
- Slot Recovery: Did the agent find the missing fact?
- Stopping Behavior: Did it stop too early or ask too many questions?
- Silent Assumptions: Did it assume something behind the scenes that wasn't verified?
- Interaction Cost: How many questions did it take?
The InterOPT Framework
The paper proposes InterOPT, a two-stage framework designed to decide what to ask and when to stop. It separates the problem of "diagnosis" from "action."
Stage 1: Dynamic Gap Search The agent maintains a "ledger" or memory of formulation-critical gaps. It scans the conversation and the initial brief to identify things that are still uncertain. Crucially, it categorizes these gaps into six types: objectives, decision scope, constraints, time boundaries, entity relationships, and hard vs. soft policies. It doesn't ask a question yet; it just records what's missing.
Stage 2: Gap-Guided Action Search Using the ledger from Stage 1, the agent generates candidate questions anchored to those specific gaps. It then selects one to ask the user. Critically, the agent is also instructed on when to stop. It’s told to ask more questions if an unresolved assumption could change the objective, constraints, or decision scope. It’s allowed to stop if the remaining uncertainty only concerns minor notational details or data collection.
The framework is designed to avoid two pitfalls: stopping prematurely (because the agent thinks it knows enough) and endlessly asking questions (even when the core facts are recovered).
3. Key Results & Benchmarks
The experiments compare InterOPT against various LLM baselines and other clarification methods. Here is the translation of the quantitative results into plain language:
- The Choice Setting (Multiple Choice Questions): This is where InterOPT shines. In the Choice-based experiments, InterOPT substantially outperforms all baselines in exact slot recovery. This means that when given options to choose from, InterOPT is much better at recovering the exact missing requirements needed to build the correct model.
- The Open Setting (Free-Form Questions): In the free-form setting, InterOPT remains competitive with strong prior methods, but it doesn’t uniformly dominate them. The paper notes that the challenge in the open setting isn't that the agent fails to ask questions—it asks about 4.8 questions per run—but that it fails to know when it has asked enough. The agent declares readiness in 98.8% of runs, yet the audit flags about 40% of those runs as "premature," meaning the agent stopped while formulation-critical gaps still remained.
- Interaction Cost: InterOPT uses more turns and atomic questions than some baselines. In the Choice setting, it asks more questions, but this appears to be a contributing factor to its better recovery rate.
- Silent Assumptions: InterOPT leaves fewer silent assumptions per run compared to baselines like MC-D (0.366 vs 0.692), meaning it's better at flagging unresolved issues rather than silently assuming a default.
A key takeaway from the results: No model exceeds 60% Core Exact recovery. This highlights that OR-Clarify is a genuinely hard benchmark, and even the best models struggle to reliably identify and recover all missing formulation-critical facts.
4. Why It Matters (Key Takeaways)
The paper concludes with several important points for anyone building or using AI for optimization:
- Clarify When Needed, Stop When Ready: The central thesis is that OR assistance should be a "selective completeness decision." The agent should ask questions to resolve formulation-critical gaps and stop only when the specification is sufficiently complete. The results show this is harder than it sounds; even advanced frameworks struggle with the stopping decision in open settings.
- The Gap Between Solving and Formulating: There is a significant gap between an LLM's ability to solve an optimization problem (given a complete model) and its ability to construct that model from scratch. Evaluations that only test the former miss the latter.
- Structured Interaction Helps: The Choice-based protocol, where the agent presents options, leads to better recovery than free-form questioning. This suggests that for critical business rules, providing structured choices can help both the AI and the user reach clarity more reliably.
- What to Watch For: Readers should watch for future work on "stopping calibration." The paper explicitly notes that stopping remains imperfect and that developing better ways for agents to know "enough is enough" is a key area for improvement. Additionally, the paper suggests that expanding the benchmark and improving question efficiency are important next steps.
In summary, this work reframes how we think about AI-assisted OR. It's not just about having the AI solve the math; it's about having the AI know what math to solve in the first place.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →