arXiv:2609.10016nvidia/nemotron-3.5-lightning-30b-a3bSeptember 9, 2026

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk RuntimesExplained for Beginners

Remco Hendriks

Machine LearningArtificial IntelligenceComputation and Language

Abstract

We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.

MetroLLM-Bench: When the Kiosk Gets Smarter

1. The Problem

Imagine you’re at a transit kiosk, ready to buy a ticket. The screen asks your origin and destination, checks if there’s a service disruption, and calculates your fare. Now imagine that instead of a fixed piece of code running the show, a language model is making those decisions in real-time. That’s the premise of MetroLLM-Bench, a new benchmark from Continker that puts language models in the driver’s seat of transit kiosk operations.

The real-world problem this addresses is tangible: transit agencies constantly update fare rules, adjust for disruptions, and modify accessibility requirements. Currently, each change triggers a software deployment cycle—developer tickets, integrator patches, regression testing, and release windows. The paper argues this is costly and slow. The alternative? Replace the hardcoded policy layer with a language model that can read updated "framebooks" (system descriptions) and adapt on the fly.

Why should you care, even if you don’t work in transit? Because this is a microcosm of a broader shift. We’re moving from writing explicit rules for AI systems to letting models learn them from descriptions. The kiosk is a clean testbed: the interactions are short, the goals are clear (buy a ticket), and the stakes are concrete (a wrong fare means a billing error). The benchmark tests whether a model can be not just smart, but usefully smart in a structured, operational context.

2. How It Works (The Technical Mechanics)

The benchmark presents the model with a "case"—a specific transit scenario. The model’s job is to interact with six structured tools (think of them as API calls) and produce a machine-renderable terminal state. This state must include:

  • An outcome (e.g., "route_and_fare_ready" or "advisory_only").
  • A per-ticket fare quote when applicable.
  • A kiosk action with a reason code and a passenger-facing message.

The model operates within a ReAct loop (Reasoning and Acting), limited to 20 rounds of tool calls. It starts with a "framebook"—a system prompt detailing the specific metro system's terminology, currency, operating hours, and cultural conventions. Crucially, the model doesn't need to know the city's map from pretraining; the framebook and tools provide all necessary facts.

The Analogy: Think of this as a sophisticated version of a SQL query or a script, but driven by natural language. Instead of a developer writing IF disruption THEN take_route_A, the model reads the disruption from a tool, consults the framebook, and decides the best route. For a software engineer, imagine the model as a dynamic stored procedure that takes inputs (origin, destination, disruption status) and produces a structured output (route, fare, action), with the "code" being the model's weights and the "system prompt" being the parameters.

3. Key Results & Benchmarks

The benchmark evaluated 26 models across six metro systems and 955 cases. The results reveal a fascinating landscape of capability and efficiency.

The Standout Performer: A 4B parameter Qwen 3.5 model, fine-tuned using Parameter-Efficient Fine-Tuning (PEFT). Despite its small size, it scored 91.32 on Tier 1 (the deterministic, tool-use focused tier). This bested both tiers of GPT-5.6 (scoring 90.63 and 90.00) and came within a hair's breadth of GPT-5.4 full at maximum reasoning effort (91.37). What's remarkable is the footprint: this performance comes from a model weighing just 2.6 GB in Q4_K_M quantization—small enough to potentially run on edge hardware.

The Scale Puzzle: The researchers trained students at four sizes: 2B, 4B, 9B, and 27B. You might expect the largest model to win. Instead, the 4B student turned out to be the "Goldilocks" zone. The 9B and 27B students provided no further Tier 1 improvement over the 4B student at this training scale. In fact, for the 27B student, PEFT (fine-tuning) actually reduced performance by 0.91 points compared to the base model—a surprising result that suggests too much capacity, combined with the specific training recipe, can sometimes be detrimental.

The "Configuration" Factor: The study found that serving configuration significantly impacts results. For instance, moving from a 4,096-token output budget to 16,384 tokens helped larger models, but the specific sampling parameters (temperature, top-p) mattered a lot. The paper notes that "serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points," highlighting that how you run the model matters as much as the model itself.

4. Why It Matters (Key Takeaways)

Here are the four most significant takeaways from the paper:

  • Frontier API ≠ Deterministic Superiority: The headline result is that a 2.6 GB open-weight model can match or exceed top-tier proprietary APIs like GPT-5.4 and GPT-5.6 on this specific, bounded task. This is a "parity claim on deterministic task performance." It suggests that for well-defined operational tasks, you don't necessarily need to pay for expensive API calls; a local, adapted model can do the job.
  • The 4B Sweet Spot: The capacity-ceiling curve is a central finding. Under the tested configuration (QLoRA rank 16, 3 epochs), 4B parameters appears to be the sweet spot for this type of PEFT adaptation on top of Qwen 3.5. Larger models (9B, 27B) didn't help and, in the case of 27B, actually harmed performance relative to their base state. If you're deploying on constrained hardware, smaller is better, and 4B seems to capture the optimal trade-off between capability and efficiency.
  • Adaptation Matters More Than Raw Size: The PEFT gain dropped dramatically as base model size increased—from a +7.03 point boost at 2B down to -0.91 at 27B. Furthermore, "every seed shows the same direction at every size": the 2B, 4B, and 9B students improved over their bases at every seed, while the 27B students regressed. This indicates that for highly capable bases, fine-tuning needs to be approached carefully; more parameters don't automatically mean better adaptation.
  • The Gap is in the "Decision" Categories: The language model's advantage over the rule-based baseline is concentrated in categories requiring judgment: temporal reasoning, policy changes, accessibility, and compound scenarios. Routing and fare calculation are relatively easy (even a scripted agent does well here). The real value of these models lies in their ability to handle the "messy," context-dependent parts of transit operations—like deciding when a disruption means you need to change routes or how to communicate an accessibility issue to a passenger.

In summary: MetroLLM-Bench provides a rigorous framework for evaluating language models not as general chatbots, but as operational policy engines. It suggests that the future of kiosks—and perhaps many other interactive systems—might not require larger and larger models, but rather smarter, more efficiently adapted models that fit the specific constraints of the task at hand.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →