arXiv:2609.07470nvidia/nemotron-3.5-lightning-30b-a3bSeptember 7, 2026

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action PolicyExplained for Beginners

Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos

RoboticsArtificial Intelligence

Abstract

Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rather than translation. Several plausible instruments produce false conclusions: a color-histogram metric rewards noise, a single-goal benchmark scores 84.6% under correct Greek and 82.6% under deliberately wrong instructions, training loss fails to predict Greek success, and single-run comparisons are dominated by seed variation. On a discriminative ninety-task suite with three seeds per arm, a multilingual text tower without Greek demonstrations remains at its wrong-instruction floor, while Greek-only training exceeds its control by at most 2.7 points. Bilingual training yields a consistent 6.7-7.1 point margin over its control and reaches about two fifths of English performance. The policy also overfits the translator's phrasing; training on seven phrasings per task approximately halves this penalty. Warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance. The results support two practical requirements for low-resource robot-policy localization: build a guaranteed null before trusting a metric, and replicate low-resource-language results across seeds.

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

1. The Problem

Imagine you’ve spent millions of dollars training a robot to understand natural language. You show it a video of a kitchen scene, type “Pick up the orange,” and the robot smoothly reaches for a fruit bowl. Now imagine the same robot, but you type the command in Greek: “Κάντε lên το πορτοκαέ.” The robot freezes. It doesn’t understand.

This scenario highlights the central problem this paper tackles: robot foundation models are overwhelmingly trained and evaluated in English. For the vast majority of the world’s languages, there simply aren’t enough robot demonstrations—videos of humans doing tasks—to go around. We have a "data desert" for non-English languages in robotics.

The authors of this study wanted to see if they could teach a robot some Greek without starting from scratch. Specifically, they took a state-of-the-art system called Cosmos3, which connects vision (seeing), language (understanding words), and action (moving the robot). They wanted to add Greek to this system.

However, there was a catch. The authors suspected that the "measurement" part—the way we tell if the robot actually learned Greek or if we’re just seeing tricks—was harder than the actual teaching. They found that standard ways of measuring success could lie to you. This paper is a reality check on how we judge language capabilities in robots, using Greek as a test case.

2. How It Works (The Technical Mechanics)

To understand the mechanics, it helps to know what Cosmos3 is. At its core, Cosmos3 is a "vision-language-action" model. Think of it as a super-advanced translator and executor combined. It takes an image (vision), processes a text command (language), and outputs movements for a robot arm (action).

The researchers weren’t allowed to change the robot’s brain architecture. They couldn’t retrain the whole model from scratch. Instead, they had to work within the existing system.

Here is the clever part of their method: Machine-rephrased instructions.

Since they didn’t have humans writing Greek robot demos, they used automated tools to translate English instructions into Greek. But not just a simple dictionary swap. They used "machine-rephrasing," which attempts to rewrite the sentence while keeping the meaning.

The authors approached this as a measurement problem first. They tested several ways to judge if the robot was doing well:

  • The Color-Histogram Metric: This sounds technical, but essentially measures colors in the image. The authors found this was "gamed" by noise. The robot might look like it’s succeeding because the image happened to have the right colors, even if it grabbed the wrong object.
  • Single-Goal Benchmark: They tested the robot on a set of tasks. Interestingly, the robot scored 84.6% when given correct Greek instructions, but scored 82.6% when given deliberately wrong instructions. This suggests the robot was mostly ignoring the language and just doing what it always does, or the metric was too easy.
  • Training Loss: Usually, we look at "loss" (error) to see if a model is learning. The authors found that low training loss didn’t necessarily mean the robot understood Greek. It might have learned English patterns so well that Greek just "felt" similar mathematically, masking a failure of true transfer.

The key technical insight is that the robot uses a "text tower"—a component that processes the words. The authors experimented with freezing this tower (keeping it fixed) or unfreezing it (letting it update). They also looked at "warm-starting" from a language-adapted world model.

3. Key Results & Benchmarks

The most striking results come from their ninety-task suite. They tested the robot on ninety different tasks, running three different random seeds (initial conditions) for each arm to see if results were consistent.

Here is the plain-language translation of their numbers:

  • The "Wrong Instruction" Floor: Even with incorrect Greek instructions, the robot scored around 82-84% on their single-goal benchmark. This reveals a harsh truth: the robot was largely oblivious to the language change. It was performing at a baseline level regardless of the words used.
  • Greek-Only Training: When they trained the system specifically with Greek (using the machine-rephrased instructions), the performance improved over their "control" (the base system) by at most 2.7 points. That’s a tiny bump. It suggests that just throwing Greek text at the system without careful design yields minimal benefit.
  • Bilingual Training: This was the winner. Training the system with both English and Greek instructions yielded a consistent margin of 6.7 to 7.1 points over the control. More importantly, it reached about two fifths (40%) of the English performance. If English performance was a score of 100, Greek-augmented performance was roughly 40. This is a significant jump from the 2.7-point gain, showing that deliberate bilingual exposure matters.
  • The Phrasing Trap: The researchers discovered the model "overfits the translator's phrasing." If they trained on only one way of saying "Pick up the cup" in Greek, the robot got confused if the phrasing changed slightly. However, when they trained on seven different phrasings per task, they approximately halved this penalty. The robot became more robust to how the Greek was worded.

4. Why It Matters

This paper is a valuable guide for anyone building robot policies or evaluating AI systems in low-resource languages. Here are the key takeaways:

  • The "Null" Metric Rule: The authors’ most practical advice is poignant: "Build a guaranteed null before trusting a metric." In plain English, this means you must establish a baseline of "random" or "wrong" performance before you celebrate any gains. If you don't know what failure looks like, you can't measure success accurately. The single-goal benchmark scoring 82.6% on wrong instructions versus 84.6% on correct ones is a perfect example of why this is necessary.
  • Replication Across Seeds: The paper emphasizes that single-run comparisons are "dominated by seed variation." If you run an experiment once and see a result, you might just be seeing luck of the draw. To trust that "adding Greek actually works," you need to run the experiment multiple times with different random starting conditions. The fact that bilingual training yielded a consistent 6.7-7.1 point margin across seeds is what makes it credible.
  • The Danger of Overfitting: The finding that training on seven phrasings per task halves the penalty is a crucial technical detail. It suggests that for low-resource language transfer, diversity of data (different ways of saying the same thing) is more important than just having more data. If you only train a robot on one phrase in Greek, it will be brittle.
  • Caution with Warm-Starting: The authors found that "warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance." This is counter-intuitive—usually, we think starting smart helps. But in this case, it seems the model got "confused" or "locked into" certain patterns. This warns researchers that simply loading a pre-trained model isn't a silver bullet; it requires careful tuning.

In summary, this study serves as a necessary corrective. It shows that adding Greek (or any language) to a robot like Cosmos3 isn't just a matter of translation. It requires specific measurement rigor, bilingual training data, and diversity of phrasing. The payoff is real—bilingual policies perform significantly better—but the path there is paved with methodological pitfalls that the field must navigate.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →