arXiv:2608.22849nvidia/nemotron-3.5-lightning-30b-a3bAugust 24, 2026

RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA ModelingExplained for Beginners

Ziyuan Wang, Bohao Tang, Fei Zhang +2 more

Machine LearningGenomics

Abstract

Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution. We present RIBOSPAN, a 1.61-billion-parameter bidirectional RNA foundation model natively pretrained with context lengths up to 10,240 nt. RIBOSPAN combines dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing to enable high-resolution modeling of complete long RNAs. We evaluate the model through nucleotide reconstruction, a controlled long-context representation benchmark, and frozen RNA-type representation analysis. Native 10K pretraining preserves strong reconstruction at 10,240 tokens, while continued pretraining with 40% masking improves recovery under heavy corruption while preserving representation quality. The long-context benchmark further shows that native 10K models maintain strong contextual responsiveness and context-specific representation separation while keeping perturbation-induced representation changes highly localized. Inference-time YaRN scaling recovers much of the contextual organization lost by direct extrapolation of short-context models, but induces substantially greater distal representation diffusion. Frozen-representation evaluations further demonstrate state-of-the-art RNA representation quality, with RIBOSPAN achieving the strongest overall performance across diverse RNA types and retaining a clear advantage on long RNAs. Building on the same backbone, we develop a multidimensionally conditioned discrete-diffusion framework for full-length mRNA generation and redesign, including synonymous-codon diffusion for protein-preserving CDS optimization. Together, RIBOSPAN establishes a powerful long-context foundation for transferable RNA representation learning and full-transcript mRNA design.

RIBOSPAN: A Long-Context RNA Foundation Model

The Problem: The Limits of "Truncated" RNA Intelligence

Imagine trying to read a novel, but every time you reached the 20th page, the book suddenly stopped. You could still understand the plot up to that point, but the ending, the character development, and the subtle foreshadowing scattered throughout the final chapters were permanently lost.

That is the reality for most RNA foundation models today. Messenger RNAs (mRNAs) can be incredibly long—often stretching over 10,000 nucleotides—but the majority of existing RNA foundation models are pretrained with a "context window" of only about 1,024 tokens. In practical terms, if a model's context window is 1,024, any mRNA longer than that has to be chopped into pieces. The model processes each piece in isolation, losing the big picture.

For a field interested in designing therapeutic mRNAs or understanding gene regulation, this is a critical gap. A mutation in the 3' untranslated region (UTR) at the very end of a transcript can influence how the protein is made at the beginning. To capture these long-range dependencies, a model needs to see the whole story at once. Existing models force a choice: either sacrifice the complete narrative (by truncating) or sacrifice the fine-grained detail (by compressing multiple nucleotides into single tokens, losing single-nucleotide resolution). Until now, no model had managed to do both: model full-length transcripts at single-nucleotide resolution while maintaining deep, bidirectional understanding.

How It Works: The Mechanics of RIBOSPAN

RIBOSPAN is a 1.61-billion-parameter bidirectional Transformer—think of it as a super-powered, bidirectional version of the architecture behind models like BERT, but specifically tuned for RNA.

Here is how it achieves the impossible:

1. Single-Nucleotide Tokenization

Most models use "byte-pair encoding" or other compression methods, where a single token might represent "CG" or "AU." RIBOSPAN assigns one token to every single nucleotide. This means the model never loses the granular detail of the sequence. If you change a single 'A' to a 'G', the model notices.

2. Attention-Isolated Sequence Packing

Training a model with 10,000-token context windows is computationally expensive. To make this feasible without losing information, the authors use a clever trick called "attention-isolated sequence packing." Imagine packing multiple suitcases into one large trunk. RIBOSPAN packs several different RNA sequences into one long training line. Crucially, it uses a special mask to ensure that the attention mechanism "knows" where one sequence ends and another begins. It attends to nucleotides within its own sequence but ignores those in the packed neighbors. This allows the authors to effectively increase the training throughput while still teaching the model to handle very long inputs.

3. RoPE Positional Encoding

The model uses Rotary Positional Embeddings (RoPE), which provide the model with a sense of where each nucleotide sits in the sequence. This is vital for long contexts, as it helps the model distinguish between the 5' end and the 3' end of a long mRNA.

Key Results & Benchmarks: The Numbers Tell the Story

The paper evaluates RIBOSPAN across three main tasks. The results are compelling, especially regarding the "native context" advantage.

Nucleotide Reconstruction (The "Fill-in-the-Blank" Test)

The models were tasked with reconstructing masked portions of mRNA sequences. Here is the plain-English summary of the results:

  • The Extrapolation Problem: If you take a model trained on short contexts (1,024 tokens) and try to use it on a long transcript (10,240 tokens) without retraining, its performance crashes. It’s like asking a reader who only knows the first 20 pages of a book to accurately predict the ending; the linguistic patterns have shifted too much.
  • The Native Advantage: RIBOSPAN trained natively on 10,240 tokens (RIBOSPAN-10K) maintains strong reconstruction even at the full 10,240 length. The short-context models that were later "scaled up" using a technique called YaRN (more on this later) still degrade significantly.
  • The Masking Trade-off: Continuing the training of the native 10K model with a heavier 40% masking rate improves its ability to recover the sequence even when heavily corrupted, without sacrificing its basic understanding.

Plain-Language Take: Native 10K pretraining preserves strong reconstruction at 10,240 tokens, while continued pretraining with 40% masking improves recovery under heavy corruption while preserving representation quality.

Long-Context Representation Benchmark (The "Context Sensitivity" Test)

The authors designed a clever experiment to see if the model truly "understands" the context of a long RNA. They took a long RNA sequence, made a small change in the middle, and checked if the model's representation of the unchanged ends of the RNA changed.

  • Localization: The native 10K models keep representation changes highly localized. If you change a nucleotide in the middle of a 10,000-nt transcript, the representation of the beginning stays mostly stable.
  • The YaRN Trap: Inference-time YaRN scaling (a method to extend short models to long contexts) recovers much of the lost contextual organization. However, it induces "substantially greater distal representation diffusion." In plain English: YaRN helps the model understand the meaning of the long context, but it makes the model "nervous"—changes in one part of the RNA ripple unpredictably to distant, unrelated parts.
  • HydraRNA Comparison: Compared to a state-of-the-art long-model called HydraRNA, RIBOSPAN strikes a balance. HydraRNA is very good at keeping changes localized (low diffusion), but it is worse at differentiating between different regional contexts (lower context separation). RIBOSPAN manages to do both: it maintains strong regional separation while keeping diffusion in check.

RNA-Type Representation (The "Classify the RNA" Test)

Here, the authors froze the model's weights and asked: "If I just look at the pattern of activations across the whole transcript, can I tell what kind of RNA this is?" (e.g., mRNA, tRNA, rRNA).

  • State-of-the-Art: RIBOSPAN achieves the strongest overall performance across diverse RNA types.
  • The Long-Rain Advantage: This is a key differentiator. On "Long RNAs" (sequences longer than 1,024 nt), RIBOSPAN retains a clear advantage. Other models struggle to organize these long sequences, but RIBOSPAN's representations remain coherent and well-separated.

Plain-Language Take: Frozen-representation evaluations demonstrate state-of-the-art RNA representation quality, with RIBOSPAN achieving the strongest overall performance across diverse RNA types and retaining a clear advantage on long RNAs.

Why It Matters: Key Takeaways

  1. A New Tool for mRNA Design: The authors built a full-generation framework on top of RIBOSPAN. This means researchers can now use this model to design full-length synthetic mRNAs. Not only can it generate a sequence, but it can optimize the "coding sequence" (the part that makes the protein) using "synonymous-codon diffusion." This allows for protein-preserving optimization—changing the DNA code without changing the protein, which can improve stability or manufacturability.
  2. Bridging the Gap: RIBOSPAN successfully bridges the divide between "generation" models (which can write RNA but might not understand deep context) and "representation" models (which understand context but can't write new sequences). It does both.
  3. The YaRN Caution: The paper provides a nuanced view of YaRN, a popular technique for extending context windows. While YaRN is effective, the results suggest it comes with a cost: it increases "distal representation diffusion." For applications where you need precise, localized understanding of a long RNA, native long-context pretraining (as done in RIBOSPAN) is superior.
  4. What to Watch For: The model is large (1.61B parameters), which means it requires significant computational resources to run. However, the release of the model and framework on Hugging Face and GitHub makes it accessible for the broader research community.

Summary

RIBOSPAN is a significant step forward for RNA bioinformatics. By natively pretraining a billion-parameter model on sequences up to 10,240 nucleotides using single-nucleotide tokens and attention-isolated packing, it solves the long-context problem that has plagued the field. It retains the detail of single-nucleotide resolution while capturing the broad, bidirectional dependencies of full-length transcripts. Its applications in mRNA design and its nuanced performance relative to existing extrapolation techniques make it a model to watch in the evolving landscape of genomic foundation models.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →