Apodex 1.1: Scaling Agentic Intelligence for Complex WorkExplained for Beginners
Apodex Team, B. An, B. Li +68 more
Abstract
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this working capability: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. Environment Scaling expands the diversity and verifiability of executable file, search, and code environments, while Agentic Coordination Scaling trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a Heavy-Duty Solver for ambitious, long-running tasks.
1. The Problem
Large language models have gotten remarkably good at answering questions, writing code, and synthesizing ideas. You can prompt them to explain a quantum physics concept, debug a tricky function, or summarize a market report, and they’ll often deliver useful, coherent text. But ask that same model to carry out a multi-step project—like “research three competing business models for a new eco-friendly product, pull the latest financial filings for each, write a comparison spreadsheet, and script a short demo video—and you’ll quickly hit a wall.
The gap isn’t reasoning ability. The gap is working capability: the capacity to sustain progress toward a real-world objective over many hours, across changing files, unreliable tools, and unexpected dead ends. Today’s models excel at producing a plausible answer in a single turn, but they struggle to maintain state when a file save fails, to recover from a botched code execution without losing valid work, to integrate evidence found at different times, or to deliver a finished artifact that another person can inspect and continue using. Prior benchmarks often measure an model’s final answer, treating the path there as a black box; if the path involves ten tool calls, a failed internet search, and a rewritten paragraph, most evaluation setups simply check whether the last line looks right—not whether the journey preserved correct dependencies, recovered from errors, or produced a verifiable result.
This matters outside academia because many valuable workflows naturally unfold over long horizons and require interaction with external environments. A financial analyst building a risk model, a researcher automating a literature review, a developer refactoring a legacy codebase, a lawyer reviewing contract clauses across dozens of documents—these tasks require the model to act like a persistent colleague, not a one-shot quizmaster. When AI can’t keep its own state straight, users end up doing the “glue work” themselves: re‑running searches, fixing lost citations, manually reconciling contradictory findings. Apodex 1.1 tackles exactly this problem by treating reliable, verifiable completion over time as the central unit of intelligence, rather than one‑shot answer quality.
2. How It Works (The Technical Mechanics)
Apodex 1.1’s approach can be understood through a analogy that should feel familiar to anyone who has managed a complex project or worked in a version‑controlled codebase. Imagine a project manager (the model) tasked with delivering a final report based on data scattered across multiple folders, spreadsheets, and a custom analytics script. The manager doesn’t just dictate the answer and walk away. Instead, they:
- Plan the work into subtasks—“gather sales data,” “check competitor filings,” “draft the comparison table.”
- Delegate each subtask to a specialist tool: a file inspector for spreadsheets, a search engine for market articles, a code interpreter for the analytics script.
- Track state in a persistent project board (the workspace). When the file specialist returns a cleaned‑up spreadsheet, the manager notes its location, version, and any assumptions. When the search specialist returns three articles, the manager records which claims are supported and which remain uncertain.
- Integrate results as they arrive. If the analytics script finishes early, the manager incorporates those numbers into the draft table immediately, rather than waiting for all subtasks to finish.
- Replan when observations change. If the competitor filings turn out to be outdated, the manager revises the research plan, perhaps asking for newer sources, without discarding the sales data that’s already validated.
- Recover from failures. If the code script crashes, the manager saves the partial output, logs the error, and re‑runs the script with a corrected input, preserving the valid parts of the earlier run.
- Deliver a verifiable artifact—the final report—accompanied by a clear trail: “Used Q3 2023 sales data from
sales/Q3.xlsx; pulled competitor info from URLs X and Y; regression run on commitabc123; output table matches verifier check X.”
In this analogy, the execution harness is the project‑management software that stores the board, records who did what, and allows the manager to rewind, replay, or intervene at any point. AgentOS is the underlying operating system for the project: it keeps the folder structure, saves intermediate files, and ensures that a failure in one sub‑task doesn’t erase the work done by others. It also supports “sticky routing”—staying on the same worker/node so that context isn’t lost between turns.
Now, how does Apodex 1.1 actually build this capability? It scales along two complementary dimensions, both powered by the same underlying harness and runtime.
Environment Scaling expands the diversity and fidelity of the worlds in which the model learns to act. The paper identifies three families of executable environments, each targeting a different bottleneck:
-
File worlds teach the model to inspect, transform, and preserve heterogeneous artifacts. Think of a workspace with nested directories, historical versions, mixed formats (PDFs, CSVs, images), and cross‑file references. The model must locate authoritative material, reconstruct relationships, and materialize that understanding as a professional deliverable. Scaling here means widening coverage across domains and occupations, and deepening structural complexity—the length of business‑logic chains the model must reconstruct—without merely adding more files.
-
Search worlds model open‑web research as discovery, acquisition, and evidence synthesis. The agent formulates and refines queries, triages candidates, follows references, and reconciles conflicting sources. The “gold” object isn’t just a final answer; it includes the relevant source set, claim‑to‑evidence alignment, and explicit uncertainty where sources disagree. Scaling changes the structure of the acquisition problem—distributed evidence, intermediate entities, heterogeneous access paths—rather than just increasing corpus size.
-
Code worlds are stateful environments where an agent changes repositories, dependencies, and processes, then receives executable feedback. Construction separates shared base images (interpreters, toolchains, tests) from task‑specific state. Coverage combines harvested worlds from real pull requests with synthesized ones that extend beyond existing distributions. Verification is sandboxed: tests must fail on the base state and pass after a reference change, or succeed in both states, and the system guards against reward hacking by isolating scoring from the solver.
The key insight is that these environments aren’t just “tools attached to a model.” They are structured, verifiable worlds with defined state transitions, resource budgets, and completion checks. By exposing a common policy to many such worlds, the model learns to handle file errors, search dead ends, and code bugs as regular parts of the task, not as out‑of‑distribution surprises.
Agentic Coordination Scaling expands how work gets organized across agents, time, and task branches. Long‑horizon work is as much a coordination problem as a computation problem. The model must decide how to decompose an objective, which branches can run independently, when partial results should reshape the shared plan, and when obsolete work should be abandoned. These aren’t properties of an inference‑time wrapper; they’re behaviors trained into the policy.
The paper’s “Agent Team” paradigm realizes these behaviors at runtime. A sub‑agent returns a useful intermediate result as soon as it’s ready. The lead agent integrates that result into shared state, revises priorities, informs or terminates other branches, and creates new workstreams if evidence changes the problem. A slow or failed branch doesn’t erase completed work elsewhere, and the user can redirect priorities while other branches continue. This isn’t about throwing more agents at a problem; it’s about the continuous, adaptive movement of results, decisions, and feedback through an active task.
Apodex trains decomposition, delegation, result integration, and replanning as part of the model’s working policy, then realizes those behaviors through Agent Team at runtime. Parallel execution supplies breadth; staged return and replanning supply feedback; shared state allows one branch to improve another’s value; explicit termination prevents obsolete work from consuming the remaining budget. The paper calls this the joint training‑and‑runtime paradigm Agentic Coordination Scaling: it scales the organization of work across agents, branches, and time, not merely the number of samples.
Connecting both dimensions is a common execution harness. This harness binds the model to File, Search, and Code environments; defines workspace, artifact, and provenance state; manages interaction budgets (turn budgets, tool‑call budgets, wall‑clock time); and supplies completion checks and verification hooks. On the coordination side, it defines sub‑agent execution state, staged returns, shared artifacts, lifecycle control, and verification hooks. Using one contract keeps trajectory collection, replay, training, and runtime execution aligned—imagine a single source of truth that both the model’s learning loop and its real‑time task execution reference.
AgentOS provides the persistent runtime beneath the harness. It maintains authoritative task state across tools, agents, context pressure, interventions, and partial failure. If a user intervenes (“Actually, prioritize the financial numbers over the market articles”), the runtime records that update in the trace, admits it into the solver‑visible state, and the model selects its next action from an updated context. Persistent state also makes intervention and recovery composable: new data may change a hypothesis, an intermediate result may expose a bad input, and the system can invalidate descendants whose dependencies changed while preserving causally independent progress. If the user materially changes the objective or delivery contract, the formalism treats the continuation as a new task, but valid prior state can be carried forward.
The same contract supports inspection and verification without exposing private chain‑of‑thought. Operational records expose plans, active and completed steps, artifacts, dependencies, failures, required decisions, and next actions. “Statement Review” then checks consequential claims against sources, computations, and artifact lineage—similar to iterative self‑feedback or learned process rewards, but attached to the persistent execution state shared by both scaling dimensions.
Training exploits both scaling dimensions in a controlled loop. A unified supervised‑fine‑tuning (SFT) mixture establishes common behavior across reasoning, tool use, recovery, delivery, and multi‑agent coordination. Agentic reinforcement learning (RL) then improves long‑horizon decisions over executable environment trajectories and coordination traces. Real tasks, benchmark errors, and runtime failures are classified into capability gaps; a “Task Pipeline” converts those gaps into new environments, coordination examples, or task specifications; Environment Scaling supplies the corresponding worlds and verifiers; and evaluation, expert review, and user feedback determine the next allocation of training effort. The paper emphasizes that self‑evolution is used only for this managed engineering loop, not for unconstrained model self‑modification.
The loop is deliberately incremental. After each training round, unsuccessful trajectories are analyzed for recurring failure modes, abstracted into capability‑level deficiencies, and fed back into the Task Pipeline. Newly generated tasks re‑enter the same family‑specific verification pipeline before contributing training signal. Model‑based diagnosis decides where to explore; environment verification decides which resulting worlds are trustworthy enough to train on. Difficulty is calibrated against what a world is intended to teach: for acquisition‑heavy file and search worlds, a first‑order pressure coordinate measures how many plausible candidates require expensive inspection relative to the tool‑call budget; for code worlds, calibration coordinates include dependency depth, state‑transition depth, test observability, and the distance between a failure and its executable verification signal.
Together, these components convert environment diversity and coordination complexity into reliable model behavior. The model doesn’t just get bigger; it gets broader in the kinds of work it can sustain, and deeper in how it organizes that work over time.
3. Key Results & Benchmarks
Across a remarkably broad spectrum of complex, real‑world domains—professional work, finance, scientific research, mathematics, coding, and search—Apodex 1.1 reaches the leading performance band, and it does so with a substantially smaller model than many frontier systems. The 35‑billion‑parameter Apodex 1.1 Mini variant retains strong working capability in a locally deployable form, reaching the performance band of selected frontier models and improving markedly over Apodex 1.0 Mini on overlapping tasks.
The evaluation design mirrors the paper’s two‑dimensional scaling framework. “ReAct” runs use a minimal scaffold to expose the underlying model policy, while “Agent Team” runs measure the system‑level lift obtained from trained coordination behaviors and additional organized computation. The main benchmark table presents a shared system snapshot, followed by capability analyses across science, finance, professional file work, IMO‑gold‑level mathematics, coding, internal long‑horizon tasks, and end‑to‑end case studies. The authors deliberately avoid reducing the system to a single aggregate score, instead presenting a richly partitioned view of where the model excels and where gaps remain.
In professional work benchmarks, Apodex 1.1 Agent Team competes directly with much larger frontier models such as GPT‑5, Claude Opus 4.6, and Gemini 1.5 Flash, despite using a fraction of their parameter count. On finance‑specific suites (FrontierFinance), the Agent Team system posts some of the strongest values in the comparative column, reflecting its ability to maintain stateful workflows, recover from tool failures, and deliver verifiable financial artifacts. In scientific‑research benchmarks (FrontierScience‑Research), similarly strong results appear, with the model achieving leading scores on tasks that require literature synthesis, experimental design, and data‑driven hypothesis generation.
Perhaps most striking is the Mini variant: the 35B‑parameter Apodex 1.1 Mini matches the performance band of selected frontier systems on many overlapping tasks, while running locally—a significant practical advantage for organizations that cannot send sensitive data to cloud APIs. Relative to Apodex 1.0 Mini, the new version shows marked improvements on long‑horizon coordination and recovery metrics, confirming that the two scaling dimensions are not just theoretical but translate into concrete gains.
To put the benchmark impact in plain language: imagine a mid‑size car that, through better driving strategy, road knowledge, and adaptive pacing, beats a fleet of much larger, more powerful limousines on a mixed‑terrain rally. The car doesn’t have the biggest engine, but its driver knows when to accelerate, when to conserve fuel, how to navigate tight corners, and how to recover from a skid without losing the overall route. Apodex 1.1 is that skilled driver. On standard professional‑work and scientific‑research tests, it answers roughly 10–15 % more questions correctly or completes roughly 10–15 % more subtasks successfully compared to similarly sized prior models, and it stays competitive with models that are several times larger. The Mini variant delivers a respectable fraction of that gain on a single GPU, opening the door for on‑premise deployment in regulated environments.
4. Why It Matters (Key Takeaways)
-
From answer generators to verifiable workers: Apodex 1.1 represents a shift in what we ask of general‑purpose models. By formalizing “working capability”—sustained, verifiable progress toward a real objective—the paper sets a new benchmark for AI utility. The implications are immediate: automated financial due‑diligence, iterative scientific data pipelines, multi‑step code refactoring, and large‑scale document preparation can now be entrusted to a single system that remembers its state, learns from mistakes, and delivers a defensible artifact rather than a disposable paragraph.
-
Locally deployable power: The 35B‑parameter Mini variant proves that you don’t need a hundred‑billion‑parameter cloud model to get serious working capability. For companies subject to data‑sovereignty rules, for research groups with limited cloud budgets, or for any organization that prefers keeping sensitive inputs on‑premise, Apodex 1.1 Mini opens the door to agentic AI that runs in‑house, on existing hardware, without sacrificing the majority of the coordination and recovery skills that the full system demonstrates.
-
Scaling that isn’t just about size: The paper’s core message—that environment scaling and agentic coordination scaling are independent, combinatorial surfaces that can yield outsized gains—offers a roadmap for the next generation of AI development. Rather than chasing ever larger parameter counts, researchers and product teams can focus on building richer executable environments and training paradigms that teach models to decompose, delegate, integrate, and recover. The harness and AgentOS design shows that reliable long‑horizon work is as much about system architecture as model size.
-
What to watch next: The Apodex team’s longer‑term goal is a “Heavy‑Duty Solver” capable of taking responsibility for increasingly ambitious, long‑running tasks. Key things to monitor include how the evaluation suite evolves (whether new benchmarks capture even richer state‑maintenance and provenance requirements), how environment coverage expands into new domains (healthcare records, legal discovery, embedded systems debugging), and whether the trained coordination behaviors generalize to truly open‑ended, user‑driven workflows where the objective itself shifts mid‑task. Also worth watching is the balance between autonomous task completion and human‑in‑the‑loop oversight: as these models become more capable, the interface design for intervention, course‑correction, and final sign‑off will determine how much of the “working capability” actually reaches end users versus remaining a laboratory curiosity.
Overall, Apodex 1.1 isn’t just a incremental update; it’s a re‑framing of what a large model can do over time. It moves the field from measuring how well a model can complete a single prompt toward measuring how well it can finish a real job—and that difference, between a clever parlor trick and a reliably useful teammate, is exactly why the paper matters.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →