arXiv 2610.00980 · Open source

Runtime AI Scientist

Can AI Scientists Coordinate at Runtime?

Zijian Liu, Yangzhixin Luo, Junyu Lu, Yi Li, Yu Chen, David Xu, William F. Shen, Xinchi Qiu, Xisen Wang
University of Oxford · King Abdullah University of Science and Technology · University of Sydney · University of Cambridge
Runtime AI Scientist in 46 seconds

What is Runtime AI Scientist?

Runtime AI Scientist is the open-source integration and evaluation code for Can AI Scientists Coordinate at Runtime? It studies Runtime Agent Coordination (RAC): choosing which agent acts next while a scientific research task is running.

The project applies the same coordination conditions to three existing AI-scientist hosts: ARK, Agent Laboratory, and EvoScientist. It preserves each host's agents, models, tools, and permissions. The code was previously called RAC AI Scientist (RAC × AI Scientist); its Python package and command remain rac-ai-scientist.

Benchmark results · Installation · Frequently asked questions · Cite the paper

Abstract

Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime?

To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets.

Same agents, different coordination: native fixed workflows versus Runtime Agent Coordination, with controlled factors, the RAC decision loop, and the N0–R3 condition ladder
Same agents, different coordination. Native transitions versus state-dependent selection by the current agent. Agents, models, tools, permissions, and run-level budget are held fixed; the cumulative conditions add communication (R1), runtime selection (R2), and contracts plus verification (R3) to the native lifecycle (N0).

01Four cumulative conditions

IDConfigurationWhat it adds
N0Native lifecycleThe host's own scheduler runs end to end, with no RAC phase loop and no extra communication.
R1Runtime communicationAgents exchange requests, results, and artifacts; every handoff still goes to the host's native successor.
R2+ Runtime selectionThe current agent selects the next host capability from a fresh checkpoint, without a separate orchestrator.
R3+ Contracts and verificationScoped work contracts, plus a supported, refuted, or inconclusive verdict passed to the next agent. It never blocks, retries, or rolls back.

02Results on ResearchClawBench

ConditionARKAgent LaboratoryEvoScientist
N0: Native lifecycle16.665.4715.99
R1: Runtime communication17.409.6312.07
R2: + Runtime selection18.4212.0818.53
R3: + Contracts and verification17.989.8815.96

Mean score per host and condition (Table 3 of the paper); tasks are equally weighted and recorded zeros retained. This is a single-seed exploratory evaluation on 10 ResearchClawBench tasks (5 for ARK) under host-calibrated budgets. Token usage, estimated cost, task-level scores, and a DiscoveryBench transfer study are in the paper.

03Run more experiments with us

The paper's ResearchClawBench experiments can be rerun from the open-source integration: one shared N0–R3 policy, isolated Docker images for each host, revision-pinned upstreams, and an external scorer that alone sees the benchmark targets. The most useful next experiments are the ones a single-seed study leaves open.

More seeds

Repeat N0 and R2 on the paper's tasks with new seeds. Does runtime selection keep its lead?

More tasks and domains

ResearchClawBench has far more tasks than the paper ran. Find where runtime selection helps and where it hurts.

Other models and budgets

Swap the evaluated model, or compare at equal spend. Is the effect model- or budget-dependent?

New AI-scientist hosts

Expose another system's roles through the HostBridge protocol and compare it under the same rules.

04Quickstart

git clone https://github.com/systemind-team/Runtime-AI-Scientist.git
cd Runtime-AI-Scientist
python -m venv .venv && . .venv/bin/activate
python -m pip install -e .
python -m unittest discover -s tests -v   # offline checks, no model calls
rac-ai-scientist --help

Live episodes need Linux with Docker Compose, an OpenAI-compatible endpoint for the evaluated model, and a judge model for scoring. See the README and the runbook (中文).

Frequently asked questions

How does runtime coordination differ from a fixed workflow?

A fixed workflow decides the next role in advance. With runtime selection, the current agent chooses the next authorized host capability using the current artifacts, open problems, execution history, and remaining budget. See the N0–R3 conditions and the architecture documentation.

Is Runtime AI Scientist a standalone AI-scientist host?

No. It is a coordination integration and evaluation layer for ARK, Agent Laboratory, and EvoScientist. The hosts supply the scientific agents and tools; this project compares how those agents coordinate.

What did the experiments find, and what are the limits?

Runtime selection (R2) had the highest observed mean ResearchClawBench score for each evaluated host. Adding contracts and verification (R3) lowered those means. The study used one seed and 10 tasks (5 for ARK) under host-calibrated budgets, so it does not establish a general improvement across tasks, models, or repeated runs. See the results table and paper.

How can I reproduce or extend the experiments?

Start with the offline installation checks, then follow the runbook for live host setup, budgets, and scoring. The contribution guide describes experiments with more seeds, tasks, models, and hosts.

Where are the official paper and source code?

The paper is arXiv:2610.00980. The official repository is systemind-team/Runtime-AI-Scientist. Use the BibTeX citation when citing the study.

05Citation

@misc{liu2026aiscientistscoordinateruntime,
  title         = {Can AI Scientists Coordinate at Runtime?},
  author        = {Zijian Liu and Yangzhixin Luo and Junyu Lu and Yi Li and Yu Chen and David Xu and William F. Shen and Xinchi Qiu and Xisen Wang},
  year          = {2026},
  eprint        = {2610.00980},
  archivePrefix = {arXiv},
  primaryClass  = {cs.MA},
  url           = {https://arxiv.org/abs/2610.00980}
}