What is Runtime AI Scientist?
Runtime AI Scientist is the open-source integration and evaluation code for Can AI Scientists Coordinate at Runtime? It studies Runtime Agent Coordination (RAC): choosing which agent acts next while a scientific research task is running.
The project applies the same coordination conditions to three existing AI-scientist hosts: ARK, Agent Laboratory, and EvoScientist. It preserves each host's agents, models, tools, and permissions. The code was previously called RAC AI Scientist (RAC × AI Scientist); its Python package and command remain rac-ai-scientist.
Benchmark results · Installation · Frequently asked questions · Cite the paper
Abstract
Multi-agent AI scientists have shown improving performance across a diverse range of tasks. Yet a common approach is design-time agentic orchestration, which typically relies on fixed workflows. In contrast, human scientists coordinate and adjust their division of labor at runtime. We therefore ask: can AI scientists also coordinate at runtime?
To this end, we introduce Runtime Agent Coordination (RAC), which selects agents from existing AI-scientist hosts during execution, assigns scoped work contracts, and provides artifact-grounded verification. Verification informs subsequent agents without blocking transitions or discarding artifacts. We conduct a single-seed exploratory evaluation across Agent Laboratory, EvoScientist, and ARK on ResearchClawBench, preserving host models, tools, and permissions under host-calibrated budgets. Four cumulative conditions separate native execution, runtime communication, runtime selection, and the combined addition of contracts and verification. Runtime selection yields the highest observed mean score for each host; adding contracts and verification reduces these means, with host-dependent outcomes relative to native execution. These results motivate runtime coordination while exposing the limits of additional coordination mechanisms under constrained budgets.
01Four cumulative conditions
| ID | Configuration | What it adds |
|---|---|---|
| N0 | Native lifecycle | The host's own scheduler runs end to end, with no RAC phase loop and no extra communication. |
| R1 | Runtime communication | Agents exchange requests, results, and artifacts; every handoff still goes to the host's native successor. |
| R2 | + Runtime selection | The current agent selects the next host capability from a fresh checkpoint, without a separate orchestrator. |
| R3 | + Contracts and verification | Scoped work contracts, plus a supported, refuted, or inconclusive verdict passed to the next agent. It never blocks, retries, or rolls back. |
02Results on ResearchClawBench
| Condition | ARK | Agent Laboratory | EvoScientist |
|---|---|---|---|
| N0: Native lifecycle | 16.66 | 5.47 | 15.99 |
| R1: Runtime communication | 17.40 | 9.63 | 12.07 |
| R2: + Runtime selection | 18.42 | 12.08 | 18.53 |
| R3: + Contracts and verification | 17.98 | 9.88 | 15.96 |
Mean score per host and condition (Table 3 of the paper); tasks are equally weighted and recorded zeros retained. This is a single-seed exploratory evaluation on 10 ResearchClawBench tasks (5 for ARK) under host-calibrated budgets. Token usage, estimated cost, task-level scores, and a DiscoveryBench transfer study are in the paper.
03Run more experiments with us
The paper's ResearchClawBench experiments can be rerun from the open-source integration: one shared N0–R3 policy, isolated Docker images for each host, revision-pinned upstreams, and an external scorer that alone sees the benchmark targets. The most useful next experiments are the ones a single-seed study leaves open.
More seeds
Repeat N0 and R2 on the paper's tasks with new seeds. Does runtime selection keep its lead?
More tasks and domains
ResearchClawBench has far more tasks than the paper ran. Find where runtime selection helps and where it hurts.
Other models and budgets
Swap the evaluated model, or compare at equal spend. Is the effect model- or budget-dependent?
New AI-scientist hosts
Expose another system's roles through the HostBridge protocol and compare it under the same rules.
04Quickstart
git clone https://github.com/systemind-team/Runtime-AI-Scientist.git
cd Runtime-AI-Scientist
python -m venv .venv && . .venv/bin/activate
python -m pip install -e .
python -m unittest discover -s tests -v # offline checks, no model calls
rac-ai-scientist --help
Live episodes need Linux with Docker Compose, an OpenAI-compatible endpoint for the evaluated model, and a judge model for scoring. See the README and the runbook (中文).
Frequently asked questions
How does runtime coordination differ from a fixed workflow?
A fixed workflow decides the next role in advance. With runtime selection, the current agent chooses the next authorized host capability using the current artifacts, open problems, execution history, and remaining budget. See the N0–R3 conditions and the architecture documentation.
Is Runtime AI Scientist a standalone AI-scientist host?
No. It is a coordination integration and evaluation layer for ARK, Agent Laboratory, and EvoScientist. The hosts supply the scientific agents and tools; this project compares how those agents coordinate.
What did the experiments find, and what are the limits?
Runtime selection (R2) had the highest observed mean ResearchClawBench score for each evaluated host. Adding contracts and verification (R3) lowered those means. The study used one seed and 10 tasks (5 for ARK) under host-calibrated budgets, so it does not establish a general improvement across tasks, models, or repeated runs. See the results table and paper.
How can I reproduce or extend the experiments?
Start with the offline installation checks, then follow the runbook for live host setup, budgets, and scoring. The contribution guide describes experiments with more seeds, tasks, models, and hosts.
Where are the official paper and source code?
The paper is arXiv:2610.00980. The official repository is systemind-team/Runtime-AI-Scientist. Use the BibTeX citation when citing the study.
05Citation
@misc{liu2026aiscientistscoordinateruntime,
title = {Can AI Scientists Coordinate at Runtime?},
author = {Zijian Liu and Yangzhixin Luo and Junyu Lu and Yi Li and Yu Chen and David Xu and William F. Shen and Xinchi Qiu and Xisen Wang},
year = {2026},
eprint = {2610.00980},
archivePrefix = {arXiv},
primaryClass = {cs.MA},
url = {https://arxiv.org/abs/2610.00980}
}