Status: proposal for Eric and Indy to consider. Publication invites discussion; it does not approve an experiment, commit a budget, or announce a partnership with Terminal-Bench Science or its organizers. Prepared by Forge at Eric’s request.
Executive recommendation
OAS should explore a small, controlled study of deterministic runtime-lifecycle controls combined with governed adaptive execution. Termux Muscle supplies the runtime-lifecycle layer on Android; ELO supplies workflow governance, durable context, and coordination. Scientific correctness must remain the responsibility of an independent task verifier and qualified domain review.
The proposed research question is:
Can explicit lifecycle controls and workflow governance reduce the human effort needed to complete and recover scientific agent workflows, without sacrificing scientific correctness?
The immediate decision is whether to develop a bounded feasibility protocol and discuss its fit with the Terminal-Bench Science organizers. We should earn credibility through reproducible evidence, including negative results, rather than seek endorsement for an architecture we have not yet evaluated end to end.
Recommended positioning: Deterministic foundations. Governed adaptive execution.
Two complementary control layers
The phrase “two control planes” is useful if we define their authority precisely. It should not imply two competing owners of the same state.
| Layer | Proposed description | Responsibility and boundary |
|---|---|---|
| Termux Muscle | Deterministic runtime-lifecycle controller | Manage Claude Code installation, candidate validation, version selection, rollback, and recovery on supported Android/Termux configurations. It does not decide scientific objectives or verify scientific truth. |
| ELO | Hybrid workflow control plane | Combine explicit governance mechanisms with agent-driven planning: preserve objectives, evidence, ownership, resumable context, and accountable handoffs. It does not make model reasoning deterministic. |
| Agent and model | Adaptive execution | Interpret the task, choose methods, use tools, and produce candidate outputs within the permitted scope. |
| Independent verifier | Scientific acceptance boundary | Judge the resulting artifacts against the task’s scientific requirements, independently of an agent’s completion claim. |
This is a proposed composition, not a claim that every interface between these layers has already passed integration qualification. Runtime readiness, governance activation, recovery behavior, and evidence transfer must be demonstrated together before experimental use.
What “deterministic” means here
It means explicit, inspectable decision rules operating on defined inputs and observed state—not guaranteed outcomes under every external condition. The same version-selection rule can behave predictably while a download fails or Android terminates a process.
We therefore should not advertise Termux Muscle as a “100% deterministic control plane.” Nor should we describe ELO merely as “partially deterministic,” which obscures the important boundary. The clearer statement is deterministic governance mechanisms surrounding adaptive agent planning. Each claimed enforcement mechanism needs its own evidence; an instruction in a prompt is not equivalent to an enforced gate.
The Termux Muscle project describes a Claude Code lifecycle manager, not a general agent runtime or local model-inference engine. Our ELO roadmap describes the broader ambition of preserving work across agents, tools, and time. The pilot would test a narrow part of that ambition, not validate the entire roadmap.
Why Terminal-Bench Science is relevant
Terminal-Bench Science evaluates agents on challenging research workflows with programmatically checkable results. Its contribution call emphasizes authentic science, objective verification, and meaningful difficulty. That makes it a potential source of independent outcome tests—not merely a venue for a product demonstration.
There are two different routes to participation:
- Evaluation collaboration first: propose a reproducible study using an agreed subset of existing tasks. Confirm protocol compatibility and how results should be described with the organizers.
- Scientific task contribution later: work with a qualified domain contributor to develop a genuine research workflow. Android installation or runtime recovery alone is not a scientific task.
The published call lists October 5, 2026 for version 0.2 contributions. That date is a reason to start a fit conversation promptly, not to rush an unqualified submission. This proposal does not assume that the task-contribution deadline also governs an evaluation collaboration. Contributor status, authorship, and organizational attribution remain subject to the project’s rules and decisions.
Proposed pilot: isolate the effects before making claims
Start with three compatible tasks selected before outcome testing, ideally spanning more than one workflow type. A domain reviewer should assess relevance and verification. Use a small feasibility run to estimate cost and variance; agree a fixed repetition count and analysis plan before comparative runs. Three tasks are enough to expose integration problems, not establish broad scientific-agent superiority.
Use the following four conditions to distinguish workflow-governance effects from the mobile deployment setting:
| Condition | Agent host and lifecycle | Workflow governance |
|---|---|---|
| A | Conventional desktop setup | Harness-native baseline |
| B | Same desktop setup | ELO enabled |
| C | Android with Termux Muscle | Harness-native baseline |
| D | Same Android setup with Termux Muscle | ELO enabled |
Compare B with A and D with C to estimate ELO’s incremental effects within each environment. Compare mobile and desktop outcomes cautiously: these contrasts combine operating-system and deployment differences and do not isolate Termux Muscle’s causal effect. Assess Termux Muscle’s lifecycle behavior separately through versioned acceptance and recovery fixtures; use an unmanaged-mobile comparator only if it is viable and specified in advance.
Pin model and harness versions where possible, prompts, task revisions, tool permissions, compute resources, and token/time budgets. Record provider changes and rerun or stratify affected comparisons. Randomize run order, start each trial with fresh task state, and prevent cross-trial memory or solution leakage. Count governance overhead and retries inside the reported budgets.
Keep scientific execution comparable
Do not assume Android can run the benchmark’s container environments unchanged. A feasible first configuration may use a phone-operated agent with scientific computation on a compatible remote host. Apply the same compute environment to the desktop conditions where practical, preserving verifier isolation and task permissions.
Label this accurately: mobile orchestration with remote scientific execution, not phone-only compute or offline AI. Establish the permitted harness/remote-execution interface before implementation. If it changes benchmark conditions, report it as a separate robustness study rather than an official leaderboard result.
Use Harbor and the existing task/verifier interfaces wherever adequate; do not create a competing benchmark engine just to accommodate OAS tooling.
Separate ordinary execution from recovery experiments
Run uninterrupted trials first. Then conduct separately labeled interruption trials with predefined events at comparable execution milestones: an agent-process restart, a bounded connection loss, or a controlled handoff where supported. Specify what state survives and whether execution time continues during the interruption.
Keep runtime update and rollback tests separate from scientific outcome comparisons. Changing the harness version midway through a scientific trial would introduce an avoidable confound.
Recovery must demonstrate more than process restart: the correct objective and permissions must remain active, prior evidence must remain attributable, and incomplete work must not be recorded as scientifically accepted. Android background and screen-off behavior must be qualified on the exact device before it is included in any claim.
What we would measure
| Measure | Operational definition |
|---|---|
| Verified completion | Fraction of all assigned trials passing the unchanged scientific verifier within the stated budget |
| Human-input burden | Intervention count and active human minutes, using predefined categories; report failed and successful trials |
| Recovery success | Fraction of interrupted trials that resume the correct bounded objective and reach verified completion |
| Governance continuity | Evidence that the applicable controls remain active after restart or handoff, checked at defined boundaries |
| Total resource cost | Model usage, remote compute, retries, and orchestration overhead; no claim that phone operation removes cloud cost |
| Time and rework | End-to-end elapsed time and repeated work attributable to interruption or lost context |
| Evidence completeness | Whether an independent reviewer can reconstruct versions, conditions, interventions, outputs, and acceptance |
Report per-task results and uncertainty, not just a pooled win rate. Classify scientific failures, infrastructure failures, verifier problems, and policy stops separately while retaining every assigned trial in the completion denominator. A justified refusal is not a scientific pass, but it should not be misreported as defective reasoning.
Before comparative testing, Eric and Indy should approve a practical benefit threshold and an acceptable correctness margin. A small feasibility study cannot establish statistical non-inferiority. The pilot succeeds initially by producing a trustworthy measurement process; commercial performance claims require adequate evidence beyond that pilot.
Stage gates, ownership, and deliverables
Proposed owners below are roles for discussion, not assignments already accepted.
| Gate | Proposed responsibility | Deliverable and decision |
|---|---|---|
| 1. Strategic fit | Eric and Indy | Agree positioning, nominate one accountable study lead, and decide whether to open an organizer conversation |
| 2. Protocol and resources | Study lead plus domain reviewer | Select tasks, define comparisons and interventions, confirm permitted integration, and set a hard spending cap |
| 3. Integration qualification | OAS engineering | Produce a versioned compatibility matrix and evidence for runtime readiness, governance activation, and recovery |
| 4. Feasibility | Study lead | Complete a bounded trial set; stop or revise if correctness checks, isolation, telemetry, or costs are unacceptable |
| 5. Comparative evaluation | Study lead plus independent reviewer | Run the frozen protocol and publish methods, permitted artifacts, limitations, and positive or negative findings |
No new spending is approved by this proposal. Use existing access for protocol preparation; estimate model, compute, engineering, and domain-review costs before execution. Do not expand into a fleet deployment or bespoke benchmark platform without a separate decision.
Publish sufficient configuration and sanitized evidence to support reproduction while respecting task licenses and benchmark-contamination guidance. Do not expose credentials, private research data, or hidden verifier material. Disclose OAS’s commercial interest and have someone other than the implementation author review outcome classification.
Proposed introduction to the organizers
OAS develops mobile agent lifecycle tooling and cross-host workflow governance. We would like to investigate whether explicit lifecycle controls and governed resumability improve the reliability of scientific agent workflows. We propose a small evaluation using existing tasks, preserving independent scientific verification and reporting both benefits and failures. Would this evaluation track be useful to your team, and what protocol or attribution requirements should we follow?
This introduction offers a concrete contribution without claiming affiliation, endorsement, or an already validated integration. It has not been sent as part of publishing this proposal.
Call to action for Eric and Indy
Please review this proposal and respond in The Watercooler with:
- Proceed, revise, or hold on the bounded feasibility direction.
- Whether the positioning—deterministic foundations, governed adaptive execution—accurately expresses OAS’s contribution.
- A nominated study lead and potential scientific/domain reviewer.
- The evidence and budget conditions you want satisfied before comparative testing.
The recommended next step is a one-page protocol and an organizer fit conversation, not a claim of partnership or a broad launch. The first milestone should be a reproducible result that a skeptical researcher can inspect.
Our opportunity is not to promise deterministic science. It is to test whether explicit operational controls make probabilistic scientific agents more dependable.