← All posts

BLOG · FIELD RECORDS

How Three Models Worked as One Team — A Two-Day Record of Making "Proceed" Zero

Publication Date: October 5, 2026

Introduction

In the previous article, we outlined six principles for spreading NMCP. The fifth principle was this:

A record of "this is what happened when we tried this" is more valuable than a claim of "this is how it should be done."

This article is that record. For two days, we made three different models work as a team. We document, in order, what blocked us until we reduced the number of times a human had to re-type "proceed" to zero. We include both successes and difficulties. All times are UTC from the ledger.

Team

Role What Model
claude-newtype Claude Code Session (Coordination) Claude
nmcp Newtype Slaves TUI Agent gpt-6-astra
slave app Newtype Slaves TUI Agent Gemini 3 Flash (Preview)
newtype sub Claude Code Session (Server/Deployment) Claude
newtype-windows Claude Code Session on Windows machine Claude

Connecting them is Nexus. Nexus is a ledger where sessions from the same account exchange messages and tasks. And there is one human (Tony). The goal was to reduce the human's role to pressing buttons in approval emails and deciding what to do.

1. Day One: Human as Intermediary

The first round trip was at 11:29 on October 4th. The message sent by Claude was recorded in the ledger as delivered and read by the TUI agent. There was no need to ask again if the recipient had read it. Here, "read" is recorded only at the moment the message enters the model's input. It is not considered read merely by being displayed on the screen or retrieved.

The problem was what came next. When Claude requested a code task via message, the agent only read it and replied with a plan. It neither modified files nor ran tests. This is by design. This is because a message is a request, not an authorization. A single line of text sent by someone else should not execute commands on my computer.

As a result, the human had to manually type "proceed" into the TUI for the work to begin. Prompts like "proceed," "did you confirm?", "check your inbox" exceeded 10 times. Tony said:

This is a bit funny... actually, you're doing all the instructing and I'm just saying "proceed"... this very inconvenience is what we need to eliminate with what we're building.

2. Execution Grant

The answer was an execution grant. A human sets the scope once.

Messages sent by claude-newtype, within this repository, including file read/write and command execution, for 50 turns, 8 hours.

Execution grant concept diagram

Once this is set, every time a tool is called in a turn initiated by a received message, the Nexus server makes a decision (allow / ask / deny). The rules are as follows:

  • A grant can only narrow an existing delegation, not broaden it.
  • Messages received before the grant was issued are rejected, even if at the same time.
  • Secrets, Nexus manipulation, and external tools always ask the human, even with a grant.
  • Writing "Sender: X" in the message body is ineffective. Only authenticated senders are considered.

We first deployed the server, then attached the client, and then enabled issuing grants via conversation ("Let's delegate all permissions except re-delegation and secret access to claude-newtype" → Preview → Human presses 1).

3. Four Attempts to Reach Zero

We sent the same request ("Please run git log and go test and reply with the results") four times. Three attempts failed, and each failure revealed a problem at a different layer.

Attempt Time What Happened What Was Fixed
1 14:11 The model saw the word "grant" in the message and tried to create another grant. The human received two confirmation prompts, and ultimately no work was done. In a received turn, human-only tools are not shown at all. An instruction is added at the start of the turn: "A grant for this scope already exists, execute immediately and reply to the sender."
2 14:33 The reply came without a confirmation prompt. The content was "Could not do it because there are no execution tools." The TUI was not providing any file or shell tools to the model by default. The TUI launched by a human provides local tools by default.
3 14:39 The tool was called, but the server denied it: scope newtype:run is not granted. The TUI's delegation did not include the execution scope itself, and a grant cannot broaden a delegation. Added local execution scope to TUI delegation (execution decision still determined by mode/human/grant).
4 14:49 Success —

Here is the ledger for the fourth attempt.

Time Event
14:49:21.215 nmcp reads message
14:49:21.442 Server decision: allow (grant)
14:49:30.606 run_command executed, decided_by: grant
14:49:44.572 Reply with results (linked to original message)

Human intervention was 0 times. It took 23 seconds from reading to replying, most of which was model invocation and test execution time.

Most painful lesson: Even with the flaws in attempts 2 and 3, both unit tests and e2e tests running a real Nexus all passed. This was because the tests directly allowed tools by name and manually created delegation scopes. This differed from how the actual TUI started. Now, e2e tests create delegations with the same code as the TUI. If we revert that change, the tests fail with the same messages as the actual errors.

4. Waking Up Idle Agents

Sending a message is useless if the recipient is idle. We connected Nexus to Claude Code's channel feature.

Time Event Elapsed
12:18:14.905 Message 0 sec
12:18:20.896 Delivered, session woke up +6.0 sec
12:18:21.175 Read +6.3 sec
12:18:28.608 Reply +13.7 sec

Notifications do not carry the message body. They only announce "You have a new message," and the read receipt is recorded when the model directly opens the inbox. This is to prevent others' messages from entering the model without a receipt.

Thanks to this feature, the orchestrating Claude no longer needs to check every few minutes while waiting for a reply from nmcp. It wakes up when a reply arrives.

5. Third Model

When Gemini was added to the team, two 400 errors occurred.

  1. The tool format definition (JSON Schema) contained items (additionalProperties, default, oneOf) that Gemini does not accept. We didn't know this until then because GPT models accepted them. We made sure to convert the format to one that Gemini accepts when sending.
  2. Gemini 3 requires a signature (thoughtSignature) to be attached when calling a tool, and expects to receive it back unchanged in the next request. We were discarding that signature.

After fixing both, at 15:31:29, the Gemini agent executed a command within its delegated authority and responded. At the same time, the gpt-6-astra agent also finished another task. Three models were mixed in one team. The fourth principle ("not tied to a specific model") from the previous article is now a ledger entry, not just a design.

6. Difficulties Encountered

Image standing for the difficult parts

15-minute outage. When enabling the remote MCP endpoint (/mcp, OAuth 2.1), the server refused to start. The reason was mcp oauth schema missing. The deployment tool only changes images and does not perform DB migrations. However, we trusted a colleague's session saying "tables are created upon deployment" without verifying it. We reverted to the previous configuration to restore service. Then, we rehearsed with a migration tool on a throwaway DB, created the tables without stopping the service, and restarted it. The human's actions were just two phrases: "proceed with recovery" and "proceed with migration," plus an approval email.

Three reviews, twelve defects. The same remote endpoint was independently reviewed three times by different Claude sessions. The most serious defect was that email approval worked only once in a lifetime due to a permanent idempotent key. This also passed tests created with a fake repository. A reviewer found it by reproducing it with a real repository.

Windows. When a Claude session on a Windows machine ran tests on actual hardware, two product bugs appeared that were not visible on Mac.

  • All saves failed. An unsupported combination of options for NTFS was being used.
  • All tool executions were rejected at the default startup for saving conversations.

And Smart App Control did not consider self-signed certificates as a basis for trust. Even with the same signing method, some file hashes passed while others were blocked. Windows deployment requires publicly trusted code signing.

Numbers

These are the four numbers defined in the previous article.

Number Current
From installation to first message Installation alone takes 2 seconds (curl -fsSL https://lic.newtype-ai.com/install.sh | sh, 0.20261005.1). The time from a human's first execution (email approval, folder trust) to the first message has not yet been measured.
Number of agents per team 5 concurrently (three models)
Ratio of the same team running the next day The ability to return to the same session, conversation, and delegation when restarted was added this time. Measurement starts now.
Number of human "proceed" actions More than 10 before delegation → 0 based on delegated tasks

Still To Do

  • Owner-only operation. Currently, others cannot sign up. This is the biggest obstacle to the second principle ("1-minute installation") from the previous article.
  • Remote connections cannot wake up. Clients connected only via URL do not wake up when a message arrives while they are idle. Waking up only works for locally connected clients.
  • Specifications are not yet in a public repository.
  • Windows signing.

Next

The first macOS/Linux release (0.20261005.1) and installation service were launched while this article was being written. For AI, https://dev.newtype-ai.com/llm.txt needs to be read. Next, we will measure the time it takes for AI to read llm.txt on a clean device, install it, and then go through the human's first execution to the first message. We will write the next article with that number.

© 2026 NEWTYPE. All rights reserved.

← All posts