← All posts

BLOG · FIELD RECORDS

Newtype Slaves Made Newtype Slaves — 5 Days of Work with OS/Role Distribution, Supervision/Review, and Testing

Development Report

Publication Date: October 5, 2026

Introduction

The previous article described how three models worked as one team. This article is a record of what that team built. The first commit to the repository was on October 1, 2026, and 266 commits accumulated over 5 days (10/1 11 · 10/2 6 · 10/3 45 · 10/4 167 · 10/5 37). The Newtype Slaves TUI and Nexus, which are now public, were built using the team formation method described on the homepage (OS-specific distribution, role-specific distribution, supervision/review, testing). Times are UTC ledger times, and identifiers are not included.

Team Composition

Team setup: how roles are split between AI agents and a person
RoleWhoWhereAssigned Task
AdministratorNexus Session (10/2–10/4)Nexus · Name changed 6 times, continuedReview and judgment only, no implementation: fix deliverables, judgment documents, write failed tests
CoordinatorClaude Code Session "claude-newtype"macOS · Terminal (channel + remote control)Planning, task distribution, review, deployment execution, requesting decisions from humans
Server/DeploymentClaude Code Session "newtype sub"macOSNexus server functions, release pipeline, server upgrades
Windows LeadNexus Session → Claude Code Session "newtype-windows"Windows 11 PCDirect build, test, and installation verification on Windows
TUI WorkerNewtype Slaves TUI "nmcp"macOS · GPTExecute tasks received under delegation
TUI WorkerNewtype Slaves TUI "slave app"macOS · GeminiExecute tasks received under delegation
Implementation WorkerMultiple Claude sub-agentsmacOSImplement feature units (each in a separate work folder/branch)
Independent ReviewerClaude sub-agentmacOSIndependently review changes made by other workers
OwnerHuman (Tony)Anywhere · Including mobile phoneDetermine goals/scope, click approval emails, key input

The server runs on one Linux (GB10) machine. Nexus and PostgreSQL are on top of it, and it is accessible from the outside via lic.newtype-ai.com.

1. Role-based Distribution — Coordinator divides, workers only their own folders

When the coordinator received a request, they broke down the work and sent it to the workers. Each worker worked in their own work folder or branch.

For the first three days (October 1-3), sessions on Nexus divided roles by name. There were separate roles for operational application, TUI verification, operational implementation, and Windows lead. 114 sessions were created on October 1 (KST) alone, most of them feature-sized tasks. The administrator session was responsible only for review and judgment, not implementation. This role continued without interruption even when sessions changed. The administrator session's name changed 6 times, and each new session read and inherited the ledger from the previous session. From the night of October 3, the administrator and Windows sessions worked in English to save tokens, and the coordinator translated for them into Korean.

For the last two days, the distribution was as follows:

  • Execution Grant: If within the scope of delegation in the received turn, it was executed without human confirmation, and the result was replied to via reply_to.
  • TUI English Screen: Approximately 950 screen texts were separated into language-specific lists. After being created in a separate branch to avoid touching the same files as other workers, they were merged onto the latest code at the end.
  • Tasks Running Simultaneously: Remote MCP endpoint, installation service, /restart modification, open-weight model support, and homepage revamp were simultaneously carried out by different workers.

There were a few moments when two workers tried to modify the same folder simultaneously. Each time, we followed the rules: if there was an uncommitted file by someone else, stop and report; commits are made by specifying file paths one by one.

2. OS-specific Distribution — Bugs only visible on that OS

All tests passed on macOS. When the Windows PC session ran the same code directly on Windows, product bugs not visible on Mac appeared.

Split by OS: finding a Windows bug that did not show on macOS
  • All saves failed: a combination of file options NTFS does not support broke credential, model-profile and team-status saves. If deployed as is, Windows users would not have been able to complete their initial registration.
  • All Tool Execution Denied: Applying Unix-style permission checks to the execution history file caused all tools to be "not executed" from the default start of saving conversations.
  • Windows Path Delegation Denied: The C:\… format was not accepted by the server.
  • Installation Script Line Breaks: If the server was built with a repository received on Windows, the installation script could be corrupted with CRLF.
  • Package App Virtualization: If AI was installed within a desktop app, files went into app-specific storage and only the PATH changed. The installation location was moved to %USERPROFILE%\.local.

On the Linux server, the server image was built twice to confirm that the bytes were identical, and clients were also verified in the same way across 6 platforms.

Remaining Issues: Windows Smart App Control blocks unsigned executables. Self-signed certificates were not trusted, and public certificates are pending.

3. Supervision/Review — Separate implementation from review

Supervision and review: a workflow where the implementing agent and the reviewing agent are separate

Administrator Session Method. When a worker produced results, the administrator first fixed the commit. Then, after verifying that all submitted hashes were correct, they reviewed it in a separately extracted work folder and documented the judgment.

  • 86 Judgment Documents: Accumulated over approximately 19 hours from the evening of October 3 to the afternoon of October 4. The most common judgment was "narrowly accepted, needs improvement, operation pending."
  • Meaning of Acceptance: Acceptance was merely a record that "this evidence meets the criteria." Deployment, secret input, and approval emails were always human responsibilities.
  • Pre-written Failed Tests: The administrator did not just verbally point out defects but directly wrote and provided failing tests. Workers had to pass these tests without lowering expectations or skipping them, and the passed tests remained as regression tests. This resulted in 85 administrator tests (55 Go, 30 Python), approximately 7,000 lines.
  • Multiple Round Trips: The custody rehearsal filter was conditionally accepted on the fourth attempt after being pending three times. Nexus custody, secret delivery plan, and control channel v2 also underwent three reviews each. Team features were fixed across four review documents.
  • Waking Up a Stalled Administrator: The administrator session often ended its turn even when there were still plans remaining. So the coordinator checked the plans every 30 minutes and sent nudges ("N incomplete steps remaining, but the turn has ended"). Decisions requested by the administrator when the owner was away were delivered via email.

Independent Review in the Last Two Days. The remote MCP endpoint (OAuth 2.1) was created by the server lead and reviewed independently three times by other Claude sub-agents.

  • 1st Review, 12 cases: The most critical issue was that email approval only worked once in a lifetime due to a permanent idempotent key. Tests created with a fake repository passed, but the reviewer reproduced it with a real repository and found the issue.
  • 2nd Review: 11 issues were fixed. A problem remained where the cumulative limit of the approval repository could be exhausted by unauthenticated requests.
  • 3rd Review: Passed after changing the hourly limit to a daily limit.

The coordinator also played a supervisory role. They did not blindly trust worker reports but directly re-ran commits and tests to verify them before merging.

4. Testing — Passing is useless if it doesn't match reality

It took four attempts to bring "proceed" prompts to zero (previous article). During the three failed attempts, unit tests and e2e tests all passed. This was because the tests were configured differently from the actual TUI.

So the server lead re-created the e2e tests with the same code path as the actual TUI. If the fixed part is intentionally reverted, the test fails with the same message as the actual error. After that, tests that run the actual executable in the terminal (whether it returns to the same session after restarting) were also added.

5. Conversation Instead of Commands

During this period, the owner did not look up or type commands directly. When speaking to the coordinator in Korean, the coordinator executed the commands or distributed them to the workers.

What the Owner SaidWhat Actually Happened
"Let's deploy the blog with the sidebar changes."The coordinator ran the GitHub Actions deployment and confirmed its success.
"Let's release the Windows version for now."The server manager built, signed, and published the release, including Windows.
"Proceed with B recovery."The coordinator reverted to the previous settings and restarted the server.
"Default to light mode when the homepage loads."The implementer fixed it, and the coordinator built and deployed it.

The same applies within the TUI. The following tasks can be done by speaking, without knowing commands.

  • Naming: "Our session name is nmcp"
  • Delegation: "Re-delegate all permissions except re-delegation and secret information access to claude-newtype" → Preview → 1
  • Update: "Latest update"
  • Language: "Change to English"
  • Enabling and disabling default model usage
  • Turning on email consent for remote connections

Statements that broaden permissions or change settings are always accompanied by a preview and approval window. And that approval is only possible in a turn directly typed by the owner. Such actions cannot be initiated by messages from other sessions.

The only actions directly performed by the owner were clicking approval emails, entering keys into hidden input fields, and allowing keychain access.

6. What the Owner Did

  • Decisions: What to build, what to postpone (e.g., Windows public certificate), wording and design
  • Approvals: Server upgrade approval email, keychain access, signing key and token input
  • Scope: Issuing delegations

Before the grant, the owner had to type prompts like "proceed," "Did you check?", and "Check your inbox" more than 10 times. Under the grant, tasks ran and were answered with zero prompts (read → replied in 23 seconds).

What Was Difficult

  • 15-minute halt: The server refused to start after a new feature was enabled without migration. We reverted to the previous settings, created a migration tool to build the table without stopping the server, and then restarted it.
  • Approval email in spam: One approval email went into spam, so the request was resent.
  • Confirmation reply ping-pong: On October 2nd, over 30 minutes in the afternoon, 114 messages were exchanged between the admin, operations, TUI verification, and Windows sessions. Most were "Confirmed" in response to "Confirmed." After that, we established a rule not to send replies that only confirm receipt.
  • Admin session interruption: On October 4th, when the Mac was rebooted, the admin session left a checkpoint and stopped. After that, the server's initial installation steps (P0–P6) were directly accepted by the owner after reviewing the coordinator's evidence.
  • Worker's misjudgment: The worker incorrectly stated that "migration will happen," and the coordinator trusted it without verification. Since then, we directly verify statements from peer sessions.

Results

This is what the team delivered in the last two days (October 4-5).

  • Execution grants (issued via server, client, and conversation)
  • Remote MCP endpoint (OAuth 2.1)
  • Installation services (llm.txt, install.sh, install.ps1) and three signed releases (0.20261005.1–.3)
  • Sessions that resume after restart, open-weight model support, English interface
  • newtype-ai.com revamp and wiki

The AI can be installed by having it read https://dev.newtype-ai.com/llm.txt. Operation records will continue to be public.

© 2026 NEWTYPE. All rights reserved.

← All posts