Quantified
Binary 0/1 verdicts and pass rates. Numbers instead of “it seems to work.” Numbers earn trust only when the judging criterion is objective.
NewtypeBench evaluation methodology
NEWTYPE · AGENT EVALUATION · ENTERPRISE AGENT EVALUATION
We quantitatively judge an agent’s performance and quality against work actually performed by people — no domain expertise required. The judge is not human opinion but execution results.
The world already has countless agent and AI-model benchmarks
But the agent you are building cannot be evaluated by any of them.
Enterprise work differs from general knowledge domains, and real workflows are far more complex and precise. That is why the evaluation standard must be the work your people actually performed — not a benchmark score.
WHY EVALUATION
A demo is not real work capability. A HumanEval-style benchmark score is not your company’s job performance. To claim results, you must measure against the work itself.
“Only by evaluating how accurately the built agent performs against work done by people can you truly measure the outcome of an agent-building project.”
Binary 0/1 verdicts and pass rates. Numbers instead of “it seems to work.” Numbers earn trust only when the judging criterion is objective.
NewtypeBench evaluation methodology
The judging standard is comparison against real work artifacts, not domain expertise. The evaluator does not need to be a domain expert.
Generalization of the “answer = actually merged PR” principle
Run it again in a different environment six months later and get the same result. A container-pinned evaluation environment guarantees reproducibility.
Docker-based reproducible environment
HOW IT WORKS
No LLM-as-judge. The only judges are the execution results of pytest and Playwright. That is why the evaluator needs no domain expertise.
Work artifacts actually produced by people become the answer key (ground truth).
Package the work order (spec) + starting state + judging criteria into a reproducible task.
The agent performs the work under identical conditions. The answers stay hidden.
New work passes ∧ existing work unbroken — judged by execution alone.
FAIL_TO_PASS PASS_TO_PASS → Did it accomplish the new work ∧ avoid breaking what already worked? Derived automatically by set operations — no manual labeling.
7-stage curation pipeline
SERVICE 01 — EVALUATION
Judged by execution against work actually performed: does it accomplish the new work, without breaking what already worked? The score is 0 or 1.
We grade only the agent’s final output. No need to reveal your agent’s internals.
SWE-Bench-compatible patch interface
Your agent connects to a real working environment and performs the work; the harness auto-extracts and grades the results.
Agent-loop interface
Agent v N vs v N−1 relative comparison — the dashboard of your improvement loop.
Relative signal across versions
Compare multiple candidates (models, vendors) under identical conditions to ground your adoption decision.
Evaluating 3–5 frontier models
| Party | Scope of definition |
|---|---|
| We define | Task specs + grading core + environment reproducibility |
| You define | Everything about the agent — prompts, tools, model choice, loop strategy |
─ This boundary becomes the service contract itself — the neutrality of an evaluation company.
SERVICE 02 — BUILD
We are not evaluation-only. We build domain-specific agents ourselves and measure our own results. We practice “building designed to be evaluable” — from day one we define together “what will judge this agent’s success,” and delivery ships with a quantitative evaluation report.
Meeting-note and report summarization, document classification and translation. What people actually wrote is the grading standard.
ground truth: summaries/classifications staff actually produced
Approval-document analysis, policy-violation detection, approval routing. Judged against past processing history.
ground truth: historical approval records
Recurring report generation, metric aggregation, anomaly detection. Accuracy measured against existing human-made reports.
ground truth: reports people already produced
Code changes, ticket handling, first-response incident analysis. Merged PRs and closed tickets are the answers.
Direct application of the NewtypeBench methodology
Not industry expertise but “anywhere work leaves a trail” — examples of the domain-agnostic principle.
─ The same engine as our evaluation methodology powers quality control of the build business. Only possible when building and evaluating live in one company.
THE FLYWHEEL
Only by measuring the agent’s accuracy against human-performed work can you speak to the ROI of a build project.
Anywhere work leaves a trail
= the answers (ground truth)
Accuracy · regression · cost
Evaluation results feed the next improvement cycle — evaluation is the heart of the loop.
The Docker-based evaluation harness runs inside the same closed network. Agents built on local LLMs are measured in-house as they are.
“The main purpose is the agent’s internal quality-improvement loop; the public leaderboard is a by-product”. Our build projects run their quality control on the same self-evaluation loop.
TRUST
Tasks unfavorable to our own agents are mandatory. We promise to publish unfavorable scores too. Vendor names are excluded from the benchmark.
Tasks are tiered public / held_out / internal_only; only post-cutoff data goes held-out, rotated by season. We control even for the chance the agent memorized the answers.
Methodology, scoring logic, and reproduction steps are fully public. Anyone can reproduce the same results by the same procedure.
BSL source-available + public dataset
“We publish results even when they are unfavorable to us. We believe that is what qualifies an evaluation company.”─ Fairness pledge