NEWTYPE · OPEN BENCHMARK · OPEN BENCHMARK
NewtypeBench — a methodology proven in public
“NewtypeBench evaluates AI coding agents on full-stack feature shipping — like SWE-Bench, but for the modern AI-native stack and judged by real production tests.”
- corpus 123 tasks
- HF v0.2.2
- license BSL
- stack FastAPI · React · Vite
leaderboard — coming soon
PROVEN IN PUBLIC
We speak in numbers
Metrics already proven on the public NewtypeBench benchmark.
We do not hide the frontier model’s measured 30%. Even the best models score at this level — which is exactly why measurement matters.
EXAMPLE
This is what an evaluation looks like
Given a natural-language spec, the agent submits a 4-line patch, and the harness auto-grades it by execution results.
@@ -141,6 +141,10 @@ def register_routes(router):
return ThreadListResponse(items=threads)
+ @router.get("/threads/{thread_id}/export")
+ def export_thread(thread_id: str, db=Depends(get_db)):
+ thread = get_or_404(db, Thread, thread_id)
+ return render_markdown_export(thread)- FAILBefore the patch — new-feature tests failing
- RUNrunning pytest · Playwright … 32,885 tests
- PASSFAIL_TO_PASS 940 passed ∧ PASS_TO_PASS 31,945 unbroken
See for yourself
The methodology, scoring logic, and dataset are all public. The company name appears nowhere in the benchmark — a neutrality principle.
WHY THIS STACK
Evaluating AI agents on the stack AI builders actually use
The evaluation stack is no arbitrary choice — each tool ranks #1 in usage for its category.
Stack survey evidence
| Technology | Usage | Source |
|---|---|---|
| React | 82% | State of JS 2024 · #1 in category |
| Vite | 78.1% | State of JS 2024 · #1 in category |
| Tailwind | 62% | State of CSS 2024 · #1 in category |
| FastAPI | 38% | JetBrains 2024 · #1 in category |
─ 16,209 Indeed postings for “fastapi react”
A unique axis in the benchmark landscape
| Benchmark | What it measures |
|---|---|
| SWE-Bench | Bug fixes in mature libraries |
| FullStackBench | General code quality |
| Terminal-Bench | CLI tasks |
| NewtypeBench | Modern AI-native stack × full-stack feature shipping |