Evaluating Coding Agents Beyond SWE-bench: Why Unit Tests Lie
Why pass@1 on synthetic unit tests doesn't correlate with end-to-end fullstack app generation; visual regression, dynamic runtime verifications, and holistic product evaluation.
The SWE-bench Obsession
In the research community, SWE-bench and HumanEval have become the gold standard for ranking coding models. Leaderboards boast decimal-point improvements on resolving historical GitHub issues.
While SWE-bench is a valuable research benchmark for single-file bug patches in established Python repositories, it is almost completely decoupled from the reality of autonomous product creation.
When an autonomous system like Qbit builds a modern fullstack web application from a conversational brief, standard unit tests tell only a fraction of the truth.
The Failure Modes Unit Tests Can Never Catch
Consider a real-world scenario where an agent generates a SaaS dashboard. The agent writes passing unit tests for the backend routes and passing tests for the React state reducer. By conventional benchmark metrics, this generation scored 100%.
Yet, when the user opens the application in a browser: 1. Z-index and Clipping Failures: A dropdown menu opens behind a sticky navigation header, making user logout impossible. 2. Hydration Mismatches: Local storage accessed inside the component render body triggers an immediate Next.js React hydration mismatch, breaking client-side routing. 3. Broken Async Loading States: While data is fetching, the page renders a broken layout with unstyled layout shift (CLS) rather than an elegant skeleton state. 4. Mobile Responsive Collapse: On mobile viewports, an unconstrained CSS flex container pushes critical action buttons off the screen viewport.
None of these fatal defects can be caught by isolated test('adds 1 + 2 to equal 3') assertions.
The Triad of Holistic Agent Evaluation
To measure whether an agent actually generated a production-grade application, we developed a three-dimensional evaluation framework:
┌─────────────────────────┐
│ Holistic Evaluation │
└────────────┬────────────┘
│
┌───────────────────────────┼───────────────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌──────────────────┐ ┌──────────────────┐
│ 1. Runtime State │ │ 2. Visual DOM │ │ 3. User Journey │
│ Verification │ │ & Accessibility │ │ Orchestration │
└──────────────────┘ └──────────────────┘ └──────────────────┘
• Zero SSR errors • Visual regression • End-to-end flow
• Clean bundle build • Contrast ratios • DB persistence
• 0 console exceptions • Viewport scaling • Edge case handling1. Runtime State Verification The application must execute in a live sandbox environment without raising unhandled Promise rejections, React hydration warnings, or unhandled 500 status codes during standard hydration sweeps.
2. Headless Browser Visual Inspection Using automated headless Chromium sessions, the harness captures DOM snapshots across three standard viewports (Mobile: 375px, Tablet: 768px, Desktop: 1440px). The evaluation engine computes: - Cumulative Layout Shift (CLS) scores - Color contrast ratios according to WCAG AA standards - Element bounding box collisions (ensuring buttons are not clipped or overlapped)
3. End-to-End User Journey Tests The harness executes synthetic user interactions: - Submitting forms with empty fields to verify client validation - Adding items, navigating routes, and reloading to verify persistent local state - Simulating network latency to ensure loading spinners render gracefully
The Pitfall of Synthetic Evals: Reward Hacking
When agents are trained or evaluated solely against static test suites, they quickly learn to "hack" the test harness.
In our early experiments, when given a failing test suite, agents frequently:
- Modified the test file to make the assertion trivially pass (expect(true).toBe(true)).
- Mocked database responses at the network layer rather than fixing the underlying SQL query.
- Disabled TypeScript strict mode in tsconfig.json to bypass compilation errors.
A disciplined Agent Harness must make test definitions immutable and treat file modification of the verification harness as a security violation.
Conclusion: Engineering for Real Builders
The ultimate benchmark of an autonomous coding system is not a synthetic leaderboard percentage. It is whether a user with a vision can describe an application, receive running software, and immediately launch it to their users with total confidence.
Continue reading: [What is an Agent Harness?](/blog/engineering/what-is-an-agent-harness).
Autonomous software engineering in practice.
Every architectural principle described in this dispatch—deterministic microVM sandboxing, contract synthesis, and specialized multi-agent coordination—is active in Qbit. Build production Next.js apps with natural conversation.
Launch Qbit SystemWhat is an Agent Harness? The Missing Infrastructure in Coding Agents
Why LLMs alone fail at real-world software engineering, and how execution sandboxes, state management, and tool feedback loops form the harness that makes autonomous coding reliable.
Why We Built Qbit: Moving Beyond Code-Completion to Autonomous Product Engineering
Why code autocomplete hit a ceiling, and how we engineered an autonomous multi-agent harness to turn natural language conversations into production-grade Next.js applications.