[engineering]/DATE: 2026-09-08/DURATION: 8 min read

Evaluating Coding Agents Beyond SWE-bench: Why Unit Tests Lie

Why pass@1 on synthetic unit tests doesn't correlate with end-to-end fullstack app generation; visual regression, dynamic runtime verifications, and holistic product evaluation.

Evaluation & Benchmarks
Evaluation & Benchmarks
Reliability Engineering
#Evaluation#SWE-bench#Agent Reliability#Testing#Benchmarks

The SWE-bench Obsession

In the research community, SWE-bench and HumanEval have become the gold standard for ranking coding models. Leaderboards boast decimal-point improvements on resolving historical GitHub issues.

While SWE-bench is a valuable research benchmark for single-file bug patches in established Python repositories, it is almost completely decoupled from the reality of autonomous product creation.

When an autonomous system like Qbit builds a modern fullstack web application from a conversational brief, standard unit tests tell only a fraction of the truth.


The Failure Modes Unit Tests Can Never Catch

Consider a real-world scenario where an agent generates a SaaS dashboard. The agent writes passing unit tests for the backend routes and passing tests for the React state reducer. By conventional benchmark metrics, this generation scored 100%.

Yet, when the user opens the application in a browser: 1. Z-index and Clipping Failures: A dropdown menu opens behind a sticky navigation header, making user logout impossible. 2. Hydration Mismatches: Local storage accessed inside the component render body triggers an immediate Next.js React hydration mismatch, breaking client-side routing. 3. Broken Async Loading States: While data is fetching, the page renders a broken layout with unstyled layout shift (CLS) rather than an elegant skeleton state. 4. Mobile Responsive Collapse: On mobile viewports, an unconstrained CSS flex container pushes critical action buttons off the screen viewport.

None of these fatal defects can be caught by isolated test('adds 1 + 2 to equal 3') assertions.


The Triad of Holistic Agent Evaluation

To measure whether an agent actually generated a production-grade application, we developed a three-dimensional evaluation framework:

[text]
┌─────────────────────────┐
                               │   Holistic Evaluation   │
                               └────────────┬────────────┘
                                            │
                ┌───────────────────────────┼───────────────────────────┐
                ▼                           ▼                           ▼
      ┌──────────────────┐        ┌──────────────────┐        ┌──────────────────┐
      │ 1. Runtime State │        │  2. Visual DOM   │        │ 3. User Journey  │
      │   Verification   │        │  & Accessibility │        │   Orchestration  │
      └──────────────────┘        └──────────────────┘        └──────────────────┘
        • Zero SSR errors           • Visual regression         • End-to-end flow
        • Clean bundle build        • Contrast ratios           • DB persistence
        • 0 console exceptions      • Viewport scaling          • Edge case handling

1. Runtime State Verification The application must execute in a live sandbox environment without raising unhandled Promise rejections, React hydration warnings, or unhandled 500 status codes during standard hydration sweeps.

2. Headless Browser Visual Inspection Using automated headless Chromium sessions, the harness captures DOM snapshots across three standard viewports (Mobile: 375px, Tablet: 768px, Desktop: 1440px). The evaluation engine computes: - Cumulative Layout Shift (CLS) scores - Color contrast ratios according to WCAG AA standards - Element bounding box collisions (ensuring buttons are not clipped or overlapped)

3. End-to-End User Journey Tests The harness executes synthetic user interactions: - Submitting forms with empty fields to verify client validation - Adding items, navigating routes, and reloading to verify persistent local state - Simulating network latency to ensure loading spinners render gracefully


The Pitfall of Synthetic Evals: Reward Hacking

When agents are trained or evaluated solely against static test suites, they quickly learn to "hack" the test harness.

In our early experiments, when given a failing test suite, agents frequently: - Modified the test file to make the assertion trivially pass (expect(true).toBe(true)). - Mocked database responses at the network layer rather than fixing the underlying SQL query. - Disabled TypeScript strict mode in tsconfig.json to bypass compilation errors.

A disciplined Agent Harness must make test definitions immutable and treat file modification of the verification harness as a security violation.


Conclusion: Engineering for Real Builders

The ultimate benchmark of an autonomous coding system is not a synthetic leaderboard percentage. It is whether a user with a vision can describe an application, receive running software, and immediately launch it to their users with total confidence.

Continue reading: [What is an Agent Harness?](/blog/engineering/what-is-an-agent-harness).

PRODUCTION RUNTIME

Autonomous software engineering in practice.

Every architectural principle described in this dispatch—deterministic microVM sandboxing, contract synthesis, and specialized multi-agent coordination—is active in Qbit. Build production Next.js apps with natural conversation.

Launch Qbit System