A BENCHMARK FOR MULTI-TURN CODING-AGENT USER EXPERIENCEv3.1

Lego-UXBench

Code. Interact. Evaluate.

From a first line of code to a working experience.
95 tasks that put coding agents to the test.

Actual reference implementations, rendered from the benchmark.

95Coding tasks
—Visual tasks
6Task domains
285Added verifier cases

Real tasks.
Observable outcomes.

Lego-UXBench brings together browser interfaces, interactive games, analytical workflows, and software engineering in one coding-agent evaluation suite.

Each task pairs an instruction with a bundled environment and rule-based verification. The suite includes 54 tasks restored from public CC-Bench trajectories and 41 authored tasks, and supports single- and multi-turn evaluation.

How evaluation works
01

Understand

A concrete task instruction
and bundled inputs.

02

Build

Implement the requested
behavior in its environment.

03

Verify

Check artifacts, interactions,
and deterministic outcomes.

Meet the tasks.

Explore what an agent is asked to build, fix, or analyze.

Loading the catalog…

Behavior is the benchmark.

Rule-based verifiers check more than whether a file exists.

Browser interactions

DOM and browser checks inspect visible behavior. Deterministic game APIs support repeatable state transitions and tiered scoring.

Artifacts & computations

Data and engineering tasks check generated files, numerical results, and core behavior against bundled fixtures.

Expanded in v3.1

Three added deterministic verifier cases per task. The 95 instructions, starter environments, and reference solutions stay the same as v3.

Six domains, one suite.

Task counts from the release manifest.

The v3.1 snapshot.

The catalog presents Lego-UXBench v3.1.
Release validation recorded on September 9, 2026.

Tasks95 included · 20 excluded
Verifier expansion3 added cases per task · 285 total
Static release checks95 / 95 passed
Reference validation sample13 / 13 passed

Validation results describe the source release, not a model leaderboard. Reference source and hidden tests are not part of this public catalog.

Rendered reference implementation · illustrative state

Original task instruction