Browser interactions
DOM and browser checks inspect visible behavior. Deterministic game APIs support repeatable state transitions and tiered scoring.
Code. Interact. Evaluate.
From a first line of code to a working experience.
95 tasks that put coding agents to the test.
Actual reference implementations, rendered from the benchmark.
Lego-UXBench brings together browser interfaces, interactive games, analytical workflows, and software engineering in one coding-agent evaluation suite.
Each task pairs an instruction with a bundled environment and rule-based verification. The suite includes 54 tasks restored from public CC-Bench trajectories and 41 authored tasks, and supports single- and multi-turn evaluation.
How evaluation worksA concrete task instruction
and bundled inputs.
Implement the requested
behavior in its environment.
Check artifacts, interactions,
and deterministic outcomes.
Explore what an agent is asked to build, fix, or analyze.
Loading the catalog…
Try a different keyword or clear the filters.
Visual cards show reference implementations, not agent submissions. Other cards show bundled input excerpts or requested outputs.
Rule-based verifiers check more than whether a file exists.
DOM and browser checks inspect visible behavior. Deterministic game APIs support repeatable state transitions and tiered scoring.
Data and engineering tasks check generated files, numerical results, and core behavior against bundled fixtures.
Three added deterministic verifier cases per task. The 95 instructions, starter environments, and reference solutions stay the same as v3.
Task counts from the release manifest.
The catalog presents Lego-UXBench v3.1.
Release validation recorded on September 9, 2026.
Validation results describe the source release, not a model leaderboard. Reference source and hidden tests are not part of this public catalog.