QwenWebBenchBrowse 296

QwenWebBench

Can the model complete substantial programming work under this benchmark's agent setup?

Version not specifiedExact reported variant
Coding8 ranked models11 reported values2 reportsHigher is better

Top rankings

One row per model, using its best reported score across effort settings.

8
RankModelBest scoreBest reported setting
1Claude Opus 4.5Anthropic1536Reported configurationQwen report1 value · 1 reportApr 21, 2026 · source
2Qwen3.6-27BQwen1487Reported configurationQwen report1 value · 1 reportApr 21, 2026 · source
3Qwen3.6-35B-A3BQwen1397Reported configurationQwen report2 values · 2 reportsApr 21, 2026 · source
4Gemma4-31BGoogle1197Reported configurationQwen report2 values · 2 reportsApr 21, 2026 · source
5Qwen3.5-397B-A17BQwen1186Reported configurationQwen report1 value · 1 reportApr 21, 2026 · source
6Gemma4-26B-A4BGoogle1178Reported configurationQwen report1 value · 1 reportApr 15, 2026 · source
7Qwen3.5-27BQwen1068Reported configurationQwen report2 values · 2 reportsApr 21, 2026 · source
8Qwen3.5-35B-A3BQwen978Reported configurationQwen report1 value · 1 reportApr 15, 2026 · source

Effort curve

Every sourced cost-linked effort value for this exact version. Lines connect complete sweeps only.

0
No cost-linked effort sweep for this version.
Definition and comparison boundaryinternal methodology

Can the model complete substantial programming work under this benchmark's agent setup?

Software implementation, debugging, or repository work under the published evaluation protocol. Higher is better. This benchmark reports points rather than percent correct.

This is a publisher-defined internal evaluation. The task set or grading details are not fully public, so treat it as directional evidence.

Internal bilingual front-end code-generation benchmark with seven categories, auto-render plus multimodal judge, and BT/Elo rating.

Method / source