llm4funAI HARDWARE, EXPLAINED.
THE WORKBENCH
MODEL BENCHMARKS / CAPABILITY EXPLORER

Set your bar.
Find your model.

Choose the capabilities that matter to you. Set minimum scores and explore which models meet every requirement.

69 models · 6 benchmarks

Open-weight models · 2026-03-28 snapshot

Source, license & limitations

These are published capability scores. Evaluation settings, tool use and reasoning budgets may differ. Use them to build a shortlist; they do not predict quality at your local quantization.

02 / EXPLORE THE LANDSCAPE

Capability, at a glance.

View as table

Publisher counts show plotted / available models for your search. Plotting needs both axis scores and follows “Show matches only”.

19meet all minimums4 below a minimum46 missing a required score
Axes fit the plotted scores
Meets all minimumsBelow a minimumMissing required score
03 / INSPECT A MODEL

Qwen/Qwen3.5-397B-A17B

Qwen · Open weights · 5 of 6 scores reported

Meets all minimums

What computer can run this model?

Compare memory fit, estimated speed and cost for this exact model.

FP8 · 8K context · 30 tok/s target. Uses FP8 because the current format is not listed for this model. Your other simulation settings are kept.

MMLU-Pro87.8%Accuracy · min 60%
GPQA88.4%Accuracy · min 50%
GSM8K—Not reported
AIME 202693.33%Score
HLE28.7%Accuracy
SWE Verified76.4%Resolved
THE SCORECARD

Every score. The same requirements.

69 models in this selection. Select a model to inspect it above. Scores are percentages; a dash means not reported.

Reported model benchmark scores and threshold results. Sorted by MMLU-Pro, highest first.
ModelYour requirementsMMLU-Pro≥ 60%GPQA≥ 50%GSM8KNo minimumAIME 2026No minimumHLENo minimumSWE VerifiedNo minimum
Missing required score88%———22.2%74%
Meets all minimums87.8%88.4%—93.33%28.7%76.4%
Meets all minimums87.1%87.6%—95.83%50.2%70.8%
Meets all minimums86.7%86.6%——25.3%72%
Meets all minimums86.1%85.5%—90.83%24.3%72.4%
Meets all minimums85.3%84.2%—93.33%22.4%69.2%
Missing required score85%—————
Meets all minimums85%82.4%—94.17%40.8%70%
Meets all minimums84.6%84.5%——23.9%71.3%
Missing required score84.4%—————
Meets all minimums84.4%83.5%—96.67%23.1%74.4%
Meets all minimums84.3%85.7%——24.8%73.8%
Meets all minimums84%71.5%————
Meets all minimums83.8%79.1%——13.6%—
Meets all minimums83.73%79.23%—90%18.26%53.73%
Meets all minimums83.73%79.23%——18.26%53.73%
Meets all minimums82.5%81.7%—92.5%——
Missing required score82%———12.5%69.4%
Missing required score81.2%—————
Meets all minimums81.02%74.43%————
Missing required score80.6%—————
Meets all minimums79.8%76.1%——17.7%—
Meets all minimums79.1%76.2%————
Missing required score78.3%———15.5%—
Missing required score78.29%—————
Missing required score78.1%———10.2%—
Missing required score75.2%—————
Missing required score75.2%—————
Meets all minimums74%65.8%—82.5%——
Missing required score72.1%———11.1%—
Meets all minimums69.6%62%————
Missing required score64.4%—89.3%———
Below a minimum55.3%—————
Below a minimum48.3%30.4%84.5%———
Below a minimum29.7%11.9%————
Missing required score——79.2%———
Missing required score—————53.9%
Missing required score—————62.4%
Missing required score—————66%
Missing required score————9.92%—
Missing required score——86%———
Missing required score——79.6%———
Below a minimum—38.89%————
Missing required score———82.5%——
Missing required score—80.5%——25.2%—
Missing required score——96.8%———
Missing required score——91%———
Missing required score——85.7%———
Missing required score——86.2%———
Missing required score—85.2%——19.4%75.8%
Missing required score————39.2%—
Missing required score————31%—
Missing required score—71.2%————
Missing required score—————74.8%
Missing required score—83.8%——12.6%—
Missing required score————37.1%—
Missing required score—67.1%——5.2%47.9%
Missing required score—56.8%——4.2%37.4%
Missing required score————19.1%—
Missing required score——89.5%———
Missing required score——79.9%———
Missing required score———87.5%——
Missing required score—————70.6%
Missing required score——50.57%———
Missing required score—————52.6%
Missing required score—————42.2%
Missing required score————22.1%—
Missing required score—75.2%——14.4%59.2%
Missing required score—86%—95.83%30.5%72.8%
EVIDENCE & REUSE

Know what the numbers say.

Scores are imported from OpenEvals leaderboard-data, whose dataset card declares an MIT license. This page includes score values and model labels, with attribution and a pinned source snapshot.

Source revision d22dd54c048a · retrieved 2026-10-08. Snapshot updates are reviewed, not live.

Read this as a shortlist, not a universal ranking.

  • The source aggregates reported results. Per-result verification, run settings, tool use and exact model revisions are not included.
  • Each benchmark is a separate measure. We do not create an overall score or treat missing results as zero.
  • We publish no benchmark questions, answers, model outputs or provider logos. The dataset’s license declaration applies to this score collection; model and underlying test materials have their own terms.
What each benchmark measures
Knowledge / Accuracy %

MMLU-Pro

Broad academic knowledge and reasoning.

Prompting and reasoning budgets can differ between submissions.
Science / Accuracy %

GPQA Diamond

Graduate-level science questions.

Reported as GPQA Diamond by the source. Run settings are not included.
Math / Accuracy %

GSM8K

Grade-school mathematical reasoning.

This older benchmark has limited coverage in the snapshot.
Math / Score %

AIME 2026

Competition-level mathematical problems.

Sampling, reasoning budget and aggregation settings are not included.
Expert reasoning / Accuracy %

Humanity’s Last Exam

Difficult questions across specialist subjects.

Tool use and text-only versus multimodal settings are not identified in this snapshot.
Coding / Resolved %

SWE-bench Verified

Resolving real software issues.

This measures a model with an agent setup. Agent tools and scaffolding can substantially change the result.