Customized benchmark
A cost-effective benchmark for instruction and skills
Running the full DeepSWE (or any major benchmark) to test or improve your instructions or skills is far too expensive. This tool selects the tasks most likely to detect a real effect within your budget, with expected score, price, and uncertainty estimated from published runs. The approach can work with any benchmark whose tasks are run multiple times.
DeepSWE uses a custom agent harness, so choose a higher reasoning level
here than in your own harness to make the comparison more representative.
For an inexpensive real-world starting point, select
gpt-5-6-luna at medium reasoning with a $2 budget; this
selects 9 tasks.
How to use this tool
Primary model and reasoning — choose the configuration you plan to benchmark. This choice determines which published tasks can be selected.
Comparison model and reasoning — optionally see how the selected tasks would perform on another configuration. It uses the same tasks, repetitions, and weights and never changes the plan.
Budget — set the expected spend limit for the primary model. The selector adds the strongest-ranked tasks that fit within the budget.
Iterations — use one while exploring. Increase repetitions when confirming a small or noisy effect: they reduce the chance that your result depends on one run, but fewer tasks fit the same budget.
Minimum headroom — exclude tasks whose published score leaves less than this much room to improve. Leave it at 0% unless baseline saturation is a concern.
These estimates help plan cost and task coverage; they cannot predict how much your instruction or skill change will help.
Uses the same tasks, repetitions, and weights as the primary model.
Expected score and cost
Ranked tasks
Exclude a task to remove it from consideration. Blue rows are in the current plan; red rows are excluded. Tasks are ranked by P5, the conservative low end of each correlation interval.
| Exclude | Task | CorrelationP5 ↓P50P95 | Weight | Price / run | First-to-pass |
|---|---|---|---|---|---|
| End of ranked tasks | |||||