What does this skill do?

The Skill Benchmark Optimization Loop transforms vague requests for performance improvements into a measurable and well-defined process. It establishes a baseline, generates variants based on specific hypotheses, eliminates those that fail in terms of correctness or safety, and promotes only the best, safe, and validated implementation, with confirmation of the performance gain.

On-Demand Optimization
When a user requests that the performance of an operation be improved without specifying how, the skill turns it into a systematic process.
Comparison of Variants
Generate alternative implementations, each with a hypothesis for improvement, and evaluate them using the same input.
Hyperparameter Tuning
Recursive search for configurations with persistent storage and comparison against the previously accepted winner.
Cost and Latency Reduction
Reduce p95 latency, execution cost, or memory usage before deploying the change to production.

Usage examples

🚀 Ambitious Acceleration
Make this job 20 times faster and show me the table of tested variants.
🔍 Recursive search
Test 50 recursive variants to reduce the p95 latency of this API within a 10-minute time limit.
⚖️ Direct Comparison
Compare batch-500 and parallel-8, and tell me which one is the best in terms of measured safety.
💾 Memory Optimization
Optimize the memory usage of this operation by generating up to 10 variants and selecting the best one.

Features

Measurable baseline Define the exact operation, the metric to improve, the current value, and the search budget before doing anything.
Variant Generation Create alternative implementations where each variant tests a single hypothesis for improvement.
Promotional Page Eliminate variants that fail in terms of accuracy, reliability, or reproducibility before declaring a winner.
Delta Confirmation Rerun the baseline and the winning variant to confirm that the performance difference is real.
Persistent Record Save each run in a ledger for recursive searches and hyperparameter tuning.

Frequently asked questions

Execution time, P95 latency, execution cost, and memory usage. The skill measures the baseline, generates variants, and recommends the best safe variant.
The fastest safe variant that passes the Promotion Gate is selected: correction tests passed, performance delta explained, evident rollback, and change committed to version control.
Yes. The skill records each run in a log, compares it to the previous accepted winner, and stops the search if the improvement is within the margin of error or the budget is exceeded.
It is not required. The skill works with Claude Code or Cowork using the project's CLI (npm run, custom scripts, etc.) to run and measure the variants.
Optimization Loop with Benchmarks — Performance Measured with Claude AI

¿Prefieres escuchar el contenido? Genera la narración de audio con un clic.