
If you build on the Claude API, the claude-api skill in Claude Code can now write an eval for your app and then improve your app against it. You start with the question you want a number for.
Step 1: build the eval
/claude-api build-eval how often does our support bot send a ticket to the right team?
- It interviews you.
- It samples test cases from your traces, tickets and codebase.
- It picks the cheapest grader that fits.
- Once you sign off, it runs a baseline.
Step 2: climb it
Then say what you want better:
/claude-api hillclimb make each ticket cheaper to handle without getting more of them wrong
It changes one thing at a time: your prompt, skills, tool descriptions, or the model and effort level. Anything that only improves the training cases gets reverted.
Why the revert rule matters
A change that only helps the cases it was tuned on is overfitting. It looks like progress on the scoreboard and fails on real tickets. Throwing those changes out means the number you end up with is one you can trust.
A good first run
With Opus 5.5 and Sonnet 5.5 both new, model is one of the levers hillclimb can pull. Sonnet 5.5 uses far fewer tokens per task, so ask whether your app holds up on it without more wrong answers, and let the eval decide.




