Evaluate
Evaluation runs your agent against a set of scenarios and grades each result.- On the Evaluate page, generate scenarios from your agent’s declared behavior, or write your own.
- Start the run. Your local worker executes each scenario against the real agent.
- Review the graded results on the same page.
Improve
Improvement is an experiment loop on top of evaluation:- Rippletide proposes an improvement zone, for example a weakness in the system prompt.
- It prepares scenarios and a deterministic check, and shows you the plan.
- After you confirm, your worker runs the real agent through candidate prompt and tool variations and compares the results.