Skip to main content
Anpord gives you a declarative way to run evals against Claude Code, Codex and any other harness you like. You write evals, validators, judges and whatever else you need in TypeScript. These are then run in a sandbox environment of your choice. Sandboxes can also be compared against one another.
hello.eval.ts

When to use Anpord

Anpord is a great fit if you’re building a product which uses or will itself be used by harnesses like Claude Code. In this case you can’t just analyze the request/response passed to a model as there’s a lot more to consider. The harness is doing real work in an environment and the evals need to be aware of this. Another common use case would be to compare how harnesses are able to use your product. Lets say you ship a change to your MCP server, you want to be sure that the end user who uses it in claude code won’t see any regressions in performance. These are the problems that Anpord is designed to solve. You write TypeScript and you’re easily able to evaluate the performance of your product with modern harnesses.

Write your first eval

Define and run an eval.

Explore the SDK

Compose evals, mocks and validators.