APIEval-20
Evaluate AI agent API testing, including multi-domain scenarios and error injection
This product is a task benchmark for evaluating AI agents in actual API tests, covering 20 scenarios in 7 domains. It can measure the ability to discover vulnerabilities from the architecture and payload. It provides 20 API scenarios from multiple domains such as e-commerce, payment, and authentication, including request patterns and sample payloads, challenging the generated test suite to find hidden errors. Each scenario has 3 to 8 errors implanted according to complexity classification, which can test the API's handling ability for different problems.

