Wallaby for coding agents

AI tests that catch real bugs

AI can write low-value tests that pass without checking much. Wallaby helps your agent strengthen assertions and verify that tests catch regressions.

Get started

  1. Install the skills

    In your terminal

    Run in your project and choose your coding agent.

    npx skills add wallabyjs/skills

  2. Give it a task

    In your agent's chat

    Ask your agent to improve an existing test suite.

    /wallaby improve all unit tests in the project

Make your agent write high-value tests

Write tests

Your agent tests expected behavior and edge cases, then verifies its assertions.

implement free shipping for orders of $50 or more as described in <GitHub issue URL> and /wallaby write unit tests

Your agent without Wallaby

Weak Duplicate Missing

Your agent with Wallaby

Strengthened Added

 Your agent strengthens assertions, covers missing cases, and removes redundancy.

Improve tests

Your agent strengthens weak assertions, adds missing cases, and removes duplicate checks.

/wallaby improve tests for src/shipping, focusing on weak assertions and missing edge cases

Your agent without Wallaby

Weak Duplicate Missing

Your agent with Wallaby

Strengthened Added

 Your agent strengthens useful tests and removes weak or duplicate checks.

Up to 40% better results with Wallaby. Measured in our internal Wallaby Bench evals.*

Let your agent check what each test does

Per-test expression coverage

Your agent finds untested code.

Your agent sees which expressions each test reaches. When needed, it targets mutation testing at plausible bugs to check the strength of its assertions.

Execution traces

Your agent follows the test's path.

Your agent follows execution across files to check the intended behavior. When needed, traces help explain why a test caught or missed a mutation.

Runtime values

Your agent inspects values.

Your agent inspects actual values when a result is unclear. If a meaningful mutation survives, it can strengthen the assertion and repeat the mutation check.

Live test results

Your agent checks each change.

Wallaby automatically reruns affected tests. Your agent can introduce one small source mutation, check whether the tests kill it, then restore the source and confirm they pass again.

More about Wallaby's test feedback

My agent already writes and runs tests. Why Wallaby?

Writing and running tests can still leave weak assertions or tests that mirror the implementation. Wallaby helps your agent focus on high-value tests: checks of intended behavior that would fail on plausible bugs.

In our internal Wallaby Bench evals, agents using Wallaby and its skills scored up to 40% better than agents running tests through Vitest or Jest alone. This combined score covers code and test quality, speed, and token efficiency.

The evals use repeated runs, hidden checks, mutation testing, and human or AI grading against predefined criteria.

My tests pass and coverage is high. Can Wallaby still help?

Yes. Checking that a shipping fee is defined can pass even when the amount is wrong. Wallaby helps your agent check the expected amount and test boundary cases.

Will this fill my project with unnecessary tests?

Your agent checks what is already covered, adds missing cases, and can remove low-value tests that duplicate other checks while preserving distinct behavior.

Can I use Wallaby in my agent loops or software factory?

Yes. Your agent can use /wallaby write and /wallaby improve in your workflows, or use the wallaby-cli skill directly to run tests and inspect coverage, traces, and runtime values.

Which coding agents can I use?

Agents that support skills, including Claude Code, Codex, GitHub Copilot, Cursor, OpenCode, and Pi. Your agent can start Wallaby in the background or connect to an existing editor instance.

Can I use a container or cloud agent?

Yes. During beta, containers and cloud agents require a CLI token, including for existing license holders. A licensed Wallaby extension includes local CLI access on the same machine without a token.

Container and cloud setup
What does a sandboxed agent need?

The agent needs write access to ~/.wallaby, the required network access, and permission to connect to the local Wallaby process.

Sandbox configuration
Which languages and testing frameworks are supported?

Wallaby supports JavaScript and TypeScript with Jest, Vitest, node:test, Mocha, Jasmine, AVA, QUnit, and Karma with Jasmine or Mocha.

Wallaby for Python is in beta, with support for pytest and unittest.

Supported frameworks

Try it on your tests

Ask your agent to improve the tests for your latest changes.

New to Wallaby?

Free local CLI access during beta.

Run this in your terminal, enter your email, and follow the activation link.

npx --package @wallabyjs/cli wallaby access

Then install the skills

Cloud or containers?

A CLI token is required, even with an existing license. Generate it with the same access command.

Cloud and container setup

Your existing license includes Wallaby for coding agents on this machine, with no extra activation or token needed.