Firedrill

Firedrill: simulation and testing for AI agents

Test your existing AI agent against synthetic tools and data. Find failures under different conditions before the agent handles customer work.

Website: https://firedrill.run/

Firedrill is a simulation and testing platform for action-taking AI agents. It runs synthetic versions of the tools an agent uses, with data, permissions, and supported failure conditions you control. Your agent, model, and runner remain yours. Connect the agent to Firedrill's synthetic tools instead of production accounts.

What problem does Firedrill solve?

An AI agent decides which tool to call, with which arguments, and when. A workflow can take different paths when data, permissions, responses, or tool availability change. Those variations can create thousands of scenario combinations to test. Maintaining test accounts, recreating starting data, and cleaning up after each attempt makes this difficult across multiple services.

Unit tests remain useful for individual functions. Simulations complement them by exercising an agent's decisions across a changing, stateful environment and checking the resulting actions and records.

How it works

  1. Prepare synthetic tools and data. Choose tools from the library, including Gmail, HubSpot, Stripe, and more, or create a tool for your own dependency. Configure the starting data and supported conditions.
  2. Connect your existing agent. Use the selected tools' supported interfaces. You configure and run your agent in your own application, development environment, or CI runner.
  3. Simulate and inspect. Exercise workflows under different conditions. Review captured tool calls, state changes, evidence, and check results to understand what happened.
  4. Test fixes and prevent regressions. Reset supported starting conditions, rerun the agent, and turn scenarios into repeatable checks in your PR workflow.

Tools, scenarios, drills, and runs

Control conditions and inspect results

Vary data, actor permissions, and faults supported by the selected tools. Advance virtual time for supported scheduled work, and reset starting state to try a fix. Resetting the environment does not guarantee identical model decisions.

Run history keeps results and their evidence available for inspection, subject to the applicable plan and retention settings. Inspect tool activity, check results, and saved snapshots to investigate failures.

Repeatable checks in pull requests

Run your agent in your CI runner against Firedrill's synthetic environment. Use results and evidence to review changes and catch regressions. Your repository's branch protection settings determine which checks must pass before merging.

Coverage and limits

The library lists each tool's declared operations, supported interfaces, and faults. Check that coverage before connecting an agent. Scenario counts describe the test space you configure, not a guaranteed concurrency or throughput figure. Plan allowances, concurrency, and pricing are listed on the pricing page.

The website walkthrough uses demonstration data. Its model comparisons and example results are not measured performance benchmarks.

Get started

Firedrill is the product and brand of Reload Tech Inc.