AI acceptance testing

Crazy Monkeys

Ship gate for AI

Let the monkeys stress it before your users do.

Your agent shipped a feature. Did it actually do what you told it to? Crazy Monkeys returns a versioned acceptance verdict with evidence: ship or don’t, before merge.

Built for teams and solo builders shipping AI apps, agents, and tool-calling features.

Proof you can trust

  • Baseline on your preview or PR
  • Catch tool-call loops and injection nonsense
  • Ship / no-ship keyed to the build (SHA)
  • Evidence you can keep: not a script dump

What you get

One run answers, for a specific workflow on a specific commit:

  1. What it was supposed to do
  2. What “correct” looks like
  3. Evidence (telemetry / traces / captures)
  4. Pass or fail
  5. Diff since the last run on the same step
  6. SHA the verdict was produced against

The runner underneath can change. The contract shape does not. The script is interchangeable. The ledger is the product.

What we are not

  • Not QA automation theater
  • Not an AI bot that authors your Playwright suite as the SKU
  • Not a coverage-as-a-Service staffing shop
  • Not enterprise SSO / on-prem (stubs stay stubs)

How it fits CI

Point us at a preview. We run acceptance. You get ship or don’t ship before merge. Limited beta. Waitlist open.

Sample run result
https://preview.acme-ops.dev/shipments
ACME OPS · STAGING
Shipments
Sync now
Search inventory…
SKU-118OrdersSynced
SKU-442ShipmentsQueued
SKU-903ExceptionsClear