How Karya verifies agent changes.

Karya is verification infrastructure for coding agents. It scans your software, builds a digital twin of it, and gives agents a place to run a maintained testing suite against that twin on every iteration. This page explains the technical approach behind each part.

Scanonce, then kept current
  1. Repositories and telemetry
  2. Scan
  3. Feature map
  4. Twin spec and testing suite
Verifyevery agent iteration
  1. Agent or CI submits a change
  2. Twin from the warm pool
  3. Impacted tests run
  4. Verdict

Fail: the trace goes back to the agent, which iterates. Pass: the pull request opens.

Maintainevery merge
  1. Merge
  2. Drift detection
  3. Review inbox
  4. Map and suite updated

The updated map and suite are what the next verification run uses.

Isolated ephemeral environments

A disposable digital twin for every run.

Agents can only verify a change if they can run it against something that behaves like production. Shared staging can't support hundreds of agents working in parallel, and production is off limits. Karya gives every verification run its own isolated copy of the stack: a MicroVM running a Kubernetes-based simulation of your production environment.

What a twin is made of

ServicesRun from your code at the change's commit.
Data storesSeeded from fixtures to a known state.
Queues and eventsReal messaging between services.
External APIsStubbed at the egress boundary.

Technical strategy

01
Derive the twin from the scan

Karya reads repositories and infrastructure configuration to build a topology of each project: its services, the data stores and queues they use, and the third-party APIs they call. That topology is the spec the twin is built from.

02
Simulate production inside a MicroVM

Each twin runs in its own MicroVM, which hosts a Kubernetes-based simulation of your production environment. Services are deployed the way they run in production for maximum fidelity, and the MicroVM boundary keeps every run isolated from every other.

03
Stub the edge, keep the inside real

Everything inside your system boundary runs as it does in your stack. Calls that leave it, such as payments, email or identity providers, are answered by stubs. Runs can't touch live third-party systems and don't depend on their availability.

04
Start every run from the same state

Data stores are seeded before each run, so a failing test points at the change, not at leftover data from another agent.

05
Keep a warm pool

Sandboxes are pre-warmed so an agent asking for an environment gets one without waiting for a build. This is what makes verifying on every iteration practical.

06
Let agents inspect the twin over MCP

Every twin exposes an MCP server. Agents use it to query the state of services, data stores and queues, and to learn more about failure modes, so a failed run tells them why and not only what.

07
One twin per run, then discard

Each agent iteration or CI job gets its own twin, so runs execute in parallel without contention. When the run ends the twin is torn down.

APM telemetry

Runtime context from the observability you already run.

Source code shows what software can do. Telemetry shows what it actually does: which paths customers use, what normal latency looks like and where errors cluster. Karya uses that runtime profile to decide what to test and how to judge the result.

Technical strategy

01
Read existing telemetry over MCP

Karya connects to the APM and tracing data you already collect, alongside codebase context, rather than asking you to add a new monitoring stack.

02
Build a runtime profile per service

From traces and metrics Karya records request volume, latency baselines, error rates and the dependencies each service was observed calling.

03
Confirm the topology with real traffic

Dependencies observed in traces are checked against the ones found in code, so the twin and the feature map reflect how the system really communicates.

04
Weight coverage by usage

Journeys that carry the most real traffic are prioritized when Karya generates and selects tests.

05
Judge runs against a baseline

Results from the twin are compared with the production baseline. A change that makes a critical path meaningfully slower is flagged, as well as one that breaks it.

Feature mapping

A live map from code to customer journeys.

To know whether a change is safe, you need to know who it affects. Karya maintains a map that connects each project's code to the features it provides and the critical user journeys customers take through them.

Technical strategy

01
Extract projects, features and journeys

The scan groups repositories into projects, identifies the features each one provides, and extracts the critical user journeys within them from routes, schemas, existing tests and code paths.

02
Record intent in plain language

Every journey carries a description of the behaviour it protects, written so a product manager can read it, for example: "A customer whose card expired keeps working."

03
Link journeys to the components they touch

Each journey is tied to the specific services, data stores, queues and external APIs it exercises. This is the graph shown in the Command Center topology.

04
Use the graph for impact analysis

When an agent changes code, Karya follows the graph from the changed components to the journeys that depend on them, and selects those tests to run. Reviewers see the affected features alongside the pull request.

05
Update on every merge

The map is revised as the codebase changes, so it describes the software as it is today.

Maintained testing suite

Playwright tests, written and kept current for you.

Most estates don't have the coverage agents need, and whatever exists goes stale as agents ship. Karya automatically authors every Playwright test needed to run a deep testing suite against the twin, generating the initial suite from the feature map and then maintaining it as the product changes.

How a test is expressed

A test is a customer journey: a set of conditions (who the user is and what state they're in), followed by steps that either act or verify. Each journey is written as a standard Playwright test. Karya doesn't add a test library of its own. The only addition is a deps fixture, built with Playwright's own fixture API, that gives the test the twin components its journey touches.

test("A customer whose card expired keeps working", async ({ page, deps }) => {
  const { api, stripe, pg } = deps;                                   // components of the twin this journey touches
  await pg.seed("billing/expired-card");                              // start from a known state
  await stripe.respond("charges.create", { error: "card_expired" });  // stubbed at the egress boundary

  await page.goto("/app");
  await expect(page.getByText("Update your payment method")).toBeVisible();
  await expect(await api.get("/v1/projects")).toBeOK();               // the customer can still work
});

Illustrative example of a generated Playwright test.

Where tests come from

SourceWhen it's created
Derived from scanGenerated from the feature map when Karya first scans a project.
Drafted from a pull requestWritten when an agent's change adds behaviour the suite doesn't cover.
Added by your teamWritten or edited by a product manager or engineer, as an ordinary Playwright test.

Technical strategy

01
Test behaviour, not implementation

Tests assert that a journey still behaves as it did against a recorded baseline, so they survive refactors and catch regressions in outcomes customers care about.

02
Detect drift after every merge

Karya compares the merged change with the journeys it touched. Tests that no longer match are updated, new behaviour gets new tests, and tests for removed features are proposed for retirement.

03
Keep people in the loop on what's correct

Every created, updated or removed test lands in an inbox. Product managers and engineers approve, reject or edit it before it becomes part of the suite.

04
Track each test's state

Tests are marked as existing, pending review, needing attention or pending removal, so the suite's health is always visible.

Connecting agents and CI

Verification inside the agent's loop.

Karya is called while the agent works, not after. That turns verification into a feedback loop instead of a gate at the end.

01
Agents call Karya on every iteration

Karya is exposed to coding agents as a tool over MCP. When an agent is ready to check its work, it requests a verification run for the change.

02
Failures return with a trace

A failing run returns the failing journey and its trace to the agent, which iterates until the suite passes.

03
A pull request opens only on pass

The same verification runs from CI, so nothing reaches review without a passing result attached.

This page describes Karya's approach at a design level. For details on a specific stack, deployment or integration, talk to us.

See it running against a twin of your stack.

Book a demo