Test agents

Learn how to use test_agent in the AgentIdem Nebutex SDK to run a baseline execution, inject fault scenarios, and inspect side effect safety results.

Use test_agent to run the AgentIdem reliability test suite against a synchronous Python target.

The function runs the target under normal execution first, then evaluates it under controlled fault scenarios.

Basic usage

from agentidem import test_agent

report = test_agent(
    "example.refund_agent",
    refund_agent,
)

The first argument identifies the target in the report.

The second argument is the callable AgentIdem should execute.

Example target

from agentidem import read, write

@read
def get_payment(payment_id: str):
    return payments.get(payment_id)

@write(identity=lambda payment_id: payment_id)
def refund(payment_id: str):
    return payments.refund(payment_id)

def refund_agent():
    payment = get_payment("payment-123")

    if payment["status"] == "paid":
        return refund("payment-123")

    return "no_refund"

Run the test suite:

from agentidem import test_agent

report = test_agent(
    "example.refund_agent",
    refund_agent,
)

Baseline execution

Before fault testing begins, AgentIdem can run a baseline execution.

The baseline checks whether the target works normally without injected faults.

Conceptually:

baseline
  ↓
normal execution
  ↓
result

If the baseline itself fails, AgentIdem should not treat the fault suite as if the target was meaningfully tested.

A baseline failure is represented separately in the report.

Fault scenarios

After the baseline, AgentIdem can run controlled fault scenarios.

The core scenarios are:

lost_acknowledgement
before_operation_failure
duplicate_delivery

Each scenario tests a different reliability condition.

Lost acknowledgement

A write succeeds, but its acknowledgement is lost.

Conceptually:

WRITE refund
SUCCESS

acknowledgement
LOST

retry

WRITE refund
SUCCESS

AgentIdem can then check whether the retry produced a duplicate successful write.

Before-operation failure

A failure occurs before a selected operation executes.

Conceptually:

failure injected

WRITE refund
never executed

This models a case where the side effect genuinely did not happen.

Duplicate delivery

The entire target invocation runs more than once.

Conceptually:

invocation 1
  ↓
WRITE refund
SUCCESS

invocation 2
  ↓
WRITE refund
SUCCESS

This simulates systems where the same job, webhook, queue message, or invocation may be delivered more than once.

Execution outcome and safety outcome

AgentIdem keeps execution outcome separate from safety outcome.

These are not the same thing.

For example:

execution failed
safety result: safe

can be valid if an injected failure occurred but no unsafe duplicate side effect happened.

Likewise:

execution succeeded
safety result: unsafe

can be valid if the target completed but repeated the same successful write.

An exception alone does not mean that the agent was unsafe.

Safety result

A fault result is considered safe when:

  • relevant invariants pass
  • no ERROR-level safety findings are present

A report should distinguish between:

safe
unsafe
execution succeeded
execution failed

These concepts remain separate.

Findings

AgentIdem uses structured findings to explain reliability problems discovered during testing.

A finding can include:

  • severity
  • category or type
  • message
  • supporting operation information

A duplicate successful write is an ERROR-level safety finding.

Invariants

AgentIdem can also evaluate user-defined invariants.

Invariants represent conditions that should remain true after execution.

Examples include:

  • at most one refund per payment
  • balance never becomes negative
  • order state remains valid
  • exactly one resource exists
  • a write must not occur after cancellation

Built-in duplicate detection and user-defined invariants are separate concepts.

Inspecting the report

The returned report is structured for both human inspection and automation.

A report can include:

  • target
  • baseline result
  • fault count
  • unsafe count
  • individual fault results
  • findings
  • invariant results
  • safety status

For example:

print(report)

You can also inspect report fields directly when needed.

Report serialization

AgentIdem reports can be serialized to JSON.

This is useful for:

  • CI
  • automation
  • build artifacts
  • later inspection
  • external tooling

The report format should use a versioned schema so future changes can be handled safely.

Testing a callable with arguments

If your agent requires arguments, wrap the call in another callable.

For example:

from agentidem import test_agent

report = test_agent(
    "example.refund_agent",
    lambda: refund_agent("payment-123"),
)

This gives AgentIdem a zero-argument callable to execute while still allowing your agent to receive the values it needs.

Synchronous targets

test_agent is for synchronous targets.

For asynchronous agents, use:

test_agent_async

See Test async agents.

When to use test_agent

Use test_agent when you want to:

  • run a normal baseline execution
  • inject AgentIdem fault scenarios
  • test retry behavior
  • detect duplicate successful writes
  • evaluate invariants
  • produce a structured reliability report

If you only want to record one execution without running the full fault suite, use run_traced.

Next steps