Use test_agent to run the AgentIdem reliability test suite against a synchronous Python target.
The function runs the target under normal execution first, then evaluates it under controlled fault scenarios.
Basic usage
from agentidem import test_agent
report = test_agent(
"example.refund_agent",
refund_agent,
)
The first argument identifies the target in the report.
The second argument is the callable AgentIdem should execute.
Example target
from agentidem import read, write
@read
def get_payment(payment_id: str):
return payments.get(payment_id)
@write(identity=lambda payment_id: payment_id)
def refund(payment_id: str):
return payments.refund(payment_id)
def refund_agent():
payment = get_payment("payment-123")
if payment["status"] == "paid":
return refund("payment-123")
return "no_refund"
Run the test suite:
from agentidem import test_agent
report = test_agent(
"example.refund_agent",
refund_agent,
)
Baseline execution
Before fault testing begins, AgentIdem can run a baseline execution.
The baseline checks whether the target works normally without injected faults.
Conceptually:
baseline
↓
normal execution
↓
result
If the baseline itself fails, AgentIdem should not treat the fault suite as if the target was meaningfully tested.
A baseline failure is represented separately in the report.
Fault scenarios
After the baseline, AgentIdem can run controlled fault scenarios.
The core scenarios are:
lost_acknowledgement
before_operation_failure
duplicate_delivery
Each scenario tests a different reliability condition.
Lost acknowledgement
A write succeeds, but its acknowledgement is lost.
Conceptually:
WRITE refund
SUCCESS
acknowledgement
LOST
retry
WRITE refund
SUCCESS
AgentIdem can then check whether the retry produced a duplicate successful write.
Before-operation failure
A failure occurs before a selected operation executes.
Conceptually:
failure injected
WRITE refund
never executed
This models a case where the side effect genuinely did not happen.
Duplicate delivery
The entire target invocation runs more than once.
Conceptually:
invocation 1
↓
WRITE refund
SUCCESS
invocation 2
↓
WRITE refund
SUCCESS
This simulates systems where the same job, webhook, queue message, or invocation may be delivered more than once.
Execution outcome and safety outcome
AgentIdem keeps execution outcome separate from safety outcome.
These are not the same thing.
For example:
execution failed
safety result: safe
can be valid if an injected failure occurred but no unsafe duplicate side effect happened.
Likewise:
execution succeeded
safety result: unsafe
can be valid if the target completed but repeated the same successful write.
An exception alone does not mean that the agent was unsafe.
Safety result
A fault result is considered safe when:
- relevant invariants pass
- no
ERROR-level safety findings are present
A report should distinguish between:
safe
unsafe
execution succeeded
execution failed
These concepts remain separate.
Findings
AgentIdem uses structured findings to explain reliability problems discovered during testing.
A finding can include:
- severity
- category or type
- message
- supporting operation information
A duplicate successful write is an ERROR-level safety finding.
Invariants
AgentIdem can also evaluate user-defined invariants.
Invariants represent conditions that should remain true after execution.
Examples include:
- at most one refund per payment
- balance never becomes negative
- order state remains valid
- exactly one resource exists
- a write must not occur after cancellation
Built-in duplicate detection and user-defined invariants are separate concepts.
Inspecting the report
The returned report is structured for both human inspection and automation.
A report can include:
- target
- baseline result
- fault count
- unsafe count
- individual fault results
- findings
- invariant results
- safety status
For example:
print(report)
You can also inspect report fields directly when needed.
Report serialization
AgentIdem reports can be serialized to JSON.
This is useful for:
- CI
- automation
- build artifacts
- later inspection
- external tooling
The report format should use a versioned schema so future changes can be handled safely.
Testing a callable with arguments
If your agent requires arguments, wrap the call in another callable.
For example:
from agentidem import test_agent
report = test_agent(
"example.refund_agent",
lambda: refund_agent("payment-123"),
)
This gives AgentIdem a zero-argument callable to execute while still allowing your agent to receive the values it needs.
Synchronous targets
test_agent is for synchronous targets.
For asynchronous agents, use:
test_agent_async
See Test async agents.
When to use test_agent
Use test_agent when you want to:
- run a normal baseline execution
- inject AgentIdem fault scenarios
- test retry behavior
- detect duplicate successful writes
- evaluate invariants
- produce a structured reliability report
If you only want to record one execution without running the full fault suite, use run_traced.

