Traces

Learn how the AgentIdem Nebutex SDK records structured execution traces for reads, writes, results, failures, identities, and acknowledgement state.

Traces record what happened during an AgentIdem execution.

A trace is a structured record of the operations performed by an agent, including reads, writes, results, failures, identities, and observation state.

Why traces matter

AgentIdem is designed to test what happens when an agent retries after external state may already have changed.

To understand whether that behavior is safe, AgentIdem needs to record more than whether the overall function succeeded or failed.

A trace helps answer questions such as:

  • which operations ran
  • whether each operation was a read or write
  • which arguments were used
  • whether a write completed successfully
  • whether its acknowledgement was received or lost
  • whether an operation failed
  • what logical identity a write used
  • what result or error was produced

Operation kinds

Each traced operation has a kind.

The two core operation kinds are:

READ
WRITE

A READ observes state.

A WRITE may change external state.

For example:

READ   get_payment
WRITE  refund

Operation status

A traced operation also records whether execution succeeded or failed.

The core statuses are:

SUCCESS
FAILED

A successful operation completed.

A failed operation did not complete successfully.

Observation state

AgentIdem keeps execution status separate from what the caller observed.

Observation states include:

RECEIVED
LOST
FAILED

This distinction is important for retry testing.

SUCCESS + RECEIVED

status:      SUCCESS
observation: RECEIVED

The operation completed successfully and the caller received the acknowledgement.

For a write, this means the side effect happened and the agent observed that it succeeded.

SUCCESS + LOST

status:      SUCCESS
observation: LOST

The operation completed successfully, but its acknowledgement was lost.

The side effect still happened.

The agent may observe an injected failure and retry the operation even though external state has already changed.

This is one of the central failure modes AgentIdem is designed to test.

FAILED + FAILED

status:      FAILED
observation: FAILED

The operation failed and did not complete successfully.

This is different from a successful write whose acknowledgement was lost.

Example trace

Consider an agent that reads a payment and then issues a refund.

from agentidem import read, write

@read
def get_payment(payment_id: str):
    return payments.get(payment_id)

@write(identity=lambda payment_id: payment_id)
def refund(payment_id: str):
    return payments.refund(payment_id)

def refund_agent(payment_id: str):
    payment = get_payment(payment_id)

    if payment["status"] == "paid":
        return refund(payment_id)

    return "no_refund"

A normal execution could produce a trace conceptually similar to:

READ
name: get_payment
status: SUCCESS
observation: RECEIVED

WRITE
name: refund
identity: payment-123
status: SUCCESS
observation: RECEIVED

Lost acknowledgement trace

Under a lost acknowledgement scenario, the write can still succeed:

WRITE
name: refund
identity: payment-123
status: SUCCESS
observation: LOST

If the agent retries, another successful write may appear:

WRITE
name: refund
identity: payment-123
status: SUCCESS
observation: RECEIVED

AgentIdem can then inspect the trace for repeated successful writes with the same logical identity.

Trace fields

A trace can contain operation-level information such as:

  • operation name
  • operation kind
  • arguments
  • logical identity
  • result
  • error
  • execution status
  • observation state
  • timestamps

The exact serialized representation may include additional structured metadata.

Logical identities in traces

When a write defines an identity:

@write(identity=lambda payment_id: payment_id)
def refund(payment_id: str):
    ...

the resolved identity can be recorded with the operation.

For example:

operation: refund
identity: payment-123

This helps AgentIdem distinguish repeated execution of the same logical side effect from unrelated calls to the same function.

Partial traces

A trace can still be useful when execution fails.

If the target raises an exception after some operations have already run, AgentIdem can preserve the operations recorded before the failure.

This is important because a partial execution may already contain successful writes.

A failure at the end of the agent does not mean that no side effects happened earlier.

Traced execution

Use run_traced when you want to execute a synchronous target once and record its operations without running the full fault suite.

from agentidem import run_traced

trace = run_traced(
    "example.refund_agent",
    lambda: refund_agent("payment-123"),
)

For asynchronous targets, use run_traced_async.

import asyncio

from agentidem import run_traced_async

async def main():
    trace = await run_traced_async(
        "example.async_refund_agent",
        async_refund_agent,
    )

    print(trace)

asyncio.run(main())

Traced execution failures

If the target fails during traced execution, AgentIdem can raise:

TracedExecutionError

The error should preserve the partial trace recorded before the failure.

That allows you to inspect what happened before execution stopped.

Trace serialization

Traces can be serialized to JSON.

This makes them useful for:

  • debugging
  • CI artifacts
  • later inspection
  • replay
  • automated analysis

Users should also be able to save a trace and load it later.

Replay

Recorded traces can be used with AgentIdem replay.

Replay reruns the original target and compares meaningful operation behavior.

It should ignore volatile values such as:

  • trace IDs
  • operation IDs
  • timestamps

and compare operation-level behavior such as:

  • name
  • kind
  • arguments
  • identity
  • result
  • error
  • status
  • observation

Traces and reports

A trace records the operations performed during an execution.

A test report summarizes the broader reliability-testing result.

A report can include:

  • baseline result
  • fault results
  • findings
  • invariant results
  • unsafe count
  • safety status

Traces provide the operation-level evidence behind those results.

Next steps