HTTP-level errors
These are returned by the API before any SSE event is emitted. The response body is typically a JSON object with adetail field (for validation errors) or a plain
error message. Your HTTP client should check the response status before consuming
the stream.
Use a small wrapper such as
response.raise_for_status() (in Python’s requests)
to convert HTTP errors into exceptions before entering the streaming loop. Trying
to consume the SSE stream from a non-2xx response will silently produce zero events.
Stream-level errors
Once the HTTP response is established with a 2xx status, errors are delivered as typed messages inside the stream. Three message types matter:LLM_RETRY
The agent’s upstream LLM call hit a transient failure (rate limit, timeout,
short-lived error) and is being retried automatically. The stream will resume
without further action.
TOOL_ERROR
A specific tool failed. The tool_name field identifies which one. The agent will
continue, possibly calling a different tool or synthesizing without that source.
Most integrations log these events and only surface a user-visible warning if the
final answer is materially degraded (for example, if every search call failed).
ERROR
The request cannot continue. The stream terminates immediately after this event;
no COMPLETE will follow. Surface the error to the user and stop reading the
stream.
error field is a human-readable string suitable for logging, but not always
suitable for direct display to end users. Wrap it in your own UX-friendly message
where appropriate.
Worked example: error-aware streaming handler
The handler below distinguishes HTTP-level failures, stream-level errors, and informational retries. It logs per-event details and only raises when the request can no longer succeed.- Pre-stream HTTP errors — raised immediately so callers do not waste cycles reading an empty stream.
- In-stream recoverable events — logged at appropriate levels but never raised.
- In-stream fatal events — raised so callers can surface the failure.
- Truncated streams — detected by the absence of a
COMPLETEevent, preventing silent “empty answer” bugs.
Asynchronous workflow runs
A submitted workflow (POST /v1/workflow/execute/async) moves the failure surface
off the HTTP request. The submit returns 202 as soon as the run is accepted, so a
2xx there says nothing about whether the run succeeds. The outcome arrives as the
run’s status when you retrieve it: completed, or error with the failing ERROR
event in its events list.
Three consequences for error handling:
- A dropped connection is not a failure. Closing or losing the event stream
never stops the run. Reconnect with the
Last-Event-IDheader, or stop streaming and retrieve the result instead. Nothing is lost either way. 409means the run is in the wrong state, not that your request was malformed. On the stream and cancel endpoints it means the run has already finished, so read its result fromGET /v1/workflow/executions/{execution_id}. On submit it means theexecution_idyou asked to continue cannot be continued, because it is still in progress or has already completed. Do not retry as-is in either case.503is worth retrying, unlike most 4xx and 5xx pairs. It means the run was accepted but is not ready to attach to or cancel yet. Retry the same request after the number of seconds given inRetry-After.
Retry and backoff
For429 and 5xx HTTP responses, exponential backoff with jitter is the safe
default:
400, 401, 403, 404, or 409: those reflect client-side,
identity, or resource-state problems that will not resolve on retry.
In-stream events do not benefit from retry at the request level. LLM_RETRY and
TOOL_ERROR already represent the agent’s own internal retry behavior; the
request is doing the right thing without your help.
When to abort versus continue
Next steps
Streaming responses
Full reference for every message type the stream may emit.
Conversation continuity
Recover from a 404 on
from_checkpoint_id by resetting the thread.