Known limitations
Places where the shipped code deliberately stops, usually because going further would mean the library owning something the host should own. None of these is a bug; for faults, see troubleshooting.
Process and deployment
No process runs for you
- Boundary:
noeta-runtimeandnoeta-sdkare libraries: no CLI, no HTTP/SSE server, no scheduler daemon. A task enqueued with no worker running just sits in the queue. - Workaround: run a
WorkerLoopyourself or callClient.start_workers(n).examples/reference-hostis the smallest host, built from the public surface only.
Multi-host needs Postgres
- Boundary: several worker processes sharing one database are safe only on Postgres (appends fenced against the live lease in the same transaction, lease expiry on the database clock). SQLite and in-memory are single-host; two processes on one SQLite file is unsafe.
- Workaround: use Postgres across machines. On one host, a worker pool is fine, and any number of clients in one process can share storage — each has its own
HostConfig.queue, children inherit it, and workers never cross queues. Fixed in this release: a secondClientopened on the same store used to re-run anotherClient's background sub-agents when it recovered at start-up; recovery now skips a sub-agent another client is still running. See ADRs multi-host lease fencing and worker queue routing.
Postgres: one connection per adapter, no pool
- Boundary: each Postgres adapter (event log, dispatcher, content store) holds one connection behind a lock, so calls on one adapter queue behind each other. A dropped connection (server restart, idle kill, network reset) is reopened automatically: a standalone statement is re-sent once, and a transaction lost before its
COMMITis re-run from the start. A connection lost duringCOMMITraises, because the write may or may not have landed — except an append with an idempotency key, which is retried safely. A server that stays down still fails the call. - Workaround: for more database throughput, run more worker processes (each opens its own connections). See ADR multi-host lease fencing.
Durability
Crash recovery cannot undo side effects
- Boundary: after a hard kill mid-step, the interrupted attempt is sealed with a
StepAttemptAbandonedmarker. The step is re-driven only if everything it recorded would have run without approval; otherwise — or after 3 seals in one turn — the task is parked: suspended with anorigin="system"notice listing each interrupted call and whether it completed. A crash during a human-approved tool always parks on the same approval. Recovery never re-runs a side-effectful call silently, but it cannot undo what already ran. - Workaround: open the parked task, check whether the listed operations applied, then type to continue (the turn restarts from the pre-attempt state) or re-approve. Normal SIGTERM does not trigger this.
Shutdown can leave a step running
- Boundary:
stop()waitsshutdown_grace_sfor the in-flight step, then abandons it. Python cannot kill the thread; it may keep writing to the event log. - Workaround: exit the process after an abandon. The lease expires and
requeue_stale()reclaims the task.shutdown_grace_s=None(or<= 0) waits forever; a stuck step then needskill -KILL <pid>.
Heartbeat has a ceiling
- Boundary: one step holds its lease for at most
heartbeat_interval × heartbeat_max(360 by default — hours in practice). Past it the lease is released and the next write fails withInvalidLease. - Workaround: treat a hit as a signal to inspect the task, not as recovery.
Observability
Reliability events are process-local
- Boundary: worker signals (
stale_requeued,suspended_without_wake,step_failed_retryable,heartbeat_invalid_lease,shutdown_abandoned,timers_fired,attempt_abandoned,attempt_parked,cap_terminal_reconciled,dispatcher_unavailable) go to a sink that defaults to structured logs. They are not event-log events and do not survive a restart. - Workaround: pass a
reliability_sinkthat forwards them to your monitoring.
Nobody is notified when a task waits on a human
- Boundary: the task suspends on a
HumanResponseReceivedwake condition andanswerdelivers the reply, but no webhook, email or inbox fires. - Workaround: subscribe an
Observer, forwardUserQuestionRequestedto your own channel, and reply withanswer.
Growth and cost
Uncatalogued models: conservative compaction, $0 pricing
- Boundary: an unknown model gets a 128,000-token window and 16,384-token output cap (compaction stays on, but may run early) and a price of
0.0, soGovernanceState.coststays zero andmax_cost_usdnever fires. Each is logged once. - Workaround: register a
ModelSpecviaHostConfig(extra_models={...})orregister_modelsfromnoeta.sdk.providers.
Content is reclaimed only when you ask
- Boundary: nothing sweeps the content store on its own.
Client.delete_taskpurges a task tree's events and dispatcher state but leaves the blobs, which may be shared by hash with other tasks (it refuses withreason="running"while any task in the tree holds a live lease). SupersededTaskSnapshotbodies — one per turn, each the whole task state — stay on disk until swept too. Model requests are the one thing that is never stored:LLMRequestStarted.request_refcarries the hash of the request the model saw, but the body behind it is written only withHostConfig(record_llm_requests=True), and a sweep drops it again. - Workaround: call
Client.collect_garbage(grace_seconds=3600.0, vacuum=False)from your own schedule (a nightly job, a health check). It keeps every blob some event still references, directly or through another blob, and deletes the rest once it is older than the grace — so it is safe while turns are running.vacuum=Trueshrinks a sqlite file afterwards but holds the write lock for the whole rewrite; run that in a quiet window. On Postgres it runs a plainVACUUM content(space reused, not returned;VACUUM FULLis yours to schedule). Without aClient,noeta.sdk.storage.collect_garbage(event_log, content_store, ...)does the same over a stack you opened yourself. Storage you plugged in from elsewhere needs asweepmethod on its content store, or the call reportsreason="unsupported".
Sandbox
No sandbox provisioner ships
- Boundary:
SandboxProvideris a protocol. The only built-in provider attaches to one already-running container (fromSandboxExecEnvConfig); itsreleaseis a no-op. Creating and reaping containers is the host's job. - Workaround: implement
SandboxProviderand pass it asHostConfig.sandbox_provider.allocatereturns aSandboxHandle;attachreconnects to theexec_env_refonTaskHostBoundwhen a task resumes. See Sandbox.
Sandbox side effects are not fenced
- Boundary: container calls go over HTTP, outside the Postgres transaction that fences log writes. A worker that lost its lease (GC pause,
SIGSTOP) can still reach the container — at-least-once, like a half-run hostBash. Damage stays inside that root task's own container. - Workaround: none automatic; the same re-drive and human review as crashed steps apply.
A stopped sandbox Bash returns no output
- Boundary: each foreground command runs in its own container shell and is killed there on interrupt, cancel, close, or
timeout(since 2026-09-25). A command stopped that way returns no partial output, where a local one returns what it printed so far. A process the command put in the background (server &) is not killed when the command finishes normally — the same as on the host. - Workaround: send long output to a file in the workspace and
Readit after a stop.
Background shell is host-only
- Boundary:
Bash(run_in_background=true)(withBashOutput/KillShell) needs the host's background runner. A sandbox returns an error instead. - Workaround: run in the foreground with a generous
timeout, or run outside the sandbox.
Sandbox browser is text-level
- Boundary: the five tools (
browser_navigate,browser_click,browser_type,browser_extract,browser_screenshot) mount only with a live browser in the container and thebrowseractivation.browser_extractreturns text plus numbered elements;browser_screenshotsaves a PNG to the workspace but is not shown to the model. The browser shares the container's lifetime and cost. - Workaround: use
browser_extractfor content,WebFetchfor pages that need no interaction, and screenshots for humans.
Closed extension points
The context composer cannot be replaced
- Boundary: swapping
ContextComposerwould break the stable prompt prefix the provider cache depends on. Only append-only hooks are open: aContentKindSpecresident or a compose-timereminder. - Workaround: use those hooks, or replace the
Policythrough thepolicysurface. See Context.