Skip to content

Known limitations ​

Places where the shipped code deliberately stops, usually because going further would mean the library owning something the host should own. None of these is a bug; for faults, see troubleshooting.

Process and deployment ​

No process runs for you ​

  • Boundary: noeta-runtime and noeta-sdk are libraries: no CLI, no HTTP/SSE server, no scheduler daemon. A task enqueued with no worker running just sits in the queue.
  • Workaround: run a WorkerLoop yourself or call Client.start_workers(n). examples/reference-host is the smallest host, built from the public surface only.

Multi-host needs Postgres ​

  • Boundary: several worker processes sharing one database are safe only on Postgres (appends fenced against the live lease in the same transaction, lease expiry on the database clock). SQLite and in-memory are single-host; two processes on one SQLite file is unsafe.
  • Workaround: use Postgres across machines. On one host, a worker pool is fine, and any number of clients in one process can share storage — each has its own HostConfig.queue, children inherit it, and workers never cross queues. Fixed in this release: a second Client opened on the same store used to re-run another Client's background sub-agents when it recovered at start-up; recovery now skips a sub-agent another client is still running. See ADRs multi-host lease fencing and worker queue routing.

Postgres: one connection per adapter, no pool ​

  • Boundary: each Postgres adapter (event log, dispatcher, content store) holds one connection behind a lock, so calls on one adapter queue behind each other. A dropped connection (server restart, idle kill, network reset) is reopened automatically: a standalone statement is re-sent once, and a transaction lost before its COMMIT is re-run from the start. A connection lost during COMMIT raises, because the write may or may not have landed — except an append with an idempotency key, which is retried safely. A server that stays down still fails the call.
  • Workaround: for more database throughput, run more worker processes (each opens its own connections). See ADR multi-host lease fencing.

Durability ​

Crash recovery cannot undo side effects ​

  • Boundary: after a hard kill mid-step, the interrupted attempt is sealed with a StepAttemptAbandoned marker. The step is re-driven only if everything it recorded would have run without approval; otherwise — or after 3 seals in one turn — the task is parked: suspended with an origin="system" notice listing each interrupted call and whether it completed. A crash during a human-approved tool always parks on the same approval. Recovery never re-runs a side-effectful call silently, but it cannot undo what already ran.
  • Workaround: open the parked task, check whether the listed operations applied, then type to continue (the turn restarts from the pre-attempt state) or re-approve. Normal SIGTERM does not trigger this.

Shutdown can leave a step running ​

  • Boundary: stop() waits shutdown_grace_s for the in-flight step, then abandons it. Python cannot kill the thread; it may keep writing to the event log.
  • Workaround: exit the process after an abandon. The lease expires and requeue_stale() reclaims the task. shutdown_grace_s=None (or <= 0) waits forever; a stuck step then needs kill -KILL <pid>.

Heartbeat has a ceiling ​

  • Boundary: one step holds its lease for at most heartbeat_interval × heartbeat_max (360 by default — hours in practice). Past it the lease is released and the next write fails with InvalidLease.
  • Workaround: treat a hit as a signal to inspect the task, not as recovery.

Observability ​

Reliability events are process-local ​

  • Boundary: worker signals (stale_requeued, suspended_without_wake, step_failed_retryable, heartbeat_invalid_lease, shutdown_abandoned, timers_fired, attempt_abandoned, attempt_parked, cap_terminal_reconciled, dispatcher_unavailable) go to a sink that defaults to structured logs. They are not event-log events and do not survive a restart.
  • Workaround: pass a reliability_sink that forwards them to your monitoring.

Nobody is notified when a task waits on a human ​

  • Boundary: the task suspends on a HumanResponseReceived wake condition and answer delivers the reply, but no webhook, email or inbox fires.
  • Workaround: subscribe an Observer, forward UserQuestionRequested to your own channel, and reply with answer.

Growth and cost ​

Uncatalogued models: conservative compaction, $0 pricing ​

  • Boundary: an unknown model gets a 128,000-token window and 16,384-token output cap (compaction stays on, but may run early) and a price of 0.0, so GovernanceState.cost stays zero and max_cost_usd never fires. Each is logged once.
  • Workaround: register a ModelSpec via HostConfig(extra_models={...}) or register_models from noeta.sdk.providers.

Content is reclaimed only when you ask ​

  • Boundary: nothing sweeps the content store on its own. Client.delete_task purges a task tree's events and dispatcher state but leaves the blobs, which may be shared by hash with other tasks (it refuses with reason="running" while any task in the tree holds a live lease). Superseded TaskSnapshot bodies — one per turn, each the whole task state — stay on disk until swept too. Model requests are the one thing that is never stored: LLMRequestStarted.request_ref carries the hash of the request the model saw, but the body behind it is written only with HostConfig(record_llm_requests=True), and a sweep drops it again.
  • Workaround: call Client.collect_garbage(grace_seconds=3600.0, vacuum=False) from your own schedule (a nightly job, a health check). It keeps every blob some event still references, directly or through another blob, and deletes the rest once it is older than the grace — so it is safe while turns are running. vacuum=True shrinks a sqlite file afterwards but holds the write lock for the whole rewrite; run that in a quiet window. On Postgres it runs a plain VACUUM content (space reused, not returned; VACUUM FULL is yours to schedule). Without a Client, noeta.sdk.storage.collect_garbage(event_log, content_store, ...) does the same over a stack you opened yourself. Storage you plugged in from elsewhere needs a sweep method on its content store, or the call reports reason="unsupported".

Sandbox ​

No sandbox provisioner ships ​

  • Boundary: SandboxProvider is a protocol. The only built-in provider attaches to one already-running container (from SandboxExecEnvConfig); its release is a no-op. Creating and reaping containers is the host's job.
  • Workaround: implement SandboxProvider and pass it as HostConfig.sandbox_provider. allocate returns a SandboxHandle; attach reconnects to the exec_env_ref on TaskHostBound when a task resumes. See Sandbox.

Sandbox side effects are not fenced ​

  • Boundary: container calls go over HTTP, outside the Postgres transaction that fences log writes. A worker that lost its lease (GC pause, SIGSTOP) can still reach the container — at-least-once, like a half-run host Bash. Damage stays inside that root task's own container.
  • Workaround: none automatic; the same re-drive and human review as crashed steps apply.

A stopped sandbox Bash returns no output ​

  • Boundary: each foreground command runs in its own container shell and is killed there on interrupt, cancel, close, or timeout (since 2026-09-25). A command stopped that way returns no partial output, where a local one returns what it printed so far. A process the command put in the background (server &) is not killed when the command finishes normally — the same as on the host.
  • Workaround: send long output to a file in the workspace and Read it after a stop.

Background shell is host-only ​

  • Boundary: Bash(run_in_background=true) (with BashOutput / KillShell) needs the host's background runner. A sandbox returns an error instead.
  • Workaround: run in the foreground with a generous timeout, or run outside the sandbox.

Sandbox browser is text-level ​

  • Boundary: the five tools (browser_navigate, browser_click, browser_type, browser_extract, browser_screenshot) mount only with a live browser in the container and the browser activation. browser_extract returns text plus numbered elements; browser_screenshot saves a PNG to the workspace but is not shown to the model. The browser shares the container's lifetime and cost.
  • Workaround: use browser_extract for content, WebFetch for pages that need no interaction, and screenshots for humans.

Closed extension points ​

The context composer cannot be replaced ​

  • Boundary: swapping ContextComposer would break the stable prompt prefix the provider cache depends on. Only append-only hooks are open: a ContentKindSpec resident or a compose-time reminder.
  • Workaround: use those hooks, or replace the Policy through the policy surface. See Context.

Next ​

Released under the Apache License 2.0.