docs/examples/streams-celld; full Streams coordination package remains draft. Generated from docs/streams/clawql-celld.md, Learn walkthrough (§3 celld v0.4.0 baseline; Lab 5 local smoke; Labs 1–3 for schedule + NATS today). Companion to ClawQL Streams v0.2 and the Durable Objects session contract. ClawQL-owned runtime: clawql-cellrt. Upstream: celld.dev/docs.title: ClawQL Celld Integration — Spec v0.1
ClawQL Celld Integration — Spec v0.1
Hands-on walkthrough: Streams getting started (Learn) — §3 celld reading order; Lab 5 for celld v0.4.0 local smoke; Labs 1–3 for schedule + NATS today.
Status: Draft · September 2026 · v0.1
celld baseline: v0.4.0 (2026-08-28) — pin with CELLD_VERSION=v0.4.0 on install; do not mix v0.3.x and v0.4.x in one fleet
Package surface: celld (self-hosted Durable Objects) for ClawQL Streams
Depends on: clawql-streams v0.2 · clawql-durable-objects.md · clawql-inference · clawql-core · mcp-api-adapter
Related: clawql-cellrt.md (ClawQL-owned Rust runtime) · aws-celld-burst.md (AWS burst / Karpenter / Istio ambient — draft) · celld docs · limitations · security · Cloudflare compat · denoland/celld (Apache 2.0)
0. Version baseline — celld v0.4.0
ClawQL Streams docs and labs assume celld v0.4.0 unless a page pins an older tag for a migration note.
| v0.4.0 change | ClawQL impact |
|---|---|
celld dev — local object store, no Docker/bucket |
Default path for Lab 5 and bundle iteration before fleet deploy |
| Workers KV, Queues, Workflows, R2 bindings (initial) | Wrangler bindings deploy from config; prefer KV for small session metadata, keep vault/training on separate buckets (§6) |
| Versioned peer tunnel — fetch/RPC/WebSocket bodies stream to owner; no retry after transmission starts | AgentSessionDO sidecars must use stable operation IDs on ambiguous retries; do not assume idempotent replay of streaming POST bodies |
| DO output gate — read-only output waits for prior writes | Audit WORM rows and inference fetch ordering match DO companion contract |
| Rolling deploy without node restart | celld deploy + in-place adoption; pair with POST /reload on internal listener for ops |
Health endpoint moved to /.well-known/celld/health |
Update load-balancer and K8s readiness probes (was /__celld/health) |
Operator CLI — celld cell list, celld kv *, celld queue * |
Fleet smoke: list cells after webhook traffic; inspect KV namespaces used by bindings |
| Request body limit — 1 GiB default; per-cell concurrency limit | Large webhook payloads OK; tune CELLD_MAX_CELL_REQUESTS under burst |
| EKS Pod Identity credentials | Helm on EKS can use injected AWS creds without static keys |
Upgrade from v0.3.0: stop every v0.3.0 node before starting v0.4.0 — peer tunnel protocol and large KV epoch references are incompatible; no rolling update. See celld docs — Shut down and roll out a node.
1. What this is
This document specifies how ClawQL Streams runs on celld — Deno's self-hosted, S3-backed Durable Objects runtime. celld exposes the same Workers / Durable Object JavaScript API as Cloudflare, with SQLite per cell and LTX replication to an operator-owned bucket (RPO=0).
ClawQL's decision (Streams v0.2): do not build a custom DO runtime on Node worker_threads. Adopt celld for Workers/DO API–compatible self-hosted Durable Objects; keep Cloudflare for hosted; keep Kubernetes HPA for regulated until a DO runtime is production-stable. The ClawQL-owned production runtime is clawql-cellrt (Rust + Wasmtime) — complementary to celld, not a Node rewrite.
Why celld vs build-own
| Option | Effort | API parity with CF DOs | Replication / WORM | Ops burden | Verdict |
|---|---|---|---|---|---|
| celld | Integrate + constrain bundle | High (Workers DO surface) | Built-in LTX → bucket, RPO=0 | Install ~58 MB binary; fleet via bucket leases | Adopt |
Custom Node worker_threads + SQLite |
Large (hibernation, ownership, WS, alarms) | Partial / drift-prone | Homegrown | Own failure detector, placement, backup | Do not build |
| Miniflare only | Low for CI | Dev approximation | Local | Not a production fleet | CI / unit tests only |
| Cloudflare only | Low for SaaS | Native | Platform | Vendor tenancy / pricing | Hosted path |
| K8s HPA only | Medium | Different model | Postgres / NATS | Familiar regulated ops | Regulated until celld GA |
Facts of record: Apache 2.0 · ~58 MB binary · hibernation model (resident / idle / hibernated / inactive) · vendor density hint ~1000 resident / 8 GB (re-measure before GTM) · RPO=0 LTX · one application per fleet (alpha). No ClawQL-owned $/cell-month figure — see Streams §9.1.
2. Runtime constraints
Source of truth: Cloudflare compatibility and limitations. Unknown keys/APIs fail loudly at deploy or first use.
Available (use these)
| API / capability | Notes for ClawQL |
|---|---|
| Module Workers + DO bindings | Gateway + named DO classes |
fetch / Request / Response |
Inference + webhooks + egress |
| DO SQLite storage | Sync storage ops; session + WORM rows |
setAlarm / alarm handler |
TTL, reconnect, api_poll, batch windows |
| Inbound hibernatable WebSockets | SubscriptionDO client channels |
Outbound ws: / wss: |
Stream sources (persist intent — §3) |
| JS RPC on DO stubs | Spawn / coordinate sessions |
| Web Crypto (partial) | digest, HMAC, AES-GCM, Ed25519/ECDSA sign, getRandomValues, randomUUID |
node:buffer, path, stream, assert, events, util, timers/promises |
Bundle-friendly subsets |
| Static assets | Optional admin UI from fleet bucket |
| Worker Loader (experimental) | 64 MiB code / 1 MiB env limits still apply |
Unavailable or unsafe (avoid + workaround)
| Gap | Behavior on celld | ClawQL workaround |
|---|---|---|
setInterval |
Throws | setAlarm + SQLite intent |
child_process / worker_threads |
Inert stub / not implemented | In-process MCP; fetch(clawql-inference) |
node:http(s), net, tls, dns |
Inert stubs | fetch / WebSocket only |
Cache API (caches) |
No | Inference semantic cache stays on clawql-inference |
deriveKey / deriveBits / wrap-unwrap |
Missing | Pre-derive outside DO; or HMAC/AES-GCM only |
| R2 / KV / Queues / Workflows bindings | v0.4.0+ initial support from Wrangler config | Fleet bucket remains authority; scope KV/R2 creds separately from vault sync (§6) |
Platform scheduled / cron |
No handler | setAlarm chains for cron sources |
| TLS on peer protocol | Plain HTTP + HMAC | WireGuard / Tailscale / private net; ingress TLS |
TCP sockets (cloudflare:sockets) |
Silent inert stub | Do not use; prefer HTTP/WS |
Facets / undeclared DO classes via ctx.exports |
Absent | Declare all DO classes in wrangler.json(c) |
| Multi-app fleet scheduler | One app per fleet | Separate bucket/fleet per ClawQL deployment |
3. Patterns: alarms, fetch, WebSocket reconnect
3.1 setAlarm instead of intervals
// api_poll / session TTL / reconnect backoff
await this.ctx.storage.put("alarm_intent", { kind: "api_poll", url, intervalMs });
await this.ctx.storage.setAlarm(Date.now() + intervalMs);
async alarm() {
const intent = await this.ctx.storage.get<AlarmIntent>("alarm_intent");
// do work…
await this.ctx.storage.setAlarm(Date.now() + intent.intervalMs);
}
3.2 fetch() to clawql-inference — not subprocess
Model calls never spawn a process. The DO holds a virtual key and calls the inference gateway:
const res = await fetch(env.INFERENCE_URL + '/v1/messages', {
method: 'POST',
headers: {
Authorization: `Bearer ${virtualKey}`,
'Content-Type': 'application/json',
},
body: JSON.stringify(payload),
})
PAL, provider SDKs, and credential vaults stay in clawql-inference — outside the 64 MiB cell bundle.
3.3 WebSocket reconnect via SQLite intent
Outbound sockets keep a cell resident and do not continue when the cell moves nodes (limitations). Persist intent and reconnect after activation:
await this.ctx.storage.put('ws_intent', {
reconnect: true,
sourceUrl,
authRef,
lastEventId,
backoffMs: 1000,
})
// on close / after constructor wake:
await this.ctx.storage.setAlarm(Date.now() + backoffMs)
Prefer pinning ingress for a cell to its owner node when latency matters; cross-node WS close/reconnect coverage is thinner than single-node.
4. DO architecture
Logical types match clawql-durable-objects.md; celld is the self-hosted runtime.
| DO class | Lifetime | Responsibility |
|---|---|---|
GatewayDO |
Long-lived / entry | Route webhooks and admin; resolve subscription names; issue spawn to AgentSessionDO |
SubscriptionDO |
Long-lived (hibernates) | Source connection, significance filter, config + rtpConsent, ambient buffer stats |
AgentSessionDO |
Ephemeral (per event) | One agent session + Audit / Inference / Training sidecars; self-exit |
Ingress (TLS terminator)
│
▼
Gateway Worker / GatewayDO
│
├─ idFromName("sub:" + subscriptionId) → SubscriptionDO
│
└─ on significance pass:
idFromName("sess:" + subscriptionId + ":" + eventId) → AgentSessionDO
├─ AuditSidecar → storage.put (LTX WORM)
├─ InferenceSidecar → fetch(clawql-inference)
└─ TrainingDataSidecar → RTP/OBT in SQLite → export
4.1 Naming for idempotency
Stable DO names make replay safe:
| Object | Name pattern | Effect |
|---|---|---|
| Subscription | sub:\{subscriptionId\} |
One cell per subscription |
| Session | sess:\{subscriptionId\}:\{eventId\} |
Same event + sub → same cell; second spawn is idempotent wake |
Gateway still allocates doInstanceId / virtualKeyId before first spawn and writes DO_CREATED once (guard with a spawned flag in session SQLite).
4.2 SQLite schemas (illustrative)
SubscriptionDO
| Table / key | Contents |
|---|---|
config |
prompt, significance, allowedTools, model alias, budgets, rtpConsent |
ws_intent |
reconnect fields (§3.3) |
last_event |
id, hash, timestamp |
buffer_stats |
pending counts for ambient delivery |
worm:* |
subscription-level reactive audit rows |
AgentSessionDO — same contract as DO companion:
| Table | Contents |
|---|---|
session_meta |
doInstanceId, subscriptionId, virtualKeyId, manifestId, startedAt, exitReason |
rtp_turns |
ordered RTP nodes as JSON rows |
inference_calls |
tier, tokens, cache, virtual_key_id |
tool_calls |
tool name, args hash, ATR result |
export_status |
pending / flushed / failed |
worm:* |
append-only forensic trail |
audit:ring |
latest isolate hash-chain snapshot (flushed on spawn) |
audit:seq:\{n\} |
per-seq WORM rows from the clawql-core audit ring |
5. Bundle architecture
celld deploy (esbuild)
└─ Worker + DO classes
├─ clawql-streams (router, filter, stream_* MCP) — planned package
├─ clawql-core/streams-slim (audit, cache, hash-chain) — Workers-safe entry
├─ fetch(CLAWQL_MCP_URL) → clawql-mcp search/execute/memory_* (Streamable HTTP)
└─ fetch(CLAWQL_MCP_ADAPTER_URL) → mcp-api-adapter REST POST /{tool}
env / vars ≤ 1 MiB
code ≤ 64 MiB
In-process today: AgentSessionDO imports clawql-core/streams-slim for hash-chained audit + session cache. Example: docs/examples/streams-celld.
Out-of-process today: search / execute / memory_* via Streamable HTTP MCP (CLAWQL_MCP_URL). Optional protocol-fabric REST via CLAWQL_MCP_ADAPTER_URL (POST /\{tool\} on mcp-api-adapter). Inference via fetch(INFERENCE_URL). Do not embed full clawql-api, clawql-memory, mcp-api-adapter, or the full clawql-core barrel (webmcp-draft / node:fs).
Why not “full core” inside the cell?
| Concern | Reality |
|---|---|
| Bundle size | Full clawql-core barrel is still ≪ 64 MiB — size alone is not the blocker. |
| Runtime APIs | celld/Workers make node:fs, Express/node:http, gRPC, and stdio inert or unavailable. clawql-api, vault clawql-memory, and Express mcp-api-adapter cannot run inside the isolate. |
| Release cadence | MCP catalogs, adapter surfaces, and model SDKs change independently of the cell — fetch sidecars keep them deployable without redeploying every DO. |
Demo the entire product by running the sidecars for real — not by stuffing Express into the Worker:
# Process smoke (CI + local): celld + clawql-mcp-http + mcp-api-adapter + inference stub
STREAMS_CELLD_SMOKE_REQUIRED=1 bash docs/examples/streams-celld/scripts/full-stack-smoke.sh
# Optional containers for MCP + adapter (celld still on the host):
docker compose -f docs/examples/streams-celld/docker-compose.full.yml up --build
Durable audit (two layers):
- DO bookkeeping — after each spawn, the isolate hash-chain is flushed to DO storage (
audit:ringsnapshot +audit:seq:\{n\}) so LTX can retain a snapshot across isolate restarts. The in-process ring does notloadTipfrom LTX — a new isolate starts at genesis. Treat this as session-resumption state, not compliance. - Host compliance — consequential events (
SESSION_START, …)fetchCLAWQL_AUDIT_WORM_URL→ hostclawql-auditWORMAuditTrail(tip-load / dual-ack). Seestreams-celld-evidence.md.
Still deferred: offline Workers-safe clawql-api slim.
Provider specs: ship a slim default set in the bundle; load additional OpenAPI/GraphQL specs via fetch into SQLite on first use or at subscription create — do not embed full enterprise catalogs in env.
5.1 esbuild 64 MiB CI check
# clawql streams celld bundle-check
clawql streams celld bundle-check --project docs/examples/streams-celld
# Fail the job if Worker/DO artifact size > 67108864 bytes
CI must fail closed on oversize bundles. Prefer:
- Externalize
clawql-inference(always) - Import
clawql-core/streams-slim, notclawql-core(full barrel) - Tree-shake unused providers
- Avoid Node polyfills that pull
fs/http(keepnodejs_compatforcrypto/Bufferonly) - Optional further Streams-slim cuts if Effect + api ever approach budget (open question in Streams §15)
6. Bucket layout
celld uses one fleet bucket as administrative authority (deployments, SQLite/LTX, ownership leases, peer secret). ClawQL still separates concerns:
| Bucket / prefix | Purpose |
|---|---|
s3://clawql-streams-state (fleet CELLD_BUCKET) |
celld deployments, cell SQLite + LTX (platform durability), ownership, node leases — not host clawql-audit |
| Team vault sync bucket (existing ClawQL R2/S3) | Obsidian vault / memory_sync — not the celld fleet bucket |
| Training export (optional) | RTP/OBT datasets (HF / dedicated prefix) — distinct from fleet authority |
Do not reuse fleet-bucket credentials for vault sync or public dataset upload. Scope each credential to one role (security).
7. Deployment
7.1 Install (pin v0.4.0)
CELLD_VERSION=v0.4.0 curl -fsSL https://celld.dev/install.sh | sh
celld --version # expect v0.4.0
gh attestation verify --repo denoland/celld # build attestation
Binary ~58 MB; replication is in-process (no external Litestream sidecar).
7.1b Local development (celld dev, v0.4.0+)
Before provisioning a fleet bucket, iterate with the upstream counter example:
git clone --depth 1 --branch v0.4.0 https://github.com/denoland/celld
cd celld/examples/counter
npm install # if the example ships a package.json
celld dev --port 9876
# Worker default: http://127.0.0.1:9876
celld dev uses a persistent local object store under .celld/dev (add .celld/ to .gitignore). It rebuilds on source changes without Docker or cloud credentials. The internal operator API stays on loopback only.
Readiness during fleet rollout: probe GET /.well-known/celld/health (v0.4.0+), not /__celld/health.
7.2 Configure object storage
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=auto
export S3_ENDPOINT=https://ACCOUNT_ID.r2.cloudflarestorage.com
export CELLD_BUCKET=s3://clawql-streams-state
celld uses the AWS credential chain (not ~/.aws profiles/SSO).
7.3 Deploy application
# esbuild on PATH; wrangler.json or wrangler.jsonc (not .toml)
celld deploy . \
--bucket "$CELLD_BUCKET" \
--endpoint "$S3_ENDPOINT" \
--region "$AWS_REGION"
Accepted config keys only: name, main, compatibility_date, compatibility_flags, durable_objects, migrations, assets, services, vars. Unknown keys abort deploy.
7.4 Start fleet
celld \
--bucket "$CELLD_BUCKET" \
--endpoint "$S3_ENDPOINT" \
--region "$AWS_REGION" \
--listen 0.0.0.0:8080 \
--advertise node-a.internal:8080
Add nodes with the same bucket and distinct --advertise addresses. Discovery is via bucket leases — no join command.
7.5 Diagnose and inspect (v0.4.0+)
celld diagnose \
--bucket "$CELLD_BUCKET" \
--endpoint "$S3_ENDPOINT" \
--region "$AWS_REGION"
# List Durable Object instances (after traffic creates ownership records)
celld cell list --bucket "$CELLD_BUCKET" --json | head
# KV / Queue ops when Wrangler bindings declare namespaces
celld kv list MY_NAMESPACE --bucket "$CELLD_BUCKET"
celld queue info MY_QUEUE --bucket "$CELLD_BUCKET"
Reports expired leases, bad advertise addresses, unreachable peers, auth failures, protocol skew. During rolling node updates, wait for restoring=0 on every node before stopping the next peer (celld docs).
7.6 Helm
streams:
scalingBackend: celld
celld:
enabled: true
bucket: s3://clawql-streams-state
endpoint: https://….r2.cloudflarestorage.com
region: auto
CLI wrappers: clawql streams celld install|deploy|start|diagnose|bundle-check (Streams §11).
8. Security hardening
| Control | Requirement |
|---|---|
| Peer traffic | HMAC + body signature + clock/replay — no TLS; private net or WireGuard/Tailscale |
| Public ingress | Terminate TLS at reverse proxy / mesh gateway; do not expose peer port |
| Bucket creds | One fleet bucket scope; rotate on suspicion; root of authority |
| Alpha caveat | Not safe for hostile multi-tenant; fixes on latest release only |
| Build attestation | gh attestation verify --repo denoland/celld on install |
| App auth | celld does not authenticate end users — ClawQL ATR / OIDC / virtual keys remain mandatory |
| Cell state / LTX | Operator bucket; sqlite3 for DO SQLite — compliance WORM is host clawql-audit |
Regulated tenants that need hostile multi-tenant isolation or certified controls should use scalingBackend: kubernetes until celld exits alpha.
9. Cloudflare vs celld
| Concern | Cloudflare Durable Objects | celld |
|---|---|---|
| API | Workers DO | Same core DO/Workers surface |
| State | Platform SQLite | SQLite + LTX → your bucket (RPO=0) |
| Hibernation | Native | Resident / idle / hibernated / inactive (same model) |
| Pricing | CF DO request/duration | No ClawQL $/mo yet — cite hibernation structure only (Streams §9.1) |
| Density | Platform | Vendor hint ~1000 resident / 8 GB (re-measure) |
| KV / R2 bindings | Available | Initial v0.4.0+ from Wrangler; fleet bucket still separate |
| Cron triggers | scheduled |
Use setAlarm |
| Peer / mesh | Cloudflare edge | Operator mesh; versioned peer tunnel (v0.4.0+) |
| Multi-tenant | CF accounts | One app per fleet (alpha) |
| Local dev | Miniflare / workerd | celld dev (v0.4.0+) + Miniflare unit tests |
| ClawQL inference | fetch |
fetch (identical contract) |
10. Testing
| Layer | Tooling | Purpose |
|---|---|---|
| Local dev | celld dev (v0.4.0+) |
Counter/Streams fixture without bucket; .celld/dev persistence |
| Unit / DO logic | Miniflare (or workerd) | Alarm, storage, significance, idempotent names |
| Bundle | clawql streams celld bundle-check |
Enforce ≤64 MiB (CI fail-closed) |
| Fetch clients | mcp-fetch / adapter-fetch unit scripts |
Streamable HTTP + adapter REST without celld |
| Fleet | celld diagnose · celld cell list |
Lease + peer health; enumerate cells after traffic |
| Smoke | STREAMS_CELLD_SMOKE_REQUIRED=1 bash docs/examples/streams-celld/scripts/smoke.sh |
Webhook → spawn → slim + MCP + adapter + LTX keys |
| Helm | make helm-celld-template-tests |
StatefulSet / probes / env injection (CI) |
| Security | Attestation verify in CI | Supply chain; pin CELLD_VERSION=v0.4.0 |
Evidence matrix (commands + honesty about gaps): streams-celld-evidence.md.
Do not treat Miniflare alone as production parity for LTX, peer HMAC, or cross-node WebSocket behavior.
11. Known gaps
Track against upstream celld alpha:
- TCP stub —
cloudflare:socketsconnect()is a silent inert stub; Streams must not depend on raw TCP. - WebSocket cross-node — thinner test coverage for close codes/reconnect across nodes; prefer owner-node ingress; always persist reconnect intent.
- Pressure shedding — off until safe defaults; tune
CELLD_MAX_RESIDENT_CELLS/ RSS manually. - Manual updates — installer immutable releases +
currentpointer; no auto-update agent; pinCELLD_VERSION=v0.4.0until the next qualified release. - v0.3 → v0.4 fleet cutover — full stop required; no mixed-version fleet (§0).
- One application per fleet — no multi-tenant scheduler; isolate ClawQL orgs with separate fleets/buckets.
- Crypto gaps — no
deriveKey; design around digest/HMAC/AES-GCM/sign. - Silent Node stubs — importing unimplemented
node:*may not fail; lint/banchild_process,http,netin the DO package. - Streaming fetch retry — v0.4.0 peer tunnel does not retry after body transmission starts; use stable operation IDs.
Further reading
docs/streams/clawql-streams.md— Streams Specification v0.2docs/streams/clawql-cellrt.md— ClawQL-owned Rust + Wasmtime cell runtimedocs/streams/clawql-tee.md— hardware TEE path on cellrtdocs/streams/clawql-tee-airgap-audit.md— QR air-gap audit transportdocs/streams/clawql-durable-objects.md— session / sidecar / virtual key contractdocs/inference/clawql-inference.md— virtual keys, PALdocs/mcp/mcp-api-adapter.md— MCP → APIs (out-of-process from cells today)docs/streams/streams-celld-evidence.md— evidence matrix + CI commands- celld.dev · docs · limitations · security · compat · GitHub