Skip to main content
PlatformcelldDraft

title: ClawQL Celld Integration — Spec v0.1

ClawQL Celld Integration — Spec v0.1

Hands-on walkthrough: Streams getting started (Learn) — §3 celld reading order; Lab 5 for celld v0.4.0 local smoke; Labs 1–3 for schedule + NATS today.

Status: Draft · September 2026 · v0.1
celld baseline: v0.4.0 (2026-08-28) — pin with CELLD_VERSION=v0.4.0 on install; do not mix v0.3.x and v0.4.x in one fleet
Package surface: celld (self-hosted Durable Objects) for ClawQL Streams
Depends on: clawql-streams v0.2 · clawql-durable-objects.md · clawql-inference · clawql-core · mcp-api-adapter
Related: clawql-cellrt.md (ClawQL-owned Rust runtime) · aws-celld-burst.md (AWS burst / Karpenter / Istio ambient — draft) · celld docs · limitations · security · Cloudflare compat · denoland/celld (Apache 2.0)


0. Version baseline — celld v0.4.0

ClawQL Streams docs and labs assume celld v0.4.0 unless a page pins an older tag for a migration note.

v0.4.0 change ClawQL impact
celld dev — local object store, no Docker/bucket Default path for Lab 5 and bundle iteration before fleet deploy
Workers KV, Queues, Workflows, R2 bindings (initial) Wrangler bindings deploy from config; prefer KV for small session metadata, keep vault/training on separate buckets (§6)
Versioned peer tunnel — fetch/RPC/WebSocket bodies stream to owner; no retry after transmission starts AgentSessionDO sidecars must use stable operation IDs on ambiguous retries; do not assume idempotent replay of streaming POST bodies
DO output gate — read-only output waits for prior writes Audit WORM rows and inference fetch ordering match DO companion contract
Rolling deploy without node restart celld deploy + in-place adoption; pair with POST /reload on internal listener for ops
Health endpoint moved to /.well-known/celld/health Update load-balancer and K8s readiness probes (was /__celld/health)
Operator CLI — celld cell list, celld kv *, celld queue * Fleet smoke: list cells after webhook traffic; inspect KV namespaces used by bindings
Request body limit — 1 GiB default; per-cell concurrency limit Large webhook payloads OK; tune CELLD_MAX_CELL_REQUESTS under burst
EKS Pod Identity credentials Helm on EKS can use injected AWS creds without static keys

Upgrade from v0.3.0: stop every v0.3.0 node before starting v0.4.0 — peer tunnel protocol and large KV epoch references are incompatible; no rolling update. See celld docs — Shut down and roll out a node.


1. What this is

This document specifies how ClawQL Streams runs on celld — Deno's self-hosted, S3-backed Durable Objects runtime. celld exposes the same Workers / Durable Object JavaScript API as Cloudflare, with SQLite per cell and LTX replication to an operator-owned bucket (RPO=0).

ClawQL's decision (Streams v0.2): do not build a custom DO runtime on Node worker_threads. Adopt celld for Workers/DO API–compatible self-hosted Durable Objects; keep Cloudflare for hosted; keep Kubernetes HPA for regulated until a DO runtime is production-stable. The ClawQL-owned production runtime is clawql-cellrt (Rust + Wasmtime) — complementary to celld, not a Node rewrite.

Why celld vs build-own

Option Effort API parity with CF DOs Replication / WORM Ops burden Verdict
celld Integrate + constrain bundle High (Workers DO surface) Built-in LTX → bucket, RPO=0 Install ~58 MB binary; fleet via bucket leases Adopt
Custom Node worker_threads + SQLite Large (hibernation, ownership, WS, alarms) Partial / drift-prone Homegrown Own failure detector, placement, backup Do not build
Miniflare only Low for CI Dev approximation Local Not a production fleet CI / unit tests only
Cloudflare only Low for SaaS Native Platform Vendor tenancy / pricing Hosted path
K8s HPA only Medium Different model Postgres / NATS Familiar regulated ops Regulated until celld GA

Facts of record: Apache 2.0 · ~58 MB binary · hibernation model (resident / idle / hibernated / inactive) · vendor density hint ~1000 resident / 8 GB (re-measure before GTM) · RPO=0 LTX · one application per fleet (alpha). No ClawQL-owned $/cell-month figure — see Streams §9.1.


2. Runtime constraints

Source of truth: Cloudflare compatibility and limitations. Unknown keys/APIs fail loudly at deploy or first use.

Available (use these)

API / capability Notes for ClawQL
Module Workers + DO bindings Gateway + named DO classes
fetch / Request / Response Inference + webhooks + egress
DO SQLite storage Sync storage ops; session + WORM rows
setAlarm / alarm handler TTL, reconnect, api_poll, batch windows
Inbound hibernatable WebSockets SubscriptionDO client channels
Outbound ws: / wss: Stream sources (persist intent — §3)
JS RPC on DO stubs Spawn / coordinate sessions
Web Crypto (partial) digest, HMAC, AES-GCM, Ed25519/ECDSA sign, getRandomValues, randomUUID
node:buffer, path, stream, assert, events, util, timers/promises Bundle-friendly subsets
Static assets Optional admin UI from fleet bucket
Worker Loader (experimental) 64 MiB code / 1 MiB env limits still apply

Unavailable or unsafe (avoid + workaround)

Gap Behavior on celld ClawQL workaround
setInterval Throws setAlarm + SQLite intent
child_process / worker_threads Inert stub / not implemented In-process MCP; fetch(clawql-inference)
node:http(s), net, tls, dns Inert stubs fetch / WebSocket only
Cache API (caches) No Inference semantic cache stays on clawql-inference
deriveKey / deriveBits / wrap-unwrap Missing Pre-derive outside DO; or HMAC/AES-GCM only
R2 / KV / Queues / Workflows bindings v0.4.0+ initial support from Wrangler config Fleet bucket remains authority; scope KV/R2 creds separately from vault sync (§6)
Platform scheduled / cron No handler setAlarm chains for cron sources
TLS on peer protocol Plain HTTP + HMAC WireGuard / Tailscale / private net; ingress TLS
TCP sockets (cloudflare:sockets) Silent inert stub Do not use; prefer HTTP/WS
Facets / undeclared DO classes via ctx.exports Absent Declare all DO classes in wrangler.json(c)
Multi-app fleet scheduler One app per fleet Separate bucket/fleet per ClawQL deployment

3. Patterns: alarms, fetch, WebSocket reconnect

3.1 setAlarm instead of intervals

// api_poll / session TTL / reconnect backoff
await this.ctx.storage.put("alarm_intent", { kind: "api_poll", url, intervalMs });
await this.ctx.storage.setAlarm(Date.now() + intervalMs);

async alarm() {
  const intent = await this.ctx.storage.get<AlarmIntent>("alarm_intent");
  // do work…
  await this.ctx.storage.setAlarm(Date.now() + intent.intervalMs);
}

3.2 fetch() to clawql-inference — not subprocess

Model calls never spawn a process. The DO holds a virtual key and calls the inference gateway:

const res = await fetch(env.INFERENCE_URL + '/v1/messages', {
  method: 'POST',
  headers: {
    Authorization: `Bearer ${virtualKey}`,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify(payload),
})

PAL, provider SDKs, and credential vaults stay in clawql-inference — outside the 64 MiB cell bundle.

3.3 WebSocket reconnect via SQLite intent

Outbound sockets keep a cell resident and do not continue when the cell moves nodes (limitations). Persist intent and reconnect after activation:

await this.ctx.storage.put('ws_intent', {
  reconnect: true,
  sourceUrl,
  authRef,
  lastEventId,
  backoffMs: 1000,
})

// on close / after constructor wake:
await this.ctx.storage.setAlarm(Date.now() + backoffMs)

Prefer pinning ingress for a cell to its owner node when latency matters; cross-node WS close/reconnect coverage is thinner than single-node.


4. DO architecture

Logical types match clawql-durable-objects.md; celld is the self-hosted runtime.

DO class Lifetime Responsibility
GatewayDO Long-lived / entry Route webhooks and admin; resolve subscription names; issue spawn to AgentSessionDO
SubscriptionDO Long-lived (hibernates) Source connection, significance filter, config + rtpConsent, ambient buffer stats
AgentSessionDO Ephemeral (per event) One agent session + Audit / Inference / Training sidecars; self-exit
Ingress (TLS terminator)
    │
    ▼
Gateway Worker / GatewayDO
    │
    ├─ idFromName("sub:" + subscriptionId) → SubscriptionDO
    │
    └─ on significance pass:
         idFromName("sess:" + subscriptionId + ":" + eventId) → AgentSessionDO
              ├─ AuditSidecar  → storage.put (LTX WORM)
              ├─ InferenceSidecar → fetch(clawql-inference)
              └─ TrainingDataSidecar → RTP/OBT in SQLite → export

4.1 Naming for idempotency

Stable DO names make replay safe:

Object Name pattern Effect
Subscription sub:\{subscriptionId\} One cell per subscription
Session sess:\{subscriptionId\}:\{eventId\} Same event + sub → same cell; second spawn is idempotent wake

Gateway still allocates doInstanceId / virtualKeyId before first spawn and writes DO_CREATED once (guard with a spawned flag in session SQLite).

4.2 SQLite schemas (illustrative)

SubscriptionDO

Table / key Contents
config prompt, significance, allowedTools, model alias, budgets, rtpConsent
ws_intent reconnect fields (§3.3)
last_event id, hash, timestamp
buffer_stats pending counts for ambient delivery
worm:* subscription-level reactive audit rows

AgentSessionDO — same contract as DO companion:

Table Contents
session_meta doInstanceId, subscriptionId, virtualKeyId, manifestId, startedAt, exitReason
rtp_turns ordered RTP nodes as JSON rows
inference_calls tier, tokens, cache, virtual_key_id
tool_calls tool name, args hash, ATR result
export_status pending / flushed / failed
worm:* append-only forensic trail
audit:ring latest isolate hash-chain snapshot (flushed on spawn)
audit:seq:\{n\} per-seq WORM rows from the clawql-core audit ring

5. Bundle architecture

celld deploy (esbuild)
  └─ Worker + DO classes
        ├─ clawql-streams (router, filter, stream_* MCP) — planned package
        ├─ clawql-core/streams-slim (audit, cache, hash-chain) — Workers-safe entry
        ├─ fetch(CLAWQL_MCP_URL) → clawql-mcp search/execute/memory_* (Streamable HTTP)
        └─ fetch(CLAWQL_MCP_ADAPTER_URL) → mcp-api-adapter REST POST /{tool}
  env / vars ≤ 1 MiB
  code ≤ 64 MiB

In-process today: AgentSessionDO imports clawql-core/streams-slim for hash-chained audit + session cache. Example: docs/examples/streams-celld.

Out-of-process today: search / execute / memory_* via Streamable HTTP MCP (CLAWQL_MCP_URL). Optional protocol-fabric REST via CLAWQL_MCP_ADAPTER_URL (POST /\{tool\} on mcp-api-adapter). Inference via fetch(INFERENCE_URL). Do not embed full clawql-api, clawql-memory, mcp-api-adapter, or the full clawql-core barrel (webmcp-draft / node:fs).

Why not “full core” inside the cell?

Concern Reality
Bundle size Full clawql-core barrel is still ≪ 64 MiB — size alone is not the blocker.
Runtime APIs celld/Workers make node:fs, Express/node:http, gRPC, and stdio inert or unavailable. clawql-api, vault clawql-memory, and Express mcp-api-adapter cannot run inside the isolate.
Release cadence MCP catalogs, adapter surfaces, and model SDKs change independently of the cell — fetch sidecars keep them deployable without redeploying every DO.

Demo the entire product by running the sidecars for real — not by stuffing Express into the Worker:

# Process smoke (CI + local): celld + clawql-mcp-http + mcp-api-adapter + inference stub
STREAMS_CELLD_SMOKE_REQUIRED=1 bash docs/examples/streams-celld/scripts/full-stack-smoke.sh

# Optional containers for MCP + adapter (celld still on the host):
docker compose -f docs/examples/streams-celld/docker-compose.full.yml up --build

Durable audit (two layers):

  1. DO bookkeeping — after each spawn, the isolate hash-chain is flushed to DO storage (audit:ring snapshot + audit:seq:\{n\}) so LTX can retain a snapshot across isolate restarts. The in-process ring does not loadTip from LTX — a new isolate starts at genesis. Treat this as session-resumption state, not compliance.
  2. Host compliance — consequential events (SESSION_START, …) fetch CLAWQL_AUDIT_WORM_URL → host clawql-audit WORMAuditTrail (tip-load / dual-ack). See streams-celld-evidence.md.

Still deferred: offline Workers-safe clawql-api slim.

Provider specs: ship a slim default set in the bundle; load additional OpenAPI/GraphQL specs via fetch into SQLite on first use or at subscription create — do not embed full enterprise catalogs in env.

5.1 esbuild 64 MiB CI check

# clawql streams celld bundle-check
clawql streams celld bundle-check --project docs/examples/streams-celld
# Fail the job if Worker/DO artifact size > 67108864 bytes

CI must fail closed on oversize bundles. Prefer:

  • Externalize clawql-inference (always)
  • Import clawql-core/streams-slim, not clawql-core (full barrel)
  • Tree-shake unused providers
  • Avoid Node polyfills that pull fs / http (keep nodejs_compat for crypto / Buffer only)
  • Optional further Streams-slim cuts if Effect + api ever approach budget (open question in Streams §15)

6. Bucket layout

celld uses one fleet bucket as administrative authority (deployments, SQLite/LTX, ownership leases, peer secret). ClawQL still separates concerns:

Bucket / prefix Purpose
s3://clawql-streams-state (fleet CELLD_BUCKET) celld deployments, cell SQLite + LTX (platform durability), ownership, node leases — not host clawql-audit

| Team vault sync bucket (existing ClawQL R2/S3) | Obsidian vault / memory_sync — not the celld fleet bucket | | Training export (optional) | RTP/OBT datasets (HF / dedicated prefix) — distinct from fleet authority |

Do not reuse fleet-bucket credentials for vault sync or public dataset upload. Scope each credential to one role (security).


7. Deployment

7.1 Install (pin v0.4.0)

CELLD_VERSION=v0.4.0 curl -fsSL https://celld.dev/install.sh | sh
celld --version   # expect v0.4.0
gh attestation verify --repo denoland/celld   # build attestation

Binary ~58 MB; replication is in-process (no external Litestream sidecar).

7.1b Local development (celld dev, v0.4.0+)

Before provisioning a fleet bucket, iterate with the upstream counter example:

git clone --depth 1 --branch v0.4.0 https://github.com/denoland/celld
cd celld/examples/counter
npm install   # if the example ships a package.json
celld dev --port 9876
# Worker default: http://127.0.0.1:9876

celld dev uses a persistent local object store under .celld/dev (add .celld/ to .gitignore). It rebuilds on source changes without Docker or cloud credentials. The internal operator API stays on loopback only.

Readiness during fleet rollout: probe GET /.well-known/celld/health (v0.4.0+), not /__celld/health.

7.2 Configure object storage

export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=auto
export S3_ENDPOINT=https://ACCOUNT_ID.r2.cloudflarestorage.com
export CELLD_BUCKET=s3://clawql-streams-state

celld uses the AWS credential chain (not ~/.aws profiles/SSO).

7.3 Deploy application

# esbuild on PATH; wrangler.json or wrangler.jsonc (not .toml)
celld deploy . \
  --bucket "$CELLD_BUCKET" \
  --endpoint "$S3_ENDPOINT" \
  --region "$AWS_REGION"

Accepted config keys only: name, main, compatibility_date, compatibility_flags, durable_objects, migrations, assets, services, vars. Unknown keys abort deploy.

7.4 Start fleet

celld \
  --bucket "$CELLD_BUCKET" \
  --endpoint "$S3_ENDPOINT" \
  --region "$AWS_REGION" \
  --listen 0.0.0.0:8080 \
  --advertise node-a.internal:8080

Add nodes with the same bucket and distinct --advertise addresses. Discovery is via bucket leases — no join command.

7.5 Diagnose and inspect (v0.4.0+)

celld diagnose \
  --bucket "$CELLD_BUCKET" \
  --endpoint "$S3_ENDPOINT" \
  --region "$AWS_REGION"

# List Durable Object instances (after traffic creates ownership records)
celld cell list --bucket "$CELLD_BUCKET" --json | head

# KV / Queue ops when Wrangler bindings declare namespaces
celld kv list MY_NAMESPACE --bucket "$CELLD_BUCKET"
celld queue info MY_QUEUE --bucket "$CELLD_BUCKET"

Reports expired leases, bad advertise addresses, unreachable peers, auth failures, protocol skew. During rolling node updates, wait for restoring=0 on every node before stopping the next peer (celld docs).

7.6 Helm

streams:
  scalingBackend: celld
  celld:
    enabled: true
    bucket: s3://clawql-streams-state
    endpoint: https://….r2.cloudflarestorage.com
    region: auto

CLI wrappers: clawql streams celld install|deploy|start|diagnose|bundle-check (Streams §11).


8. Security hardening

Control Requirement
Peer traffic HMAC + body signature + clock/replay — no TLS; private net or WireGuard/Tailscale
Public ingress Terminate TLS at reverse proxy / mesh gateway; do not expose peer port
Bucket creds One fleet bucket scope; rotate on suspicion; root of authority
Alpha caveat Not safe for hostile multi-tenant; fixes on latest release only
Build attestation gh attestation verify --repo denoland/celld on install
App auth celld does not authenticate end users — ClawQL ATR / OIDC / virtual keys remain mandatory
Cell state / LTX Operator bucket; sqlite3 for DO SQLite — compliance WORM is host clawql-audit

Regulated tenants that need hostile multi-tenant isolation or certified controls should use scalingBackend: kubernetes until celld exits alpha.


9. Cloudflare vs celld

Concern Cloudflare Durable Objects celld
API Workers DO Same core DO/Workers surface
State Platform SQLite SQLite + LTX → your bucket (RPO=0)
Hibernation Native Resident / idle / hibernated / inactive (same model)
Pricing CF DO request/duration No ClawQL $/mo yet — cite hibernation structure only (Streams §9.1)
Density Platform Vendor hint ~1000 resident / 8 GB (re-measure)
KV / R2 bindings Available Initial v0.4.0+ from Wrangler; fleet bucket still separate
Cron triggers scheduled Use setAlarm
Peer / mesh Cloudflare edge Operator mesh; versioned peer tunnel (v0.4.0+)
Multi-tenant CF accounts One app per fleet (alpha)
Local dev Miniflare / workerd celld dev (v0.4.0+) + Miniflare unit tests
ClawQL inference fetch fetch (identical contract)

10. Testing

Layer Tooling Purpose
Local dev celld dev (v0.4.0+) Counter/Streams fixture without bucket; .celld/dev persistence
Unit / DO logic Miniflare (or workerd) Alarm, storage, significance, idempotent names
Bundle clawql streams celld bundle-check Enforce ≤64 MiB (CI fail-closed)
Fetch clients mcp-fetch / adapter-fetch unit scripts Streamable HTTP + adapter REST without celld
Fleet celld diagnose · celld cell list Lease + peer health; enumerate cells after traffic
Smoke STREAMS_CELLD_SMOKE_REQUIRED=1 bash docs/examples/streams-celld/scripts/smoke.sh Webhook → spawn → slim + MCP + adapter + LTX keys
Helm make helm-celld-template-tests StatefulSet / probes / env injection (CI)
Security Attestation verify in CI Supply chain; pin CELLD_VERSION=v0.4.0

Evidence matrix (commands + honesty about gaps): streams-celld-evidence.md.

Do not treat Miniflare alone as production parity for LTX, peer HMAC, or cross-node WebSocket behavior.


11. Known gaps

Track against upstream celld alpha:

  1. TCP stub — cloudflare:sockets connect() is a silent inert stub; Streams must not depend on raw TCP.
  2. WebSocket cross-node — thinner test coverage for close codes/reconnect across nodes; prefer owner-node ingress; always persist reconnect intent.
  3. Pressure shedding — off until safe defaults; tune CELLD_MAX_RESIDENT_CELLS / RSS manually.
  4. Manual updates — installer immutable releases + current pointer; no auto-update agent; pin CELLD_VERSION=v0.4.0 until the next qualified release.
  5. v0.3 → v0.4 fleet cutover — full stop required; no mixed-version fleet (§0).
  6. One application per fleet — no multi-tenant scheduler; isolate ClawQL orgs with separate fleets/buckets.
  7. Crypto gaps — no deriveKey; design around digest/HMAC/AES-GCM/sign.
  8. Silent Node stubs — importing unimplemented node:* may not fail; lint/ban child_process, http, net in the DO package.
  9. Streaming fetch retry — v0.4.0 peer tunnel does not retry after body transmission starts; use stable operation IDs.

Further reading