Skip to main content

Operations

Health and readiness​

Control exposes these endpoints on its metrics port (default 9090):

  • /healthz: process liveness;
  • /readyz: service readiness, including Redis reachability in HA mode;
  • /metrics: Prometheus text format.

Egress exposes /healthz, /readyz and /metrics on its health port (default 8090). A worker has one listener; there is no separate egress metrics port.

Metrics​

Prometheus metrics use bounded labels. Counters are cumulative process-lifetime values; gauges are instantaneous. none means the metric has no variable labels.

MetricTypeLabelsUnitProfileInterpretation
straw_requests_totalcountererror_coderequestsallcompleted Control requests; empty code is success
straw_request_duration_secondshistogramnonesecondsallend-to-end Control request latency
straw_active_requestsgaugenonerequestsallrequests currently executing in this Control
straw_routing_duration_secondshistogramnonesecondsallworker-selection latency
straw_assignment_duration_secondshistogramnonesecondsallassignment request/ack latency
straw_nats_request_duration_secondshistogramnonesecondsallNATS request/reply latency
straw_nats_errors_totalcountererror_codeerrorsallNATS transport failures by stable error code
straw_worker_sessionsgaugenonesessionsallregistered live worker sessions
straw_workers_availablegaugenoneworkersallworkers currently eligible for assignment
straw_worker_heartbeat_age_secondsgaugenonesecondsallage of the stalest registered heartbeat; rising age indicates worker/NATS trouble
straw_runtime_state_availablegaugenoneboolean (0/1)HA1 only while Redis coordination is reachable
straw_runtime_state_operations_totalcounternoneoperationsHAshared Redis coordination operations attempted
straw_runtime_state_errors_totalcounternoneerrorsHAfailed shared Redis coordination operations
straw_receipts_created_totalcounternonereceiptsreceiptsdurable receipts created
straw_receipt_parts_uploaded_totalcounternonepartsreceiptsparts uploaded or replaced
straw_receipts_verified_totalcounternonereceiptsreceiptsreceipts passing size and checksum verification
straw_receipts_rejected_totalcounternonereceiptsreceiptsreceipts rejected for size or checksum mismatch
straw_receipt_assignments_totalcounternoneassignmentsreceiptsassignment-scoped receipt references issued
straw_receipts_consumed_totalcounternonereceiptsreceiptsrequest receipts consumed successfully
straw_receipts_expired_totalcounternonereceiptsreceiptsexpired receipts removed by cleanup

Egress workers publish their own series. They are prefixed straw_egress_ because both services are usually scraped into one Prometheus, and a worker measures a different thing from Control even where the name would be the same.

MetricTypeLabelsUnitProfileInterpretation
straw_egress_readygaugenoneboolean (0/1)allthe same readiness /readyz reports for this worker
straw_egress_sessions_activegaugenonesessionsall1 while a registered session is being served; a worker serves one at a time
straw_egress_active_requestsgaugenonerequestsallassignments currently executing on this worker
straw_egress_concurrency_limitgaugenonerequestsalladmission ceiling for the live session; compare active requests against it for saturation
straw_egress_assignments_totalcounteroutcomeassignmentsallassignments this worker finished, success or error
straw_egress_request_duration_secondshistogramnonesecondsallworker-side duration of decoded outbound requests
straw_egress_bytes_totalcounterdirectionbytesalldecoded body bytes out to the upstream and in from it; raw tunnel traffic is not included
straw_egress_upstream_errors_totalcountercodeerrorsallfailed assignments by the same canonical error code Control reports
straw_egress_nats_errors_totalcountercodeerrorsallasynchronous NATS failures seen by the worker's client

A worker cannot see what it never receives: assignments rejected for capacity or draining are counted by Control as straw_requests_total{error_code="executor_capacity_exhausted"}, and raw CONNECT tunnel volume is streamed by the worker SDK and is absent from straw_egress_bytes_total.

Upstream CONNECT failures use the existing implemented worker series straw_egress_upstream_errors_total{code="upstream_proxy_failure"}; the corresponding Control series is straw_requests_total{error_code="upstream_proxy_failure"}. Do not query straw_egress_requests_total{error_code="upstream_proxy_failure"}: that metric name is not exported by this runtime. Neither upstream proxy profile IDs nor provider session IDs are metric labels, and they must not be added as labels in deployment recording rules.

Expose metrics only to your monitoring network. No telemetry database is required.

Scrape and alert examples​

Scrape the Control metrics port (9090 by default) from the monitoring network. For HA, add one target per Control instance; the load balancer's API listener is not the metrics endpoint:

scrape_configs:
- job_name: straw-control
metrics_path: /metrics
static_configs:
- targets: ["control.example.internal:9090"]
- job_name: straw-egress
metrics_path: /metrics
static_configs:
- targets: ["egress.example.internal:8090"]

The checked-in starter rules are deliberately small. Their PromQL and windows are:

AlertExpressionForSuggested response
StrawNoAvailableWorkersstraw_workers_available == 02mCheck worker readiness, pool eligibility, and NATS before adding capacity.
StrawNATSErrorsrate(straw_nats_errors_total[5m]) > 05mInspect NATS health, credentials, reconnects, and payload limits.
StrawHAStateUnavailablestraw_runtime_state_available == 01mRestore Redis coordination before admitting HA traffic.
StrawReceiptRejectionsincrease(straw_receipts_rejected_total[15m]) > 0noneInspect declared size/checksum, part uploads, storage, and clock/credential errors.
StrawEgressSaturatedstraw_egress_active_requests / straw_egress_concurrency_limit > 0.9 and straw_egress_concurrency_limit > 05mAdd workers or raise egress.capabilities.max_concurrency once upstream latency is ruled out.
StrawEgressUpstreamErrorsratio of straw_egress_upstream_errors_total to straw_egress_assignments_total over 5m > 0.2510mSplit by code to separate DNS, TLS, reset, and policy denials from upstream outages.
StrawEgressNotReadystraw_egress_ready == 05mCheck the worker's NATS reachability, credentials, and registration rejections.

Copy and adapt deploy/monitoring/prometheus-alerts.yml; choose notification routing and objective thresholds in your monitoring system. These rules do not replace a production alert policy.

Logs​

Control and Egress write structured JSON to stdout. Collect container stdout with the logging system already used by your environment. Request IDs connect client errors, Control logs, and worker activity. Common fields are timestamp, level, msg, and service; event-specific bounded fields include request_id, worker_id, addr, and durations. The startup NATS url field is credential-redacted. Never emit bearer tokens, service credentials, upstream or signed receipt URLs, headers, or bodies. A safe escalation bundle contains versions, profile, redacted config shape, health/readiness, metric names/values, and request IDs only.

For upstream proxies, safe event fields are the bounded profile ID, fixed failure fact, numeric CONNECT status, phase, and bounded duration. Startup validation may include endpoint host and port. Never log a rendered provider username, password, Proxy-Authorization, provider session ID, raw sticky ID, full destination URL/query, or proxy response header value. Do not enable credential/session debug logging while troubleshooting; use the fixed facts and status described in Troubleshooting.

A representative startup event is:

{"timestamp":"2026-07-14T00:00:00Z","level":"INFO","msg":"listening","service":"control","addr":"0.0.0.0:8080"}

Treat event-specific fields and messages as diagnostic context rather than a stable machine API; use metric names and documented error codes for automation.

Suggested service indicators and alerts​

Track successful request ratio, p95/p99 latency, assignment availability, worker saturation, NATS errors, Redis availability for HA, and receipt rejection/expiry. Choose objectives from your traffic and upstreams; 99.9% success for eligible requests is an example, not a promise. Starter alerts live in deploy/monitoring/prometheus-alerts.yml and are intentionally not a universal production policy.

Scaling and shutdown​

Scale Egress workers for outbound concurrency. A single Control is the simplest deployment. For Control HA, place at least two Redis-backed Controls behind a readiness-aware load balancer. On SIGTERM, Control fails readiness and gives active HTTP requests up to the configured maximum request timeout plus five seconds to finish before draining NATS.

Operate executor pools​

Create or update pools in the runtime snapshot before adding workers that claim them. Verify that each worker's pools, executor type, tags, countries, regions, and IP types match the intended pool. A disabled pool is a safe cutover control: it stops new assignments without cancelling requests already running. Re-enable it only after the worker rollout and capability claims are verified. If a pool has no eligible worker, requests matching its rule return route_unavailable; a full eligible fleet returns executor_capacity_exhausted.

Verify a routing rollout through all enabled ingress modes: send one REST request, one absolute-form proxy request, and one CONNECT request with equivalent hints and confirm they select the intended pool/capabilities. For sticky sessions, repeat through the same Control and through each HA Control instance; the selected worker should remain pinned until the rule TTL expires or the configured sticky fallback is exercised. Destination-policy denials must remain consistent across all three modes.

Keep pool definitions and official-worker allowed_pools in the same deployment trust boundary. Pool membership is not tenant authorization, and Straw does not provide cross-deployment or per-user pool permissions; run a separate deployment when isolation is required.

Roll out upstream proxy pools​

Protocol minor 2 requires a Control-first rollout because an old Control rejects minor-2 workers, silently drops the new pool object while decoding snapshots, and can strip proxy capability from shared worker rows. Keep all active pools direct throughout the version transition:

  1. Release the canonical protocol and v0.4.0 Go/Python bindings, then the protocol-coupled SDKs.
  2. Deploy the new Control binary first while existing minor-0/minor-1 direct workers continue serving direct pools.
  3. Remove every old Control instance. After the last old shared-state writer stops, expire/delete shared worker rows and force workers to register again.
  4. Deploy protocol-minor-2 official workers against direct pools and verify fresh registration with the upgraded Controls.
  5. Add fresh disabled proxy pools, roll the intended workers with worker-local upstream_proxies and exact claims, verify registry eligibility/profile parity, then add and enable one canary route.

Canary REST, absolute-form HTTP proxy, and raw CONNECT paths. Validate exit country/type, same-pool worker fallback, provider session retention, proxy error rate, and direct-pool regressions without logging session or credential values. Do not claim exact-IP persistence; effective affinity remains bounded by both the active Straw route pin and provider retention.

Rollback by disabling proxy routes or pools and selecting an existing/fresh direct pool. Never mutate a proxy pool into a direct pool under the same ID, because that obscures resolution and affinity semantics and can make stale claims unsafe. Worker-local profiles may remain during a coordinated rollback window, but ordinary startup validation rejects unused profiles.

Upgrades​

Read CHANGELOG.md, test the new version against a staging deployment, and keep protocol-coupled custom workers pinned until verified. For protocol minor 2, follow the Control-first sequence above; do not apply the older generic Egress-first sequence.

When the runtime-administration profile is enabled, include the NATS JetStream configuration bucket in backup and recovery drills. Inspect rollout status after upgrades; an official worker reports applied after receiving the current snapshot. See Runtime administration.

When receipt storage is enabled, back up durable receipt records and verified bodies according to their explicit retention, monitor rejection/expiry counters, and test cleanup plus interrupted-upload recovery. The S3 bucket or local volume is optional application data; NATS and Redis never contain receipt bodies.

The owned local backup/restore drills are make state-backup-smoke PROFILE=admin and make state-backup-smoke PROFILE=receipts. They stop a uniquely named disposable stack, archive its named volume, delete/recreate the volume, restore it, restart services, and verify runtime configuration or receipt content. Never point these commands at shared or production Compose projects. Production backup tooling must provide equivalent application-consistent snapshots, encryption, retention, and restore verification.

Run make ha-smoke for the owned HA failure drill. It creates a uniquely named disposable stack, scales Egress to two workers, verifies service through one Control loss, confirms both Controls become unready during a Redis outage and recover afterward, then stops one worker gracefully and verifies requests continue. It removes all namespaced containers, networks, and volumes on exit.