Skip to content

Scaling to 100 Streamers — 3-Month Production Review ​

Audit of the Scaling to 100 Streamers plan (written 1 May 2026, when the platform had ~9 active streamers on stage) against production reality on 28 Aug 2026 — roughly three months and one real launch later. Every claim below was verified against the live cluster, the operations repo, and the codebase on the review date, not recalled from memory.

TL;DR — the plan aged well on architecture, got one load assumption 4× wrong without consequence, and its most specific prophecy (secret drift causing an outage) came true. Not one production incident in the period was a capacity incident.


Scale reached vs planned ​

392 signed-up streamers (all with verified channels), 77 beta, 33 ever paid, 29 with active deals. By the plan's unit — active streamers — we sit near the "streamer #50" waypoint. By traffic, we already passed the plan's endgame:

SignalPlan "today" (9)Plan @100Actual @ ~30 active (28 Aug)
Probe rows / day~6.5 k~45 k38–40 k — 85 % of the 100-projection
Kick chat msgs / day~3 k~865 k1.1 M — 127 % of the 100-projection

The chat assumption (20 msgs/min per live streamer) was ~4× too low — casino chats run much hotter than modelled. Nothing broke, which vindicates the plan's core judgment that the webhook receiver is comfortable and RPS-bound: the system absorbed the 100-streamer chat load with a third of the streamers.


Phase 0 scorecard — "don't ramp without this" ​

ItemVerdict
0.1 Probe deadman alert✅ Done — ViewerPollStale + IngestScrapeDown live in the data namespace (where the poll moved)
0.2 CronJob failure alerts✅ Done — CronJobNotSucceeding, RecentJobFailed, DailySummaryNotDelivered
0.3 Hand-mirrored secrets❌ Not done — and it bit exactly as predicted. The dm vf role password drifted from its mirrored secret and crashlooped the ingest tier. No reflector / external-secrets adopted. Still the top predicted-and-realized outage source.
0.4 ch_dual_write_failures counter❌ Not implemented — void chInsert(…) is still fire-and-forget with no metric. A CH outage still produces silent gaps.
0.5 Worker resource bump⚠️ Requests still 50m / 128Mi; limits (1 CPU / 512Mi) were added instead. Zero restarts observed, so no harm to date.
0.6 PG probe retention✅ Made moot by a better move — probes left the OLTP database entirely (shared data-namespace Postgres + ClickHouse, the strangler-fig migration). The bloat concern no longer applies to the app DB.

Phase 1 scorecard ​

1.6 Job substrate — the plan's biggest call, executed as designed. pg-boss on the existing Postgres (Option B, exactly as recommended over Temporal), UPDATE … RETURNING atomic claims in the viewer poll, K8s CronJobs for hourly batch. The one unfinished piece: the worker still runs replicas: 1 with node-cron for the slow tier — the pod remains a single point of failure; the anti-affinity two-replica design exists only in comments.

1.2 ScraperAPI — resolved harder than planned. The "$0 daily-poll" option proved insufficient in practice: vendors were quota-exhausted in production. ScraperAPI was removed entirely, ZenRows moved to a paid tier, and quota pollers + alerts (ScraperQuotaExhausted, ScraperQuotaNearExhaustion, ScraperUsagePollStale) now exist. Directionally right; cost estimate optimistic.

1.3 Operator inbox — solved by buy, not build. The in-house inbox the plan wanted to restructure into master/detail was deleted and replaced with TalkJS, which ships the recommended shape out of the box.

1.4 Admin streamers pagination — ❌ not done, now overdue. The plan called 100 rows "tolerable"; the panel renders 392 in one unpaginated fetch today.

1.1 Pusher / 1.5 Resend — no tier changes, nothing broke. Resend's actual problem was one the plan never imagined: deliverability (subdomain verification, DMARC hardening, the suppression list silently dropping all sends after one bounce) — not volume.


What the plan could not see ​

In three months of production, not one incident was a capacity incident. What actually hurt: the invite-limit counting bug, deal-cancel stranding pending deliveries, the migration/deploy ordering race, the ClickHouse monitoring blind (operator ACL vs Cilium SNAT), the admin CF-cookie clobber, and email deliverability. All correctness and process — the failure class a capacity plan structurally cannot cover. Meanwhile the plan's infrastructure bets (Postgres as coordinator, ClickHouse as the analytical sink, no Temporal, kill fragile vendors) all held under 4×-hotter-than-modelled chat load.


Remaining before a real 100-active ramp ​

Roughly a day of combined work, in priority order:

  1. Secrets mirroring (0.3) — the one proven bleeder. Reflector remains the smallest change.
  2. ch_dual_write_failures_total (0.4) — a one-line counter + alert; CH gaps are currently invisible.
  3. Worker replicas: 2 — the claim/SKIP LOCKED plumbing already supports it; only the deployment and anti-affinity are missing.
  4. Admin streamers pagination (1.4) — 392 rows and growing.

Phase 2 items (probe cadence knob, webhook batch inserts, CH query caching, PG read replica) remain unstarted and — on the evidence of 1.1 M chat messages/day passing without strain — correctly deprioritized. The webhook-receiver batch-insert audit (2.3) is the first one worth revisiting, since chat already exceeds its trigger threshold.

Verifluence Documentation