# CodeRunner durable queue — design brief

**Status:** design agreed, not implemented. The FPM pool half shipped in `eecfbb1`.
**Written:** 2026-08-28, from the analysis of the 2026-08-24 incident.
**Revised:** 2026-08-28 — every open decision closed, §4 expanded, and three corrections to
the first draft (see §13). Nothing in here is now waiting on a ruling; §15 is the build order.
**Audience:** whoever implements this. It is self-contained — you do not need the
conversation that produced it.

---

## 1. Why

CodeRunner executions are customer-authored DSL scripts fired by Businessmap business
rules. They call the Businessmap API, and `BusinessmapConnect::send()` answers a 429 with
`sleep($retry_after + 5)` and re-sends **recursively, with no attempt cap**
([BusinessmapConnect.php:270-280](../app/Controllers/Components/BusinessmapConnect.php#L270-L280)).
A sleeping worker holds its slot while consuming no CPU, so one API key hitting its rate
limit can occupy every worker in whatever pool it lands in.

On 2026-08-24 that happened three times. In the 12:08 window, **143 of 309 `runAsync`
requests (46%) died at the 300 s wall**, holding ~24 of 40 `[www]` workers continuously for
executions that were never going to succeed. `[www]` is shared with the portal backend,
devportal and the ELB health check, so `healthcheck.php` degraded to a 191 s p95 and the ALB
pulled the instance. **A customer's rate limit took the box offline.**

There is a second-order effect that matters for sizing: **the concurrency is causing the
rate limits.** At peak those executions were competing against each other for the same key's
quota. Cap concurrency and the API pressure drops roughly in proportion, so most executions
finish at their healthy service time instead of sleeping at all.

## 2. What already exists (do not rebuild)

`eecfbb1` added a dedicated `codeRunner` FPM pool — `deploy/php-fpm/codeRunner.conf` and
`deploy/apache/businessmap-sa-coderunner.conf`, both with the reasoning inline. Read those
two files before touching anything.

**The pool is a hard ceiling, not the concurrency policy.** It protects the health check and
the portal even if this queue has a bug, or someone POSTs `/codeRunner/runAsync` directly.
`pm = ondemand`, `pm.max_children = 8` — double the queue's target of 4, so it never binds
in normal operation but stops a runaway. It should never be reached once the queue owns
concurrency; if it is, that is a queue defect.

Routed to the pool: `/codeRunner` and `/codeRunner/runAsync` only. `/codeRunner/logs`,
`/logs/data`, `/test` and `/internal/codeRunner` deliberately stay on `[www]` so the log
viewer and the debugger's linter stay responsive while executions are queued.

The kernel accept queue (`listen.backlog = 4096`) already behaves as a crude queue today,
because nothing waits on a `runAsync` response. **It is not a substitute for this work**:
an FPM restart drops it silently, it has no fairness so one tenant starves another, it
loses requests silently past the backlog, and Apache abandons a proxied request after 310 s
so waiting is capped at `wait + execute < 310 s`. Those four gaps are the queue's reason to
exist.

## 3. Current execution path

```
Business rule
  → POST /codeRunner              CodeRunner::run()      app/Controllers/Tools/CodeRunner.php:191
      lexes + parses to validate the script
      inserts the codeRunner_executions row
      ignore_user_abort(true); echo 'OK'; closeConnectionEarly()   ← fastcgi_finish_request()
      Utils::executeAsyncPostCurl(base_url().'codeRunner/runAsync', [jobId, asyncNonce])  :388
  → POST /codeRunner/runAsync     CodeRunner::runAsync() :399
      verifies the single-use per-job nonce, then interprets the DSL
```

`executeAsyncPostCurl` uses `CURLOPT_TIMEOUT 2` and treats cURL errno 28 as success, so the
dispatch is genuinely fire-and-forget — **nobody is waiting for a `runAsync` response.**
That is what makes an out-of-band drain possible without changing the caller's contract.

Error-code family: `106` async dispatch failed, `112` nonce missing, `113` nonce invalid,
`201` lexer/parser, `301` caught exception, `302` fatal, `303` reaped.

### Landmine: the async nonce expires in 5 minutes

`CodeRunnerModel::consumeAsyncNonce()`
([CodeRunnerModel.php:193](../app/Models/Tools/CodeRunnerModel.php#L193)) defaults to
`$ttlSeconds = 300`, measured against `codeRunner_jobs.started_at`. Today that is generous,
because dispatch follows admission within milliseconds. **Under a queue, any execution that
waits longer than five minutes will be rejected as error 113.** A gated account (§8) can
easily wait longer than that.

Fix by minting the nonce at **dispatch** time rather than at admission: the drainer writes a
fresh nonce to the job row immediately before the dispatch cURL. That keeps the single-use
property intact, keeps the TTL meaningfully short, and removes the coupling between queue
latency and nonce validity. Do not simply raise the TTL — that widens the replay window for
every execution to cover the worst-case one.

### What actually bounds a runaway execution

Three mechanisms, and it is worth being precise about which covers what, because §9 depends
on it:

| bound | value | what it actually catches |
|---|---|---|
| `Interpreter::MAX_EXECUTION_SECONDS` | 180 s | Checked at the **top of each loop iteration**, in `for` and `foreach` — the only two loop constructs in the DSL, so every loop is guarded ([Interpreter.php:498](../app/Libraries/Businessmap/CodeRunner/src/Interpreter/Interpreter.php#L498), [:545](../app/Libraries/Businessmap/CodeRunner/src/Interpreter/Interpreter.php#L545)). It bounds how many further iterations start. It does **not** bound time spent inside one iteration, and it never runs at all in straight-line code. |
| `CURLOPT_TIMEOUT` | 180 s | A hung network call ([HTTPConnect.php:164](../app/Controllers/Components/HTTPConnect.php#L164)). |
| `request_terminate_timeout` | 300 s | Everything else — and in particular a retry ladder, which lives entirely inside one statement and therefore crosses no iteration check. |

**A 429 retry ladder is bounded only by the FPM wall.** That is the reason §9 keeps FPM as
the executor rather than interpreting the DSL inside the drainer.

### `codeRunner_executions`

Created by `2026-05-26-133832_CreateCodeRunnerExecutionsTable.php`; `source_code` added by
`2026-06-26-120000`, `idx_status_started_at` by `2026-06-18-100000`.

| column | type | note |
|---|---|---|
| `execution_id` | INT UNSIGNED PK AI | |
| `job_id` | INT UNSIGNED NULL | |
| `subdomain` / `subdomain_hash` | VARCHAR(500) / CHAR(64) | hash is the lookup key |
| `identifier` | VARCHAR(255) NULL | |
| `status` | VARCHAR(16) NOT NULL default `running` | |
| `error_code` / `error_message` | INT NULL / TEXT NULL | |
| `source_code` | MEDIUMTEXT NULL | |
| `debug` | TINYINT(1) default 0 | |
| `started_at` | **DATETIME(3) NOT NULL** | see §5 |
| `finished_at` | DATETIME(3) NULL | |

Indexed on `execution_id`, `subdomain_hash`, `started_at`, `job_id`, `(status, started_at)`.
This is already most of a job table, which is why §5 recommends extending it rather than
adding a second one.

### The reaper

`BaseModel::clearJobs()` ([BaseModel.php:208](../app/Models/Tools/BaseModel.php#L208)) marks
rows still at `status='running'` after **15 minutes** as `error_code 303`, and purges the
table on a 3-month window keyed on `started_at`.

## 4. Settled decisions

| Decision | Why |
|---|---|
| **Global concurrency target 4** | Pool ceiling of 8 sits above it. §6 shows 4 is 580× average demand |
| **Concurrency cap is per account (2 in flight); the rate-limit gate is per API key** | Different jobs — see §7 and §8. Capping per key alone would let a multi-key account take all four slots |
| **Per-account limits are mandatory, not optional** | See §6 — a global cap alone *creates* head-of-line blocking |
| **Drained by a long-running CLI daemon under systemd** (see §9) | No `request_terminate_timeout` on the drainer itself, no Apache or ELB coupling |
| **The drainer gates; FPM still executes** (architecture A, §9) | Keeps the 300 s wall, which per §3 is the only thing bounding a retry ladder, and leaves the execution path untouched |
| **1 Hz poll, ~500 ms median dispatch latency** | Accepted explicitly. §9 shows the cost is ~0.03 % of a core |
| **The minutely/hourly split in `send()` collapses into the gate** | §8. An hourly limit becomes "this key is gated for 30 minutes" instead of "the next N executions all fail" — strictly better for the customer, and only possible because the queue gives the work somewhere to wait |
| **Backpressure: reject at admission; per-account 500, global 2,000** | §11.1. Sized as a runaway stop, not as routine backpressure — normal depth is single digits |
| **Thin command + a floored `CodeRunnerQueueService`** | §12. Mirrors Scheduler/SchedulerService, but joins the coverage floors because it is the component that decides tenant occupancy of the box |
| **No retry on worker death** | DSL scripts commit side effects — cards created, comments posted. Re-running a half-finished script duplicates them. Today's `303` is terminal-without-retry; keep that. A queue makes retrying tempting; resist it |
| **The sync/debug path stays synchronous** | There is a human in a UI on `/codeRunner/test`; it must not queue |
| **"Queued" must be distinguishable from "failed"** in customer-visible state | Otherwise support reads a backlog as an outage |

## 5. Schema change

Three columns and one index, in one migration:

- **Add `queued_at DATETIME(3) NOT NULL`.**
- **Make `started_at` nullable.**
- **Add `apikey_hash CHAR(64) NULL`** — see §8; the gate is per key, and the drainer must be
  able to compute the gate key straight off the claim query. The apikey itself lives
  encrypted on the job row and is scrubbed to NULL after two hours by `clearJobs()`, so
  decrypting per candidate row on a query that runs every second is not an option.
- **Add an index on `(status, queued_at)`** for the claim query, and **`(status,
  subdomain_hash)`** for the per-account admission count (§11.1). The claim wants
  `WHERE status='queued' ORDER BY queued_at`; the count wants
  `WHERE status='queued' AND subdomain_hash=?`. Four indexes on a table taking ~400 inserts
  a day is not a write-cost concern.

Then the row's meaning is unambiguous at every stage, which it is not today:

| state | `queued_at` | `started_at` | `finished_at` |
|---|---|---|---|
| queued, not yet picked up | set | NULL | NULL |
| running | set | set | NULL |
| finished | set | set | set |
| **never dispatched** | set | **NULL** | set, with an error code |

`started_at IS NULL` is then the exact discriminator between "never dispatched" and "killed
mid-run" — the two states `303` conflates today.

Prefer this over a separate `codeRunner_queue` table: the row is already created in `run()`
at exactly the moment an execution is admitted, so a second table would need reconciling
against it for no gain.

### Nullable `started_at` breaks both halves of the reaper

Both must be fixed in the same commit as the migration
([BaseModel.php:250-290](../app/Models/Tools/BaseModel.php#L250-L290)):

- The **303 reap** filters `status='running' AND started_at < cutoff`. NULL-safe, so queued
  rows are correctly skipped — but that means nothing ever reaps a queued row if the drainer
  dies. Add a `status='queued' AND queued_at < cutoff` arm.
- The **3-month purge** keys on `started_at <`, so never-dispatched rows are immortal.
  Key it on `queued_at` instead, which is set in every state.

## 6. Measurements — use these, do not re-derive them

14 days to 2026-08-26, from `codeRunner_executions`.

**Healthy service time (successes only — this is the sizing input):**

| | |
|---|---|
| Successful executions | 5,200 |
| Mean duration | **1.6 s** |
| Under 5 s | **97.2 %** (5,054 of 5,200) |
| Over 240 s | 9, max 317 s |
| Total successful work | 8,320 execution-seconds of 1,209,600 s wall clock |
| **Average useful concurrency** | **0.0069 workers** |

**Volume and burstiness:**

| | |
|---|---|
| Executions, 14 d | 5,856 (~418/day, ~17/hour) |
| Peak hour | 414 — 2026-08-24 00:00 (next highest in 14 d: 164) |
| Peak minute | 137 — 2026-08-24 14:23 |
| Peak concurrent | 120 — 2026-08-24 12:12:59 |
| Recent week vs prior | `last_7d` = 4,284 of 5,856 → **2.7×**. Size for growth. |

At the 1.6 s healthy service time, 4 workers clear ~2.5/s and 8 clear ~5/s, so the worst
burst ever recorded drains in well under a minute. **4 is 580× average demand.** This is not
a throughput problem — it is burst absorption and drain latency.

Note that the peak hour on record sits inside the 00:00–01:00 UTC maintenance window, where
apt restarts services at 00:45. The drainer will be bounced at roughly the time load peaks —
which is why acceptance criterion 2 must be tested against a *kill*, not a clean stop.

**Concentration — the argument for the per-account cap:**

Top 3 accounts are **89 %** of volume, the top one alone 45 %. The shape matters more than
the share (these `avg_s` figures are contaminated per the trap below — directional only):

| account | n | avg_s |
|---|---|---|
| `c79be0…` | 2,630 | 41.2 |
| `8bbbed…` | 1,858 | **2.9** |
| `3ec13b…` | 720 | **223.5** |

`3ec13b` is ~57 % of all execution-seconds from 12 % of the executions — almost certainly
the rate-limit sleeper. Put these in one 4-wide FIFO and `8bbbed`'s 2.9-second scripts sit
behind `3ec13b`'s 223-second ones: the well-behaved tenant's automation is delayed by
minutes because another tenant is being throttled. **A global cap alone creates this.**

This grouping is by `subdomain_hash`, i.e. by tenant. That is the right unit for the
*fairness* argument above. It is **not** the unit of the API quota — see §8.

### Traps in this data, each of which gave a wrong answer first

- **Never size from observed peak concurrency (120).** That peak is executions sleeping on
  rate limits *caused by their own concurrency*. Sizing to it rebuilds the problem at a
  smaller scale. Size from `arrival_rate × healthy_service_time`.
- **Durations on reaped rows are fabricated.** The reaper sets `finished_at` to when the
  janitor noticed, not when the worker died — inflated by up to a full purge cycle, and
  `clearJobs()` itself can be slow. `max_s` read 2,865 s / 3,574 s on a pool whose wall is
  300 s. Any `AVG` over all rows blends real durations with janitor artifacts. **Always
  segregate by outcome** (`status='success'` / `error_code=303` / other).
- **The other outcome classes, for completeness:** `error_code 301`, n = 364, mean 111.3 s;
  reaped, n = 187. Those two are where the worker-seconds went.
- **`runAsync`'s 60-second cap does not exist.** [CodeRunner.php:402-403](../app/Controllers/Tools/CodeRunner.php#L402-L403)
  calls `set_time_limit(60)` and `ini_set('max_execution_time', '60')`. Both silently fail:
  the pool sets `php_admin_value[max_execution_time] = 300`, and `php_admin_value` cannot be
  overridden by `ini_set()` in *either* direction. It would not fire anyway — on Linux
  `max_execution_time` counts CPU only, so `sleep()` and network waits never accrue, which
  is why 364 executions averaging 111 s produced zero "Maximum execution time" fatals. **The
  only effective wall is `request_terminate_timeout = 300 s`.** Do not rely on those two
  lines.
- **One data-quality oddity:** 2 rows carry `subdomain_hash = e3b0c442…7852b855`, the
  SHA-256 of the empty string, meaning both the salt and the subdomain were empty. Implies
  `CredentialService::getCredentialValue('salt')` returned empty at that moment. Two rows
  only, but a missing salt would silently break subdomain lookups elsewhere.

## 7. Two limits, two keys

These are separate mechanisms protecting separate things. Conflating them is the easiest
mistake to make here.

| | keyed on | purpose | where it lives |
|---|---|---|---|
| **Concurrency cap** — max 2 in flight | account (`subdomain_hash`) | Fairness between tenants (§6) | In-memory in the drainer, plus the claim query |
| **Rate gate** — "not before T" | API key (`apikey_hash`) | Respecting the quota | Memcached, §8 |

The Businessmap rate limit is **per API key**, not per account. So the gate must be per key.
But if an account runs several keys, a per-key cap alone would let it occupy all four slots
while every individual key stays under its limit — so the *fairness* cap stays per account.
In the common one-key-per-account case the account cap bounds the key as a side effect.

## 8. The rate gate

Instead of sleeping inside a running execution, record the cooldown and refuse to *start*
work for that key until it passes.

This is not a new pattern here — ReportingProxy already ships it. On a 429 it stores
`time() + retry_after` in memcached under a per-key key with a TTL
([ReportingProxy.php:1885-1893](../app/Controllers/Tools/ReportingProxy.php#L1885-L1893)),
and `waitUntilReportingApiAvailable()`
([ReportingProxy.php:1808](../app/Controllers/Tools/ReportingProxy.php#L1808)) refuses to
start until the stored timestamp has passed, deleting the key when it has. Mirror it:

- Key: `cr_api_avail_<apikey_hash>`. Note ReportingProxy keys on
  `md5(subdomain|apikey)`; for CodeRunner the key alone is the quota unit.
- Value: `time() + retry_after + 5`. TTL `max(300, retry_after + 120)`.
- Written by the executing `runAsync` request when it sees a 429, at both the `< 100` and
  the `>= 100` branch.
- Read by the drainer before dispatch. A gated key's rows are skipped, not failed.

**The minutely/hourly split collapses.** Today `retry_after >= 100` terminates the execution
outright ([BusinessmapConnect.php:288-294](../app/Controllers/Components/BusinessmapConnect.php#L288-L294)),
and that is only because there was nowhere to park the work. With a queue there is. An
hourly limit becomes a 30-minute gate on that key, not N failed executions. This is the
single biggest customer-visible improvement in the design.

### What the gate cannot do

**It protects executions that have not started.** It cannot help the one that is mid-script
when the 429 arrives — that script has already created cards and posted comments. At that
instant there are only three options: wait, fail and lose the partial work, or requeue and
duplicate the side effects. The third is forbidden by §4. So the gate does not remove the
decision at the 429, it makes it rare — which is the actual win, because a gated execution
never reaches that fork at all.

It is also not authoritative. The key's quota is shared with the portal backend, the other
SA tools and the customer's own integrations, so the gate predicts remaining quota rather
than measuring it. **You will still see 429s you did not cause.** Gate proactively, and keep
the capped in-execution wait from §9 as the mid-script fallback.

### Sizing consequence

The gate is what makes queue depth grow, so the backpressure limit must be derived from the
longest gate — up to ~60 minutes of one key's arrivals. At the top account's average rate
that is ~8 rows; at the 2026-08-24 peak, ~200. Size the cap from that, not from a round
number.

## 9. The drainer

**Architecture: the drainer gates, FPM executes.** The drainer claims rows and POSTs
`/codeRunner/runAsync` exactly as `run()` does today, but blocking rather than
fire-and-forget, at most 4 in flight and at most 2 per account — `curl_multi` with 4 handles.

The alternative — forking children that interpret the DSL in-process — was rejected. Per §3
the 300 s FPM wall is the only thing bounding a retry ladder, and a drainer that executes
in-process throws it away. That is strictly worse than today: a blocked script currently
holds 1 of 40 `[www]` workers for 300 s; in-process it would hold 1 of 4 slots indefinitely,
and four of them would kill the queue permanently. `pcntl` and `posix` are both loaded on
this box so the wall could be rebuilt with `pcntl_alarm`, but that is new safety-critical
code replacing a mechanism already proven in production. Architecture A also makes
`pm.max_children = 8` structurally unreachable, which is acceptance criterion 3.

**Process shape: a systemd service, not a cron.** Follow
[deploy/monitoring/fpm-pool-metrics.service](../deploy/monitoring/fpm-pool-metrics.service),
whose header makes this same argument. `Type=simple`, `Restart=always`, `RestartSec=1` —
lower than the metrics sampler's 10, because a 10-second dispatch gap matters here in a way
it does not for monitoring. Keep a `flock(LOCK_EX|LOCK_NB)` guard as protection against a
manual double-start, following [Scheduler.php:19-30](../app/Commands/Scheduler.php#L19-L30),
but systemd — not the lock — is the supervisor.

A cron-per-minute shape was considered and rejected: its only advantage was a fresh DB
connection each minute, and at a 1 Hz poll the connection is never idle for more than a
second, so MySQL's `wait_timeout` (default 28800 s) never fires. It also costs a ~5 s
dispatch gap at each minute boundary for nothing.

You still need a reconnect arm around the poll — a MySQL restart or a server-side kill will
drop the handle — but that is one `catch`, not a lifecycle.

**Poll cost.** A 1 Hz `WHERE status='queued' AND …` against the `(status, queued_at)` index
is an index dive returning zero rows on an idle queue: ~0.1–0.3 ms including the local round
trip. 86,400 queries/day, ~30 seconds of DB time per day, ~0.03 % of one core. Against an
average arrival rate of ~17 executions/hour, polling is not the expensive part of anything.

**Dispatch latency** is therefore p50 ~500 ms, worst case ~1 s, against today's effectively
immediate fire-and-forget. Accepted. A FIFO in `writable/` that `run()` pokes would close
that gap, and with the 1 Hz poll retained as the durability backstop it would carry no
silent-failure mode — but it is extra moving parts for 500 ms. Do not build it in v1;
measure the real latency first.

**Memcached from CLI:** instantiate `MemcachedHandler` directly with the cache config, as
ReportingProxy does at [ReportingProxy.php:102-105](../app/Controllers/Tools/ReportingProxy.php#L102-L105).
The `cache()` service defaults to the `file` handler and is not what you want.

**Single drainer.** One process holding the in-flight counts makes the per-account cap
in-memory bookkeeping with no distributed coordination. If a second drainer is ever needed,
the claim query must become `SELECT … FOR UPDATE SKIP LOCKED` — check the MySQL version
first; it needs 8.0+.

## 10. Blocking dependency — cap the retry ladders first

**The queue does not work without this.** With failures still holding a slot for 111–300 s,
a 4-wide queue has 4 slots against executions that can each occupy one for five minutes, so
a single throttled key stalls the drain — just behind a queue instead of in the pool.

It is also worth shipping on its own, ahead of everything else: an attempt cap plus the
pool from `eecfbb1` is already a complete fix for the *box-offline* failure. 143 executions
dying at the 300 s wall become 143 dying at ~35 s, in a pool that cannot exceed 8 workers.
Everything after that is fairness, durability and visibility.

### The 429 ladder

The existing semantics are deliberate and **must be preserved until §8's gate replaces
them**:

- `retry_after < 100` → **wait and execute.** Getting the execution done matters more than
  the delay.
- `retry_after >= 100` (the hourly limit) → **terminate.**

So the change is an **attempt cap inside the `< 100` branch**, not a change to that split.
Note that a wait is at most 104 s and the wall is 300 s, so a budget check alone cannot trip
until roughly t≈208 — two full waits already spent. **At low `retry_after` it is the attempt
cap, not a budget check, that does the work**: with `retry_after: 10` you get ~19 attempts
across 300 s, and a cap of 3 fails at ~35 s. The actual `retry_after` values are recoverable
— `addToLogFile` records "Auto proceed after N seconds" on every wait, so `codeRunner_logs`
has them.

### The 401 ladder is worse, and the first draft missed it

[BusinessmapConnect.php:252-258](../app/Controllers/Components/BusinessmapConnect.php#L252-L258)
does an **unconditional** `sleep(60)` and recursive re-send on a 401 "API limit reached for
this method" — no `getRequestRetryEnabled()` guard, no `retry_after` split, no cap. Under
FPM that is ~5 attempts before the wall. Cap it in the same change.

### Implementation note

The retry is recursion — `return $this->send(...)`. Do **not** use an instance counter:
`BusinessmapConnect` is long-lived across a script's many API calls, so a counter would
never reset and the third 429 anywhere in a script would silently disable retries for the
rest of it. Add a defaulted attempt parameter to `send()` instead. It is self-resetting,
because each fresh logical call starts at 0.

`HTTPConnect` already carries the adjacent machinery — `requestRetryEnabled`,
`setRetryAfter()`, and a `$maxAttempts = $this->requestRetryEnabled ? 3 : 1` transport-level
retry at [HTTPConnect.php:473](../app/Controllers/Components/HTTPConnect.php#L473) — so
match its shape and its default of 3.

**The portal backend has its own copy of both ladders and needs the same fix separately.**

## 11. Resolved — the three questions the first draft left open

1. **Backpressure: reject at admission, never age out. Per-account 500, global 2,000.** A
   business rule fired because something actually happened in Businessmap, so a silently
   aged-out execution is a silent functional failure for the customer — the worst of the
   three outcomes even though it looks like the tidiest. Reject with a new error code so the
   row is visible in the execution log and the caller sees a failure.

   Both numbers are sized like `pm.max_children = 8`: far above normal operation, present as
   a runaway stop rather than as routine backpressure. Depth grows because of the gate, so
   the input is the longest gate — an hourly limit, ~60 minutes of one key's arrivals:

   | | rows accumulated in a full 60-minute gate |
   |---|---|
   | Top account at its average rate (2,630/14 d ≈ 7.8/hr) | ~8 |
   | Worst hour ever recorded (414, 2026-08-24 00:00) | ~414 |
   | All three concentrated accounts gated at once, at peak | ~1,200 |

   So 500 per account is the worst hour on record plus ~20 %, and 2,000 global sits above
   all three concentrated accounts being gated simultaneously at peak. **Normal depth is
   single digits.** If either cap ever fires, treat it as a defect signal, not as the system
   working as designed.

   Both counts run at admission, in `run()`, on the index added in §5. At the peak minute
   ever recorded that is 137 counts a minute against an index-only scan — not a hot-path
   concern.

   Note that the cap bounds depth, not staleness, and staleness is the thing that actually
   hurts: a business rule's action delivered 40 minutes late may be worse than useless. That
   is covered by making oldest-queued-age observable (§16), not by a lower cap — rejecting
   at admission and aging out after admission are different things, and the second stays
   forbidden.
2. **Ordering: FIFO plus the per-account cap.** Cheapest correct version, and free under the
   single-drainer shape in §9. Round-robin across accounts is a comparator change if drain
   latency proves unfair in practice; do not build it speculatively.
3. **Splitting `303`: infer from `started_at IS NULL` at read time.** Do not add a code. The
   distinction is derivable from §5's schema, and the error *message* can differ without the
   code differing — a new value in a customer-visible family is not free.

### Still open

Nothing blocking. Two things that can only be settled once code exists:

- The **floor number** for `CodeRunnerQueueService` (§12) — set it from the first coverage
  run, a little under what the tests actually achieve, per the convention in
  [tests/coverage-floors.json](../tests/coverage-floors.json).
- Whether drain latency proves unfair enough to want round-robin over FIFO (§11.2). Decide
  from observed oldest-queued-age per account, not speculatively.

## 12. Testing obligations

The first draft said nothing about this, and in this repo that is not optional — the commit
gate will block on it.

- `CodeRunner.php` (floor 99.0), `CodeRunnerModel.php` (98.0) and `BaseModel.php` (93.0) are
  all in the tested scope. The admission change, the nonce change, the claim/gate methods
  and both reaper arms land inside those floors and need tests in the same commit.
- `BusinessmapConnect.php` is **not** in the floors, so §10's attempt cap lands unmeasured.
  That is the status quo for that file; note it in the commit message rather than expanding
  the scope mid-change.
- **The drainer is a thin command plus a tested service, and the service joins the floors.**
  Mirror [Scheduler.php](../app/Commands/Scheduler.php) → [SchedulerService.php](../app/Libraries/SchedulerService.php):
  a ~40-line `app/Commands/CodeRunnerQueue.php` that does flock and delegates, and an
  `app/Libraries/CodeRunnerQueueService.php` holding the poll loop, the claim, the gate
  check, the per-account accounting and the dispatch — with injectable models and a
  `protected function dispatch()` seam, exactly as `SchedulerService` has.

  This goes one step further than the precedent: `SchedulerService` has
  [tests/unit/SchedulerServiceTest.php](../tests/unit/SchedulerServiceTest.php) but is *not*
  floored. The drainer is, because the stakes differ — if `SchedulerService` misfires you get
  a missed cron; if the drainer misfires you get 2026-08-24 again. It is the component that
  decides how much of the box one tenant can occupy, so it carries the standing obligation.

  **Structure it this way from the first line.** The loop needs a single-tick entry point
  (`runOnce()` or equivalent) that a test can drive without the `while (true)`, and the
  clock, the memcached handle and the models all need to be injectable. Retrofitting that
  onto a finished `curl_multi` poll loop is far more work than building it in.
- The suite must never reach the network. The drainer's dispatch is an HTTP call, so it goes
  through the `Transport` seam or the loopback echo server, never raw cURL — `curl_multi` in
  the drainer will need a seam of its own, and [scripts/seam-budget.php](../scripts/seam-budget.php)
  will fail the gate otherwise. Design that seam before writing the dispatch loop, not after.

## 13. Corrections to the first draft

- **§6 said `retry_after` is per-account. It is per API key.** The concentration analysis is
  still correct as a fairness argument, but it is not a quota argument. §7 splits the two.
- **§3's account of the execution wall was wrong.** `MAX_EXECUTION_SECONDS` is not a
  wall-clock cap on an execution; it is a per-iteration check that never runs in
  straight-line code and never interrupts a statement in progress. The table in §3 replaces
  it.
- **The 401 ladder was missing entirely.** See §10.

## 14. What not to do

- Do not raise `pm.max_children` to absorb bursts. Every worker you add is another slot a
  throttled key can sleep in.
- Do not retry a script whose worker died. Side effects are already committed.
- Do not route `/codeRunner/test` through the queue.
- Do not treat the accept queue as durability.
- Do not size anything from the contaminated `avg_s`/`max_s` columns.
- Do not raise the async nonce TTL to cover queue latency. Mint the nonce at dispatch (§3).
- Do not use an instance counter for the attempt cap (§10).

## 15. Implementation order

1. **Attempt caps on the 429 and 401 ladders** (§10). Self-contained, independently valuable,
   no schema. Ship and observe before anything else.
2. **Migration + reaper fixes** (§5). `queued_at`, nullable `started_at`, `apikey_hash`, the
   `(status, queued_at)` and `(status, subdomain_hash)` indexes, and both arms of
   `clearJobs()`. Still no behaviour change —
   `run()` keeps dispatching inline, `queued_at` is just written alongside `started_at`.
3. **The rate gate** (§8). Written by `runAsync`, read by nothing yet. Observable in
   memcached, zero risk.
4. **The drainer** (§9), with `run()` switched from inline dispatch to admission-only, the
   nonce minted at dispatch, and the gate consulted. This is the step that changes behaviour;
   everything before it is reversible on its own.
5. **Backpressure** (§11.1). The numbers are settled, so this is implementation rather than
   investigation — but it lands in `run()`, which step 4 already rewrites, so fold it into
   step 4 rather than touching admission twice.

## 16. Acceptance criteria

- Sustained submission of 500 executions from one account never puts more than 2 of them in
  flight, and never delays a second account's execution by more than one service time.
- An FPM restart or a drainer **kill** (`SIGKILL`, not a clean stop) mid-backlog loses
  **nothing** — the surviving rows are still `queued` and get picked up. Test this against
  the 00:45 service restart specifically, per §6.
- A key under an hourly limit gates rather than failing: its executions stay `queued` and
  drain after the cooldown, and no other key's executions are delayed by it.
- `pm.max_children = 8` is never reached; `max_children_reached` stays 0 on the pool (it is
  `ondemand`, so unlike `[long]` and `[reporting]` that counter is a real signal there).
- Neither depth cap fires under a replay of any burst in the 14-day window. If one does,
  that is a defect to diagnose, not a limit working as intended (§11.1).
- An execution rejected at admission for depth is visible in the execution log with its own
  error code, and the caller receives a failure rather than a silent drop.
- A queued execution is visibly "queued", not "failed", in the tool's own log view.
- Queue depth, oldest-queued-age and the set of currently gated keys are observable, so a
  stalled drainer is distinguishable from an idle one and from a gated one.
- Replaying the 2026-08-24 12:08 window's arrival pattern produces no `[www]` impact and no
  health-check degradation.

## 17. Reference

- `deploy/php-fpm/codeRunner.conf` — the pool, with the measurements that sized it
- `deploy/apache/businessmap-sa-coderunner.conf` — routing, and the duplicate-authority trap
  that cost three days on this host
- `deploy/php-fpm/www.conf` — what the shared pool carries and why it must be protected
- `deploy/FPM-LIMITS-AND-CONCURRENCY.md` — what queues vs what fails in the FPM limit chain
- `deploy/monitoring/PORTAL-UI-PLAN.md` §13 — how to read the pool dashboard while testing
- `deploy/monitoring/fpm-pool-metrics.service` — the systemd daemon precedent (§9)
- `app/Commands/Scheduler.php` — the CLI-command and flock precedent
- `app/Controllers/Tools/ReportingProxy.php` §`waitUntilReportingApiAvailable` — the working
  rate-gate implementation this design copies
