# Plan: FPM pool monitoring UI in the portal admin section

Handoff spec for whoever builds this in `portal.businessmap.io`. The data already
exists and is being written.

**Revised 2026-08-19.** The 2026-08-17 version is stale in ways that matter, because a
week of production analysis changed both the data and the conclusions. If you built
against it, these are the deltas:

- **The queue chart it specifies plots a field that is always zero.** `listen_q` and
  `max_listen_q` read 0 on UNIX-socket pools no matter how deep the queue is. Real depth
  comes from `sock_recvq`, which the sampler now emits. See §5.
- **There is now a permanent hourly archive** (`rollup-YYYY-MM-DD.json`) that outlives the
  45-day prune, plus three fields that exist only there: `real_reqs`, `pss_suspect`,
  `cfg_changed`. Wide ranges should read it instead of raw ticks. See §3, §4, §6.
- **An errors chart was missing.** §4 of the old version said to surface `err_oom`
  prominently and then listed five charts, none of which showed it.
- **Pool capacities changed again**: `[long]` 80 → 12 workers with a 512 MB limit,
  `[reporting]` → 40. Any example in the old doc citing 80 or 15 is wrong. This is the
  third change in six days, which is the whole argument for §5's "read capacity from the
  data".
- **New section §13, "Reading the dashboard."** Added 2026-08-27, after the first build was
  read as showing an idle server during a two-hour `[www]` stall. Operator-facing: what each
  tooltip term means, the triage order, and — §13.4 — the workloads that are *supposed* to
  hold a worker for minutes, which the §6 ladder misclassifies as stalls.
- **New section §6, "Reading server load correctly."** The most important addition. On
  2026-08-18 `[reporting]` sat on its 40-worker ceiling for six hours with average
  occupancy of 2–3 workers, and an occupancy chart alone would have called that a
  capacity problem. It was not. 39 of those 40 workers were asleep.

**Read §5 first.** The numbers in these files are easy to render *correctly* and easy to
render *misleadingly*, and several of the traps below cost real hours during the incidents
that produced this data. "Semantics you must encode" is not advisory — a dashboard that
ignores it will confidently show a healthy pool during a saturation.

---

## 1. Goal

An admin-only page that answers, without SSH:

1. Is any of the three FPM pools saturated right now, and was it at time T?
2. **When a pool is full, is it full of work or full of waiting?** These need opposite
   responses — more workers, or fix the stall — and they look identical on an occupancy
   chart. See §6.
3. Is worker memory growing, and how close is it to the ceiling?
4. Is a pool receiving the traffic it is supposed to receive?
5. Which routes and which clients consume the pools?
6. When did FPM last reload, die, or have its limits changed?

## 2. Where the data is

Same host as the portal (`10.0.2.176`). The portal backend runs as `www-data`; the files
are `root:www-data`, mode `640`, in a `2750` directory — so they are readable directly.
**No network hop, no agent, no database.**

| Path | Content | Written by |
|---|---|---|
| `.../metrics/pool-YYYY-MM-DD.jsonl` | one JSON object per pool per 15s | `fpm-sample.sh` (systemd `fpm-pool-metrics.service`) |
| `.../metrics/pool-YYYY-MM-DD.jsonl.gz` | same, days 3–45 | `fpm-metrics-prune` (cron.daily) |
| `.../metrics/rollup-YYYY-MM-DD.json` | **hourly aggregate of a finished day, kept forever** | `fpm-daily-rollup` (cron.daily) |
| `.../metrics/traffic-YYYY-MM-DD.json` | **per-route durations, statuses and per-pool concurrency for a past day, kept forever** | `fpm-traffic-snapshot` (cron.daily) |
| `.../metrics/traffic-YYYY-MM-DD.log` | the same day as a human-readable report, pruned at 45 days | `fpm-traffic-snapshot` (cron.daily) |
| `.../metrics/www-saturation-YYYY-MM-DD.log` | `?full` worker tables, only when the fast pool is ≥50% busy | `fpm-sample.sh` |

Base path is `/var/log/php8.3-fpm/metrics/`.

Raw ticks older than 45 days are **deleted**; from day 3 they are gzipped, so the backend
must handle both (`gzopen()` reads plain files too — use it unconditionally). Rollups are
never pruned, and are deliberately named `.json` not `.jsonl` so the prune glob cannot
match them.

Volume: ~590 bytes/line, 3 pools × 5,760 ticks = 17,280 lines/day, **~10.2 MB/day raw**,
~30 MB steady state at 45-day retention. Rollups are 550 B per pool-hour → 39 KB/day →
**~14 MB/year**. **Aggregate server-side; never send raw lines to the browser.**

## 3. Two tiers, and which to read

This is the main architectural change since the last revision.

| Range | Source | Why |
|---|---|---|
| today, or last few days | `pool-*.jsonl(.gz)` | 15-second resolution; needed to see a spike shorter than a bucket |
| anything older, or wider than ~a week | `rollup-*.json` | already hourly; a day of raw ticks is 17,280 lines, the same day as a rollup is 72 rows |
| **beyond 45 days** | `rollup-*.json` **only** | the raw ticks no longer exist at any price |

Measured on this data shape, PHP aggregating raw ticks streams at ~3.2 µs/line with flat
2 MB memory: 54 ms for a day, 392 ms for a week, 3.16 s for the full 45 days. So raw is
fine for recent ranges, and the rollup is what makes a multi-month range instant — and
possible at all.

A rollup row is **not** interchangeable with a bucket you computed yourself: it carries
three fields raw ticks cannot give you (§4), and it applies reload/counter-reset semantics
that are easy to get wrong. Prefer it whenever the requested bucket is ≥1h.

## 4. Field reference

### Raw tick — one line per pool per 15s

```json
{"ts":"2026-08-18T12:45:23Z","pool":"reporting","up":true,"accepted":8231,
 "listen_q":0,"max_listen_q":0,"sock_recvq":79,"sock_backlog":511,
 "active":40,"idle":0,"total":40,"max_active":40,
 "max_children_reached":0,"slow_requests":61,"start_since":9214,"workers":40,
 "rss_peak_mb":126.4,"rss_avg_mb":58.2,"pss_total_mb":1846,"oldest_s":9214,"oldest_req_s":63.4,
 "err_oom":0,"err_timeout":0,"warn_busy":0,"child_signal":0,
 "max_children":40,"pm":"static","memory_limit_mb":256,"terminate_timeout_s":300,
 "load1":1.9,"mem_avail_mb":2942,"swap_used_mb":0,"pswpin":0,"pswpout":0}
```

An unreachable pool writes `{"ts":…,"pool":…,"up":false}` and nothing else — render those
as an explicit outage band, never as a gap or a zero.

### Rollup row — one per pool per hour, in `hours[]`

Same information plus these, which you cannot reconstruct from a single tick:

| Field | Meaning |
|---|---|
| `samples` | ticks that contributed to the hour — the denominator for everything |
| `reqs` | accepted delta over the hour, reset-safe |
| **`real_reqs`** | `reqs` minus `samples`. The sampler's own status polls are *inside* `accepted`; subtracting them leaves real traffic. §5 |
| `active_mean` / `active_max` | mean and **peak of concurrent `active`** over the hour |
| `sockq_max` / `backlog` | peak pending connections, and the configured backlog |
| `pss_total_mb` / **`pss_suspect`** | pool memory, and whether a reload inflated it. §5 |
| **`oldest_req_s`** | longest-running request **in flight**, peak over the hour. This is the stall signal — see §6 |
| `oldest_s` | age of the oldest worker *process*. **Not a request metric.** Kept only because older rollups carry it; do not build anything on it — see the warning below |
| `slow`, `oom`, `timeout`, `busy`, `signal` | reset-safe deltas |
| `reloads`, `down` | reload count and unreachable ticks |
| `cfg_*` + **`cfg_changed`** | the ceilings each measurement was taken against. §5 |

## 5. Semantics you must encode

**`listen_q` is always 0 here — use `sock_recvq`.** FPM only reports a listen queue for
TCP listeners; all three pools use UNIX sockets, where it reads 0 regardless of actual
depth. Real depth comes from `ss -lx` Recv-Q, which the sampler emits as `sock_recvq`
(pending) and `sock_backlog` (the ceiling). On 2026-08-18 at 12:00, `sock_recvq` reached
**79** while `listen_q` sat at 0 — a queue chart built on `listen_q` would have been a
flat, empty, reassuring line during the worst queueing of the day.

**Saturation on a static pool is `active == max_children`, not `max_children_reached`.**
FPM increments that counter only when it must *spawn* a child and cannot, which never
happens under `pm = static` — every worker exists from startup. `[reporting]` read
`max_children_reached = 0` through all six hours it spent pinned at its 40-worker ceiling
on 2026-08-18. The counter stays useful for `www`, which is `dynamic`.

**Use `max`, not `mean`, for `active` and `sock_recvq`.** A pool saturated for four
minutes in an hour averages ~7% busy. `[reporting]` on 2026-08-18 hit 40 of 40 in six
separate hours with hourly means of 1.8–3.5. Show mean as a secondary series if useful,
never as the primary.

**`real_reqs` near zero means the pool is not receiving traffic.** A pool that is up,
idle, and *unreachable* looks exactly like a pool that is up, idle, and quiet — every
occupancy and memory field is identical. `real_reqs` separates them, because the only
requests a misrouted pool still serves are the sampler's own status polls. This is not
hypothetical: `[reporting]` existed and was healthy for hours before anything was routed
to it, and on 2026-08-17 the hourly `real_reqs` sat at 0–1 until 14:00 and then jumped to
319. Flag any pool with routes whose `real_reqs` is ~0 while `reqs` ≈ `samples`.

**Never multiply `rss_peak_mb` or `rss_avg_mb` by worker count.** Forked workers share
most pages copy-on-write; summed RSS massively overstates usage — on this box 36 workers
were added and `used` memory went *down* 257 MB. `pss_total_mb` divides shared pages by
the number of processes mapping them and **is** a valid total. It is `null` on most ticks
(sampled every 20th — it costs a page-table walk per worker), so carry the last known
value forward or plot sparse points, but never treat `null` as zero.

**`pss_suspect: true` means that hour's PSS is inflated.** During a graceful reload both
worker generations are alive at once, so the pool total roughly doubles for no real
reason. Mark those points rather than letting them read as a memory spike.

**Counters reset when FPM reloads.** `accepted`, `slow_requests`, `max_children_reached`,
`max_active` are cumulative since the pool last started. Detect a reload by
**`start_since` decreasing between consecutive samples**, and at that boundary: do not
compute a delta across it; draw a vertical marker on every chart; break the line series
rather than joining across it. A decrease means the source reset, so the new value *is*
the count since the reset.

**The four error counters reset on a different event.** `err_oom`, `err_timeout`,
`warn_busy`, `child_signal` come from FPM's master log, so they reset when *that log
rotates*, which has nothing to do with a pool reload. Treat any decrease as a reset and
use the new value as the delta. `dlt()` in `pool-metrics-summary.pl` is the reference
implementation.

**One request-terminate kill writes TWO log lines, so `err_timeout` and `child_signal`
double-count it.** FPM logs the timeout and then logs the child's death separately:

```
[26-Aug-2026 16:16:48] WARNING: [pool www] child 158659, script '…/index.php'
    (request: "POST /internal/index.php") execution timed out (310.954587 sec), terminating
[26-Aug-2026 16:16:48] WARNING: [pool www] child 158659 exited on signal 15 (SIGTERM)
    after 440.018615 seconds from start
```

Same event, same second, same pid. A stacked errors chart summing `timeout` and `signal`
will therefore show **two bars per kill**. Either stack only one of them, or pair them on pid
and count once. Note also that the SIGTERM line's duration is the **child process** lifetime
(440 s here), not the request duration (310.954 s) — reading it as a request time is wrong
and makes the kill look far worse than it is. Only the `execution timed out` line carries the
request duration. Conversely a SIGTERM with no preceding timeout line is a reload or a
`pm.max_requests` recycle, not a kill.

**Surface `err_oom` prominently, and give it its own chart.** 105 of these were the root
cause of the 2026-08-13 incident and 108 landed on `[long]` on 2026-08-14, yet they are
invisible in every other field: a request killed by `memory_limit` looks identical to a
fast successful one in the occupancy and memory charts, because the worker frees its
memory and returns to idle immediately.

**`cfg_*` is intent on disk, not what the master is enforcing.** An edit without a reload
shows the new number while FPM still runs the old one. Two cross-checks catch it, but only
one works everywhere: peak `active` above `cfg_max_children` means the file is behind the
master, on any pool; `max_children_reached` climbing while peak `active` sits below
`cfg_max_children` means the master's cap is lower than the file, and that one is
**dynamic-pools-only** for the reason above. Treat either as "needs a reload", not as a
measurement. Also: `cfg_*` values recorded **before 2026-08-19 are unreliable** — the
sampler read `pool.d` once at startup, so they froze at whatever was on disk when the
service last restarted. Do not draw a capacity line from `cfg_*` for 2026-08-14…08-18.

**`cfg_changed: true` invalidates comparison across that hour.** Occupancy either side of
a `max_children` change is not the same measurement.

**`swap_used_mb` is not swapping.** It sat at ~1.4 GB while the box was 95% idle: pages
parked during an old peak and never faulted back. Plot the `pswpout` delta if you want a
swap indicator. Do not colour `swap_used_mb` red.

**A dead sampler must look different from a healthy pool.** If the newest sample is older
than ~3× the interval (45s), show a stale-data banner. Silence otherwise renders as a
flat, calm chart — the most dangerous possible failure mode for this page.

**`active` can miss a spike shorter than 15s; `max_active` cannot.** When a bucket's
`max_active` exceeds the largest `active` you sampled in it, the true peak was higher than
plotted — surface that, e.g. as a faint band above the line. Note the naming trap: in the
rollup, `active_max` is the honest per-hour peak of concurrent `active`, while FPM's own
`max_active` is a high-water mark since the master started and is monotonic within a
master's life. Do not plot the latter as an hourly value.

**The traffic JSON carries no client IPs, by design.** Not truncated, not hashed —
absent, in both address families. Attributing load to a *customer* is done from the tool
logs' `Subdomain` column, which is how one account was identified as 24,765 of 24,920
worker-seconds on 2026-08-17; the access log could not have answered that anyway, since
every request arrives from the ALB. What it does carry is cost **by user agent**, which is
what distinguishes Looker Studio from Power BI from a script. The privacy section below
therefore has nothing to redact in Phase 2 — do not add it back.

**A pool figure in the traffic JSON counts only requests that reach PHP.** Static assets
are split into a separate `static` object and excluded from every pool bucket and from
`combined_peak`. This is not tidiness: measured on 2026-08-19, the non-pool bucket peaked
at 51 simultaneous requests, of which 50 were 1,521 static hits holding **2.9 seconds
between them** — one browser opening a page and firing fifty asset requests in the same
instant. Excluding them left a peak of 7, exactly what FPM's own sampled `active_max`
reported for `[www]` that day. So plot `static.peak` if you like, but never next to a
capacity line, and never as `[www]` occupancy: a high peak beside near-zero seconds is a
page load, not load.

**`--until` filters on request START, so the data runs past the requested end.** A request
arriving at 23:59:58 and running 200s is counted in full. The document reports `requested`
and `window` separately for this reason; every percentage in it divides by `window`, and a
consumer that divides by the requested range instead will be wrong by however far the tail
overhangs.

## 6. Reading server load correctly

This section exists because occupancy is not load, and on this box the difference is the
whole story.

`[reporting]` on 2026-08-18: pinned at 40 of 40 workers for six hours, hourly mean
occupancy 1.8–3.5, `sock_recvq` up to 79, **zero** OOMs, timeouts and kills. An occupancy
chart says "add workers". That would have been wrong. 39 of the 40 workers were asleep
inside the application, waiting on another worker to finish extracting the same report:
26,581 worker-seconds of sleeping against 3,364 seconds of actual work, a ratio of 7.9:1.
Adding workers buys more sleepers.

**A full pool has two very different causes, and they need opposite responses.**

| | full of work | full of waiting |
|---|---|---|
| `active` at cap | yes | yes |
| `oldest_req_s` | low — requests complete and recycle | **high** — the same requests sit there |
| `slow` delta | proportional to traffic | high relative to `real_reqs` |
| `reqs` per busy worker-second | high | **near zero** |
| right response | more workers, or faster code | fix the stall; more workers makes it worse |

So: **plot `oldest_req_s` next to occupancy, and never show occupancy alone.** It is the
single field that discriminates. A pool at its ceiling with `oldest_req_s` under a few
seconds is doing its job. A pool at its ceiling with `oldest_req_s` in the tens of seconds
is stalled on something, and the ceiling is a symptom.

> **Do not use `oldest_s` for this.** An earlier version of this document defined
> `oldest_s` as "longest-running request seen in the hour". That was wrong. It is computed
> from `ps` `etimes` — the age of the oldest worker *process* — and because these workers
> are never recycled it collapses into a duplicate of `start_since`. Measured 2026-08-25:
> all three pools reported `oldest_s == start_since == 14896` to the second, on an idle
> box, while `[www]` actually had a request 238 s into its life (79% of its 300 s
> terminate timeout). A ladder built on `oldest_s` would have scored rule 4 as STALLED
> during every genuine saturation event and never once reached rule 5. `oldest_req_s`,
> added to the sampler on 2026-08-25, is the correct field; rollups from before that date
> do not have it, so render it as absent rather than zero.

Suggested per-pool verdict for the header strip, in priority order — first match wins:

1. `up:false` samples in range → **DOWN**
2. newest sample older than 45s → **NO DATA** (never "healthy")
3. `oom` or `signal` delta > 0 → **KILLING REQUESTS** (this outranks saturation; it is
   silent everywhere else)
4. `active_max == cfg_max_children` **and** `oldest_req_s` high → **STALLED**
5. `active_max == cfg_max_children` **and** `oldest_req_s` low → **SATURATED**
6. `sockq_max > 0` → **QUEUEING**
7. pool has routes and `real_reqs` ≈ 0 → **RECEIVING NOTHING**
8. otherwise → **OK**

Pick the `oldest_req_s` threshold from the pool's own `terminate_timeout_s` rather than a
constant: these pools range from 300s to 11,000s, so a fixed number is meaningless across
them. Something like 5% of the terminate timeout is a reasonable starting point, tuned
against known-good hours. Note that a request approaching `terminate_timeout_s` is about
to be SIGTERMed, so the same field also drives a "request about to be killed" warning —
`[www]` at 238 s of 300 s was 62 seconds from exactly that.

**A percentage of `terminate_timeout_s` is not sufficient on its own, and rule 4 will
false-positive without §13.4.** Some workloads on this host are *designed* to hold a worker
for most of its pool's wall clock: the portal's `jira2bmap-worker` runs a 240 s budgeted
loop in `[www]` and refires itself back-to-back for the length of a migration. Measured
completions are 241–310 s, a median 259 s — **87% of that pool's 300 s timeout, for hours,
while nothing is wrong.** Read §13.4 before tuning
this threshold or the page will report STALLED through every Jira migration.

**What this page cannot tell you.** Whether a stalled worker is waiting on a lock, a slow
upstream API, or a database is not in these files — it is in the application's own logs.
Do not try to infer the cause here; the honest output is "stalled, and here is when",
with a link to the slowlog path for that hour so the backtrace is one step away.

## 7. Architecture

Read → aggregate → cache, all in the portal backend. No new database table: the files are
already the time-series store, with retention already enforced.

```
rollup-*.json ──┐
                ├→ selector (§3) → bucket aggregator → cache → JSON API → charts
pool-*.jsonl ───┘
```

**Stream, don't slurp.** Use `gzopen()`/`gzgets()` in a loop (works on plain files too). A
45-day range is ~460 MB raw; `file_get_contents` will exhaust `memory_limit`.

**Cache aggressively, keyed on file mtime.** Any `pool-<date>.jsonl` for a past day never
changes again — its aggregate can be cached indefinitely. Rollups likewise. Only today's
file is live; give it a 60s TTL. Key on `basename + mtime + bucket size` so a stale entry
is impossible.

**Never write to the metrics directory.** Read-only, always. Retention belongs to the cron
jobs; a UI that deletes files will fight them.

### Endpoints

```
GET /api/admin/fpm/now
    → live snapshot: last sample per pool, staleness, and the §6 verdict

GET /api/admin/fpm/series?from=<iso>&to=<iso>&bucket=5m|15m|1h|1d&pool=<name>|all
    → aggregated buckets per pool, reload markers, and which tier served it

GET /api/admin/fpm/traffic?date=YYYY-MM-DD          (Phase 2)
    → per-route durations, statuses, concurrency, top consumers

GET /api/admin/fpm/saturation?date=YYYY-MM-DD       (Phase 3)
    → parsed www ?full dumps
```

Clamp `to - from` and reject `bucket=5m` on ranges over ~3 days — 5-minute buckets over 45
days is 12,960 points per pool. Return `"source": "ticks" | "rollup"` per range so the UI
can say "hourly resolution only" instead of silently drawing a smoother line.

### `/series` response shape

```json
{
  "from": "2026-08-18T00:00:00Z",
  "to":   "2026-08-19T00:00:00Z",
  "bucket_seconds": 3600,
  "source": "rollup",
  "pools": {
    "reporting": {
      "config": { "max_children": 40, "pm": "static", "memory_limit_mb": 256,
                  "terminate_timeout_s": 300, "stale_before": "2026-08-19" },
      "buckets": [
        { "ts": "2026-08-18T12:00:00Z", "samples": 240,
          "active_max": 40, "active_mean": 2.71, "fpm_max_active": 40,
          "sockq_max": 79, "sock_backlog": 511,
          "reqs": 866, "real_reqs": 635,
          "max_children_reached": 0,
          "rss_peak_mb": 126.4, "rss_avg_mb": 58.2,
          "pss_total_mb": 1846, "pss_suspect": false,
          "oldest_req_s": 63.4,
          "slow": 61, "oom": 0, "timeout": 0, "busy": 0, "signal": 0,
          "mem_avail_min_mb": 2942, "swap_out_pages": 0, "load1_max": 1.9,
          "down": 0, "reload": false, "cfg_changed": false,
          "verdict": "STALLED" }
      ]
    },
    "long": { "config": { "max_children": 12, "pm": "static", "memory_limit_mb": 512, "terminate_timeout_s": 11000 }, "buckets": [] },
    "www":  { "config": { "max_children": 40, "pm": "dynamic", "memory_limit_mb": 256, "terminate_timeout_s": 300 }, "buckets": [] }
  },
  "reloads": [],
  "stale": false,
  "newest_sample": "2026-08-19T09:31:12Z"
}
```

**Read `config` from the data — never hardcode it.** Hardcoding was tried and failed
inside two days: `[long]` went 40 → 80 workers on 2026-08-16 and the dashboard kept
drawing its capacity rule at 40, so a pool that had reached its ceiling looked like one
running at 200% of a limit that no longer existed. It has since gone to 12. `[reporting]`
went 32 → 40. Assume it will change again.

**Capacity is `max_children`, not `total`.** For a static pool the two coincide, but `www`
is `dynamic`: `total` is the workers currently *alive* — often 8 — while the ceiling is 40.
Drawing `total` as capacity would understate it fivefold and make an idle pool look
permanently saturated.

## 8. UI

### Header strip — "now"

Per pool: the §6 verdict as the primary element, then `active/max_children` as a gauge,
`sock_recvq`, `oldest_req_s`, `pss_total_mb`. Host-wide: `mem_avail_mb`, load. Plus a
prominent stale/outage banner.

Three traps this strip has already fallen into, all live as of 2026-08-25.

**"Oldest worker" was reading `oldest_s`**, so it showed 4.0h on three idle pools while
`[www]` had a request 238 s into a 300 s terminate timeout. Use `oldest_req_s` (§6).

**"Pool memory (PSS)" showed `—`** on all three pools, because `pss_total_mb` is computed
only every `PSS_EVERY` ticks (5 minutes) and is `null` in between. Carry the most recent
non-null value forward and label its age; never blank.

**A "Listen queue" tile must read `sock_recvq`, never `listen_q`.** FPM can only measure
accept-queue depth for TCP listeners, so on this host's UNIX sockets `listen_q` and
`max_listen_q` are hard-wired to 0 forever. Proven from the archive on 2026-08-24, all on
`[www]` at `active:40/40`:

| ts (UTC) | `listen_q` | `sock_recvq` |
|---|---|---|
| 12:08:39 | 0 | 119 |
| 12:13:08 | 0 | 96 |
| 14:24:40 | 0 | 54 |
| 14:50:34 | 0 | 38 |

A tile fed from `listen_q` displayed a healthy 0 while 119 connections waited unaccepted.

Also show load against `nproc`: this host has 2 cores, so 0.26 is 13% of capacity, and the
bare number reads very differently on 2 cores than on 8. The same archive has `load1:
17.19` at 12:16:56 — 860% oversubscribed — which no unqualified load figure conveys.
`max_children_reached` deserves a place too: it is a real signal on `[www]` (dynamic), hit
4 on 2026-08-24, and resets to 0 on reload, so a current 0 does not mean it never happened.

The verdict is the point of the strip. A number someone has to interpret is worse than a
word plus the number that justifies it.

### Charts (shared x-axis, shared zoom, reload markers on all)

1. **Occupancy + oldest request** — `active_max` line, capacity line at
   `cfg_max_children`, `active_mean` faint, and **`oldest_req_s` on a second axis**. Per §6
   these two together are the only honest read of load. Faint band up to `fpm_max_active`
   where it exceeds `active_max`. Y axis fixed 0..`cfg_max_children` — autoscaling makes 9
   busy workers look like a crisis.
2. **Queue** — `sockq_max` bars, **not** `listen_q`. Non-zero means requests waited for a
   worker; the ALB gives up at 600s.
3. **Memory** — `pss_total_mb` area with `pss_suspect` points marked, `rss_peak_mb` line
   (largest single worker) with a threshold at `cfg_memory_limit_mb`, `mem_avail_mb` on a
   second axis. This is the chart that answers "are workers leaking".
4. **Errors and kills** — stacked `oom`, `timeout`, `signal`, `busy` deltas. Zero most of
   the time and that is fine; it is the chart you need on the one day it is not. Do not
   fold these into another chart, and do not hide it when empty.
5. **Traffic** — `real_reqs` per bucket, with total `reqs` faint behind it so a
   receiving-nothing pool is visibly different from an idle one.
6. **Slow requests** — `slow` per bucket. Link each point to the slowlog path for that
   hour so the backtrace is one step away.

Presets: last hour / 6h / 24h / 7d / 45d / 6 months, with `bucket` and tier chosen
automatically. Mark the 45-day boundary on any range that crosses it — beyond it the data
is hourly forever, and a user should know why the line got coarser.

### Phase 2 — traffic table

Sortable: route, count, p50/p95/p99/max duration, peak concurrency, status breakdown. Plus
concurrency ("≥N for X seconds") and top consumers by **worker-seconds held**, not request
count — one 300s request costs the pool as much as 300 one-second requests, and only the
former can exhaust it.

## 9. Auth and privacy

- **Admin-only.** This exposes internal route names, client IPs and capacity.
- **Client IPs are personal data.** Truncate the final octet by default and put full values
  behind an explicit reveal, or drop them entirely if the route/UA breakdown is enough. Do
  not cache them beyond the 45-day file retention.
- **Do not render raw request paths.** `reportingProxy/advancedSearch/<token>` path
  segments are encrypted credential payloads. The Phase 2 JSON will arrive pre-redacted to
  the collapsed route name; do not undo that by re-parsing the access log yourself.

## 10. Constraints and non-goals

- **Do not parse the Apache access log on request.** Phase 2 consumes a nightly pre-built
  JSON snapshot only.
- **Do not parse that log positionally if you ever touch it.** The custom `LogFormat` puts
  duration (`%{ms}T`) at field 9 and status at field 11, so `awk '$9==502'` selects
  requests that took 502 *milliseconds*. Doing exactly that produced a completely
  fabricated error profile during the investigation. `X-Forwarded-For` also arrives as
  `ip1, ip2`, shifting every later column. The timestamp is the request *start*, and the
  interval is `[t, t+duration]`.
- **A pool absent from the sampler's `POOLS` list leaves no rows at all** — not even
  `up:false`. So a gap is missing coverage, not an idle pool, and the UI must not fill it
  with zeros. `[reporting]` has no rows before 2026-08-17T10 for exactly this reason.
- **Single instance assumed.** These files are local to `10.0.2.176`. If the ALB ever
  fronts more than one target, every chart silently becomes "one arbitrary instance". The
  fix is an `instance_id` dimension plus shipping files to S3 — design the JSON so adding
  that key later is not a breaking change.
- **No alerting here.** The ALB already emits `HTTPCode_ELB_5XX_Count`,
  `UnHealthyHostCount` and `TargetConnectionErrorCount` to CloudWatch for free; alarms
  belong there, not in a page someone has to be looking at.

## 11. Phases

| Phase | Scope | Blocked on |
|---|---|---|
| **1** | `/now`, `/series`, header strip with §6 verdicts, charts 1–6, both tiers per §3 | nothing — data is being written now, and 2026-08-14 onward is already archived |
| **2** | `/traffic` + route table | nothing, as of 2026-08-20. `pool-traffic-report.pl --json` exists and `fpm-traffic-snapshot` writes `traffic-YYYY-MM-DD.json` nightly, backfilled over the 14 days logrotate keeps. Read the file; do not shell out to the script |
| **3** | `/saturation` viewer for `www-saturation-*.log` | nothing, but low value until the fast pool actually saturates (peak so far: 10 of 40) |

Phase 1 is the whole point. Do not block it on Phase 2.

## 12. Acceptance checks

- [ ] A range spanning a reload shows a marker, a broken series, and **no negative**
      `reqs` bucket.
- [ ] Stopping `fpm-pool-metrics.service` makes the page show stale data within a minute,
      not a calm flat line.
- [ ] Stopping `php8.3-fpm` produces a visible outage band from `up:false`.
- [ ] A 6-month range responds from rollups without touching a `.jsonl` file.
- [ ] A range crossing the 45-day boundary is visibly marked as changing resolution.
- [ ] A past-day range served twice hits cache the second time.
- [ ] `pss_total_mb` renders as sparse points or held values, never as zeros, and
      `pss_suspect` hours are marked.
- [ ] The occupancy chart's y-axis is driven by `cfg_max_children` and stays fixed when the
      pool is quiet.
- [ ] **2026-08-18 12:00 for `[reporting]` renders as STALLED, not SATURATED**, and the
      queue chart shows 79 — this is the regression test for §5 and §6 together.
- [ ] **2026-08-17 before 14:00 for `[reporting]` renders as RECEIVING NOTHING**, not OK.
- [ ] No capacity line is drawn from `cfg_*` for 2026-08-14…08-18.
- [ ] Nothing in the UI displays a full `advancedSearch` path or an untruncated client IP
      by default.

## 13. Reading the dashboard — the operator's guide

Everything above is for whoever *builds* the page. This section is for whoever *looks at*
it, and it exists because the first build was read as "it always seems like there is no load
on the server" on a day when `[www]` had a request sitting at 250–300 s for two hours.

### 13.1 Why it reads as idle

Two reasons. The first is a measurement artefact you must know about; the second is a design
mistake in the page.

**The sampler counts itself.** FPM's `active processes` includes the worker serving the
status request that measured it, so a completely idle pool reports `active: 1`, never 0.
Three pools reporting exactly `1` in the same tick is the signature. Read `1/40` as "idle"
and `2/40` as "one real request in flight". Corroborate with `oldest_req_s`: `<1 s` is the
poll itself. **The builder should subtract the poll, or label the tile** — a floor of 1 that
nobody explains is why the strip reads as permanently-slightly-busy and simultaneously as
never-loaded.

**Occupancy is the wrong headline for this host.** It is the largest, first element on the
page, and it is the least informative one here. This box's failure mode is not throughput
saturation — it is workers *held* while consuming no CPU (§6). A pool can sit at 5%
occupancy and be in real trouble. Observed 2026-08-27: `[www]` reported 2 of 40 busy — 5%,
apparently nothing — while `oldest_req_s` had been at 250–300 s since 09:10, slow requests
were logged every bucket, and one request had been killed by signal. Every number was
correct and the layout buried all three that mattered.

### 13.2 The six questions of §1, and where each is answered

| Question | Field | Where | Caveat |
|---|---|---|---|
| How many workers are busy now? | `active` **− 1** | header gauge | the status poll is inside `active` (13.1) |
| What was the peak, and when? | `active_max` per bucket | Occupancy line | rarely the interesting peak here — see 13.4 |
| How much memory? | `pss_total_mb`, and `rss_peak_mb` vs `cfg_memory_limit_mb` | Memory chart | never sum RSS; `pss` is `null` on 19 of 20 ticks |
| How many requests waited or were delayed? | `sockq_max` | Queue chart | `listen_q` is hard-wired to 0 on UNIX sockets (§5) |
| How many were served? | `real_reqs` | Traffic chart | `reqs` includes the sampler's polls |
| Was anything killed? | `oom` / `signal` / `timeout` | Errors chart | invisible in every other chart (§5) |

The memory question has two different answers and they are not interchangeable.
`pss_total_mb` is what the pool costs the host. `rss_peak_mb` is the largest single worker,
and it is the only one the `memory_limit` line applies to — a pool total of 766 MB against a
256 MB limit is not a breach, it is 40 workers of ~19 MB each.

### 13.3 Triage order — the four questions that actually matter

Ask them in this order. On a normal day the first three are all zero and you can stop.

1. **Is anything stuck?** `oldest_req_s` ÷ that pool's `terminate_timeout_s`. Read 13.4
   before acting on it.
2. **Is anything queueing?** `sockq_max > 0` means requests waited for a worker.
3. **Is anything dying?** any of `oom` / `signal` / `timeout` > 0. This outranks everything
   else and is silent in every other chart.
4. **Only then, is it busy?** `active_max` against `cfg_max_children`.

### 13.4 Long requests that are supposed to be long

`oldest_req_s` discriminates work from waiting (§6). It does **not** discriminate a stall
from a workload that is legitimately long. Four of those run on this host, and a reader who
does not know them will call a healthy migration a stall.

| Workload | Pool | Normal in-flight duration | Why |
|---|---|---|---|
| `jira2bmap-worker` (portal) | `[www]` | **~260 s, back-to-back, for hours** | budgeted worker loop, self-refiring |
| DW ingest | `[long]` | minutes to hours | designed for it; 11,000 s terminate timeout |
| reportingProxy extraction | `[reporting]` | tens of seconds | followers wait on the leader (§6) |
| CodeRunner | `[codeRunner]` | 1.6 s median, 97 % under 5 s | past 60 s is anomalous by definition |

**`jira2bmap-worker` is the one that will fool you.** `php spark jira2bmap:dispatch` runs
from cron every minute, but that cron is only the fallback: a worker whose **240 s** budget
expires with work still to do **immediately POSTs its own next tick**, so an active
migration is an unbroken chain of ~260-second requests in `[www]` for as long as the
migration lasts — hours. One worker per concurrent migration; an account runs one migration
per Businessmap API key, so 1–3 per account. **This is the designed behaviour, not a leak.**

Measured 2026-08-27 from the access log, 38 consecutive ticks of two concurrent migrations:
**241.6 s min, ~259 s median, 310.1 s max.** The budget is 240 s but it is only checked at
*step* boundaries, so every tick overshoots by however long the step that crossed the line
takes — 1.6 s to 70.1 s here. Do not expect a flat 240 s; expect 240–285 s with a tail past
300 s. Two of the 38 (5.3 %) exceeded `[www]`'s 300 s `terminate_timeout_s`, and one reached
Apache's 310 s `ProxyTimeout` and was returned to the caller as **504**. The 504 itself is
cosmetic — `Utils::executeAsyncPostCurl` fires the worker with `CURLOPT_TIMEOUT 2` and
nobody is reading the response — but it is the visible edge of a margin that is too thin.

Three consequences for this page:

- **§6 rule 4 false-positives on it.** 240 s against `[www]`'s 300 s `terminate_timeout_s`
  is 80 %, far past §6's suggested 5 % threshold, so a healthy migration scores **STALLED**
  and trips the "about to be killed" warning on every tick for hours. Uncorrected, the page
  cries wolf and the one signal that discriminates stops being read. Either exempt the
  worker route, or drive the threshold from a per-pool expected-duration baseline rather
  than from `terminate_timeout_s` alone.
- **The `[www]` slow-request count is not a signal.** `request_slowlog_timeout` is 30 s
  there and every tick runs 240 s, so every tick is logged. Each entry costs a `ptrace` stop
  of the worker while its stack is dumped — the same amplification that left 31 workers in
  `T` state on 2026-08-24. During a migration the floor is one slow request per tick per
  migration. Do not present a non-zero `[www]` slow count as evidence of a problem.
- **The headroom is 60 seconds, the median tick already spends 19 of it, and one call can
  spend the rest.** The budget is checked at *step* boundaries, while the portal's
  `BusinessmapConnect` answers a per-minute 429 with
  `sleep(retry_after + 5)` **inside** a single call — up to ~104 s. A step entered at 239 s
  elapsed that then sleeps is SIGTERMed by FPM at 300 s. Measured 2026-08-27: 4 of 49 ticks
  (8.2 %) crossed 300 s and three were killed that day. Each kill redoes that tick's
  un-checkpointed step and holds a `[www]` worker for the full ~305 s. It does not appear to
  stop the migration — migration 97 absorbed all three of those kills and kept running — so
  treat the kills as waste, not as an outage. This is what `signal` on the Errors chart is
  catching, and the reason that chart must not be hidden when it looks empty.

**Two liveness traps, both of which look like a dead migration and are not.** Recorded
because both caught this document's author on 2026-08-27.

- **`heartbeat_at` age is not liveness.** It is stamped at *step* boundaries, and a step is
  not bounded by the tick budget — an attachment batch is 10 cards of download, validate,
  zip and upload, and can legitimately run for many minutes. A heartbeat older than a
  260 s tick is therefore normal mid-step, not evidence of a crash: 411 s was observed while
  the migration was demonstrably uploading files. Only a heartbeat age far beyond the
  slowest plausible step means anything.
- **`jira2bmap:dispatch` logging "0 candidate run(s)" every minute is the healthy steady
  state**, not a failure. A `running` row with a fresh-enough heartbeat is deliberately
  skipped (§6.3): the worker chain refires itself and owns the work, and the cron dispatcher
  is only the fallback for when that hand-off fails. Expect zero candidates for the entire
  duration of a healthy migration.

### 13.5 The Occupancy tooltip, term by term

These four are **Occupancy** series. They are easy to mistake for memory series because the
charts sit next to each other; the Memory chart's series are `pss_total_mb` (pool total),
`rss_peak_mb` (largest single worker, the one the limit applies to) and `mem_avail_mb`.

| Label | Field | Meaning |
|---|---|---|
| Peak busy | `active_max` | highest simultaneously-busy workers **in that bucket**. The honest occupancy figure. |
| FPM peak (high-water) | `max_active` | FPM's own high-water mark **since the last reload** — not per-bucket, monotonic within a master's life. Rendered as a band; where it sits above the line, the true peak was higher than any 15 s sample caught. |
| Mean busy | `active_mean` | average over the bucket. Never read alone: a pool saturated for four minutes of an hour averages ~7 % (§5). |
| Oldest request (s) | `oldest_req_s` | age of the longest-running request **in flight**, on the second axis. The stall signal (§6). |

Worked example — `[long]`, 24 h view, 2026-08-26 17:00: *Peak busy 5, FPM peak 10, Mean busy
2.7, Oldest request 891.1*. Five of twelve workers is unremarkable and is not the content of
that tooltip; a request running just under fifteen minutes is. Against `[long]`'s 11,000 s
terminate timeout that is 8 %, and entirely normal for a DW ingest. The identical 891 s on
`[www]` would mean a request that had already been SIGTERMed twice over.

### 13.6 What the page still cannot answer

- **"How many requests were delayed" only has a per-tick answer.** `sockq_max` is the peak of
  a 15-second sample; a queue that formed and drained between two ticks leaves no trace, and
  there is no cumulative "requests that waited" counter anywhere in FPM. Treat a non-zero
  value as proof queueing happened, never a zero as proof it did not.
- **Why a worker is stuck is not in these files** (§6) — lock, upstream API, or database is
  in the application's own logs. The honest output here is "stalled, and here is when", with
  the slowlog path for that hour.
- **Which customer caused it is not here either.** The traffic JSON carries no client IPs by
  design (§5); attribution is done from the tool logs' `Subdomain` column.

## 14. Reference

- [`README.md`](README.md) — the same traps from the operator's side, plus install steps.
- [`pool-metrics-summary.pl`](pool-metrics-summary.pl) — a working reference implementation
  of the aggregation, including reload detection and the counter-reset rules. Read it
  before writing the aggregator; `--json` emits exactly the rollup shape.
- [`../php-fpm/reporting.conf`](../php-fpm/reporting.conf) — why that pool is sized the way
  it is, with the measurements.
- `REPORTING_PROXY_429_RETRY.md` at the repo root — why `[reporting]` was full of sleeping
  workers, and the change that stops it. Useful context for §6: after it ships, that pool's
  occupancy should fall while its request count rises. **Lands with the
  `feature/reporting-proxy-429-reply-to-looker` branch**, so it is not on `dev` yet; read
  it there if this link does not resolve.
