# PHP-FPM limits, queueing & concurrency — go-live reference

How requests are limited and queued under Apache (mod_proxy_fcgi) + PHP-FPM, why a request
*fails* vs. just *waits*, and how to tune the chain so bursty external connectors
(**Google Looker Studio**, **Power BI**) hitting the Data Warehouse gateway don't drop reports.

Read this alongside `README.md` (the cutover runbook) and `php-fpm/*.conf`.

---

## 1. The core fact: workers queue, they don't reject (until a limit is hit)

`pm.max_children` is a **concurrency** limit — how many requests run *at the same time* — not a
"connection cap" that refuses clients.

Example: 40 requests arrive, pool has 20 workers.
- **20 run immediately.**
- **The other 20 wait in the socket listen backlog** (`listen.backlog`) and are served FIFO as
  workers free up. They are **not** refused; the client does not have to retry.
- FPM logs: `server reached pm.max_children setting (20), consider raising it` — the signal the
  pool is undersized for the load.

A request only **fails** when one of the limits in section 2 is exceeded.

---

## 2. The full limit chain (a failure at ANY layer drops the request)

A request flows: **client → Apache (MPM) → FPM socket (backlog) → FPM worker**. Each hop has a
ceiling:

| # | Limit | Where set | Default | What happens when exceeded | Client sees |
|---|-------|-----------|---------|----------------------------|-------------|
| 1 | `MaxRequestWorkers` (MPM) | Apache mpm config | event: 400 | Apache can't accept the connection to proxy it | `503` |
| 2 | `pm.max_children` | FPM pool | — (we set it) | request waits in the backlog (no failure) | none (slower) |
| 3 | `listen.backlog` | FPM pool | 511 | kernel refuses new connections to the socket | `502` / conn refused |
| 4 | `net.core.somaxconn` | kernel sysctl | 4096 (modern) / 128 (old) | **caps `listen.backlog`** regardless of FPM value | `502` / conn refused |
| 5 | `ProxyTimeout` / `timeout=` | Apache proxy | core `Timeout` (300s) | request waited (queue + processing) too long | `504` |
| 6 | `request_terminate_timeout` | FPM pool | 0 (off) | worker hard-killed mid-request | `502` + truncated |

Key implications:
- Raising `pm.max_children` alone is not enough — if `listen.backlog` (3) or `somaxconn` (4) is
  small, bursts past the worker count are **refused**, not queued.
- `somaxconn` silently truncates `listen.backlog`. On older kernels (128) a `listen.backlog=1024`
  is effectively 128. Always raise the sysctl too.
- A request can sit in the backlog AND then process; both count against `ProxyTimeout` (5).

---

## 3. The Data Warehouse gateway gotcha (Looker / Power BI)

External BI connectors fetch from the **gateway** route `/dw/v1/...`
(`Routes.php` → `External\DwGateway::serve`). Two things make this the sensitive path:

1. **It streams the whole artifact via `readfile()`** — a worker is held for the **entire transfer
   duration** (seconds for a multi-GB parquet), not milliseconds. Concurrent downloads each pin a
   worker for as long as the transfer takes.
2. **Connectors fire many requests near-simultaneously.** If one of them gets a `502`/`503`/`504`,
   Looker/Power BI typically **fails the whole report** rather than retrying that one request.

⚠️ **Routing note:** in `apache/businessmap-sa.conf` the long-pool `LocationMatch` matches
`external/dw` (the *ingest* runner at `/external/dwIngestRunner`), **NOT** `/dw/v1`. As written,
Looker traffic lands on the **fast `www` pool** and competes with all UI/tool traffic for its
workers. For burst resilience, give `/dw/v1` its **own pool** (section 4).

---

## 4. Recommended setup to minimize connector failures

The goal for Looker/Power BI is **"serve everything, just slower"** (deep queue, generous timeout)
rather than **"fail fast"** — because one dropped request kills the whole report.

### 4a. Dedicated `dw` pool for the gateway

`php-fpm/dw.conf`:
```ini
[dw]
user = nfsnobody
group = nfsnobody
listen = /run/php/php8.3-fpm-dw.sock
listen.owner = www-data
listen.group = www-data
listen.mode = 0660

pm = static                  ; instant availability — no spawn lag when a burst lands
pm.max_children = 30         ; >= connector peak parallel requests; bounded by RAM (4d)
listen.backlog = 1024        ; deep queue: bursts past max_children WAIT instead of 502

request_terminate_timeout = 600s
php_admin_value[max_execution_time] = 600
php_admin_value[zlib.output_compression] = Off
catch_workers_output = yes
```

Route `/dw/v1` to it (in `apache/businessmap-sa.conf`):
```apache
<LocationMatch "^/dw/v1/">
    ProxyPassMatch "unix:/run/php/php8.3-fpm-dw.sock|fcgi://localhost/var/www/html/php8.3/production/businessmap-sa/public/index.php" timeout=600
</LocationMatch>
```

### 4b. Raise the kernel backlog cap (otherwise `listen.backlog` is truncated)
```bash
# /etc/sysctl.d/99-fpm.conf
net.core.somaxconn = 1024
# apply: sudo sysctl --system
```

### 4c. Use mpm_event with enough MaxRequestWorkers
So Apache can *hold* all the concurrent connections while they queue/stream. `prefork` (forced by
mod_php) needs one heavy process per held connection — a strong reason to switch MPM during the
cutover (see README runbook step 6). Ensure `MaxRequestWorkers` >= expected peak concurrent
connections across all pools.

### 4d. Sizing & the trade-off
```
pm.max_children  =  min( connector peak parallel requests ,  RAM_for_PHP / peak_process_size )
```
- Too low → bursts queue (slow) or, past the backlog, get refused.
- Too high → risk of OOM (every static worker reserves RAM even when idle).
- `pm = static` trades idle RAM for zero spawn latency — right for predictable burst traffic.

---

## 5. Monitoring signals (watch these at go-live)

- **FPM error log**: `server reached pm.max_children setting (N), consider raising it` → bump
  `pm.max_children` for that pool.
- **FPM status page** (enable `pm.status_path = /fpm-status` on a protected location): watch
  `listen queue`, `max listen queue`, `active processes`. A non-zero **`max listen queue`** means
  requests are queueing in the backlog — fine until it approaches `listen.backlog`, then you're
  near refusals.
- **Apache logs**: `502`/`503`/`504` on `/dw/v1` correlate to limits 3/1/5 respectively.
- **Kernel**: `ss -ltn` shows `Recv-Q`/`Send-Q` on the FPM socket; `nstat -az TcpExtListenOverflows`
  rising = backlog overflowing (raise `listen.backlog` + `somaxconn`).

---

## 6. To measure before final sizing

These two numbers drive `pm.max_children` and the timeouts — capture them from real traffic:

1. **Connector parallelism** — how many requests Looker Studio / Power BI fire per report refresh,
   near-simultaneously. (Set `pm.max_children` >= this for the `dw` pool.)
2. **Artifact size / transfer time** on `/dw/v1` — how long a worker is held per download. (Drives
   `timeout=` / `request_terminate_timeout`, and how quickly the queue drains.)

Until measured, the values in section 4 are reasonable starting points; revisit after observing the
FPM status `max listen queue` under real connector load.
