# deploy/ — Apache + PHP-FPM configuration for solutions.businessmap.io

**Mirror of what is live on the host.** Nothing here is read by the application at runtime.
It is tracked so the web-tier config is reviewable, has history, and survives the box —
previously it existed only in `/etc` on one EC2 instance, which is how it silently drifted
from the templates that were supposed to describe it.

These files are copies, not the source of truth. **Run the drift check below before trusting
them.**

## Files and where they live

| Repo | Host |
|---|---|
| `php-fpm/www.conf` | `/etc/php/8.3/fpm/pool.d/www.conf` |
| `php-fpm/long.conf` | `/etc/php/8.3/fpm/pool.d/long.conf` |
| `php-fpm/reporting.conf` | `/etc/php/8.3/fpm/pool.d/reporting.conf` |
| `apache/php-fpm-handler.conf` | `/etc/apache2/conf-available/php-fpm-handler.conf` |
| `apache/businessmap-sa-long.conf` | `/etc/apache2/conf-available/businessmap-sa-long.conf` |
| `apache/businessmap-sa-reporting.conf` | `/etc/apache2/conf-available/businessmap-sa-reporting.conf` |
| `apache/businessmap-sa-coderunner.conf` | `/etc/apache2/conf-available/businessmap-sa-coderunner.conf` |
| `apache/portal-coderunner.conf` | `/etc/apache2/conf-available/portal-coderunner.conf` — the **portal** vhost's Code Runner paths onto the same `[codeRunner]` pool; see below |
| `apache/businessmap-sa-portal-redirects.conf` | `/etc/apache2/conf-available/businessmap-sa-portal-redirects.conf` |
| `apache/mpm_event.conf` | `/etc/apache2/mods-available/mpm_event.conf` — **not** conf-available |
| `monitoring/logrotate-php8.3-fpm-slowlog` | `/etc/logrotate.d/php8.3-fpm-slowlog` |
| `monitoring/logrotate-update-solutions` | `/etc/logrotate.d/update-solutions` |
| `maintenance/apt-daily.timer` | `/usr/lib/systemd/system/apt-daily.timer` — **vendor file, edited in place** |
| `maintenance/apt-daily-upgrade.timer` | `/usr/lib/systemd/system/apt-daily-upgrade.timer` — **vendor file, edited in place** |

Each file carries its own rationale in comments. This README covers only what spans them.

## Drift check

```bash
diff /etc/php/8.3/fpm/pool.d/www.conf                      deploy/php-fpm/www.conf
diff /etc/php/8.3/fpm/pool.d/long.conf                     deploy/php-fpm/long.conf
diff /etc/php/8.3/fpm/pool.d/reporting.conf                deploy/php-fpm/reporting.conf
diff /etc/apache2/conf-available/php-fpm-handler.conf      deploy/apache/php-fpm-handler.conf
diff /etc/apache2/conf-available/businessmap-sa-long.conf  deploy/apache/businessmap-sa-long.conf
diff /etc/apache2/conf-available/businessmap-sa-reporting.conf deploy/apache/businessmap-sa-reporting.conf
diff /etc/apache2/conf-available/businessmap-sa-coderunner.conf deploy/apache/businessmap-sa-coderunner.conf
diff /etc/apache2/conf-available/portal-coderunner.conf        deploy/apache/portal-coderunner.conf
diff /etc/apache2/conf-available/businessmap-sa-portal-redirects.conf deploy/apache/businessmap-sa-portal-redirects.conf
diff /etc/apache2/mods-available/mpm_event.conf            deploy/apache/mpm_event.conf
diff /etc/logrotate.d/php8.3-fpm-slowlog                   deploy/monitoring/logrotate-php8.3-fpm-slowlog
diff /etc/logrotate.d/update-solutions                     deploy/monitoring/logrotate-update-solutions
diff /usr/lib/systemd/system/apt-daily.timer               deploy/maintenance/apt-daily.timer
diff /usr/lib/systemd/system/apt-daily-upgrade.timer       deploy/maintenance/apt-daily-upgrade.timer
```

The last two matter more than the rest. Every other file here lives in `/etc` and survives
package upgrades; those two are **vendor files owned by dpkg**, and systemd units are not
conffiles — an upgrade of the `apt` package overwrites them silently, with no prompt and no
`.dpkg-dist`. When that happens the maintenance window reverts to a random minute between
06:00 and 07:00 and nothing announces it. Run the diff after any apt upgrade, and re-copy:

```bash
sudo cp deploy/maintenance/apt-daily.timer         /usr/lib/systemd/system/apt-daily.timer
sudo cp deploy/maintenance/apt-daily-upgrade.timer /usr/lib/systemd/system/apt-daily-upgrade.timer
sudo systemctl daemon-reload
sudo systemctl restart apt-daily.timer apt-daily-upgrade.timer
systemctl list-timers 'apt-daily*' --no-pager
```

## The 00:00–01:00 UTC maintenance window

Reserved. `apt-daily-upgrade` installs packages at 00:45 and needrestart then restarts every
service linked against them — apache2, php8.3-fpm, mysql, redis, memcached — killing anything
mid-request with no PHP shutdown handler. DW ingest runs 1403 and 1459 died exactly that way
when the window still floated across 06:00–07:00.

**No scheduled report may start between 00:00 and 01:00.** The dashboard hour pickers render
00:00 as a disabled "Server maintenance" option rather than omitting it, so the rule explains
itself where the choice is made instead of looking like a missing entry someone should restore.

Two gaps worth knowing:

- `every_6h` / `every_12h` derive later fires from their anchor, so an anchor of 06:00 still
  produces a 00:00 run. The pickers cannot express that and the server-side validators only
  check `0 <= hour <= 23`. No active report uses those presets today.
- Proxy-mode reports fire on client request at any hour, so one can still be running at 00:45.
  The gateway serves the existing artifact and refreshes in the background, so a client sees
  stale data rather than an error.

## Restarting PHP-FPM by hand without destroying in-flight work

`systemctl reload php8.3-fpm` keeps the master and lets workers finish their current request
(`Reloading in progress ...` in the FPM log, master pid unchanged). `systemctl restart` kills
them outright (`Terminating ... / exiting, bye-bye!`, new master pid) — that is how DW ingest
run 1334 died on 2026-08-13 at 14:04:20, two minutes into report 74's 20-minute run, during the
pool-split work on this very config. A reload is enough for pool and php.ini changes; only a
swapped shared library needs the restart. Before either, check that nothing long is running:

```bash
sudo find /var/www/production/writable/dw -name '*.tmp' -newermt '-3 minutes'   # empty = no writer
sudo env SCRIPT_NAME=/fpm-status SCRIPT_FILENAME=/fpm-status REQUEST_METHOD=GET QUERY_STRING=full \
  cgi-fcgi -bind -connect /run/php/php8.3-fpm-long.sock | grep -A1 'state:.*Running'
```

## One ProxyPassMatch trap that costs days

**Every `ProxyPassMatch` on this host must use a distinct authority in its `fcgi://` URL.**
Apache keys reverse-proxy workers on the URL after the `|` and stores the unix socket as an
attribute of that worker, so two directives sharing `fcgi://localhost/...` collapse into one
worker and the second silently serves from the first's socket.

That is how `reportingProxy` ran in `[long]` from 2026-08-14 to 2026-08-17 while a correct,
`configtest`-clean `<LocationMatch>` pointed it at `[reporting]`, and the reporting pool sat
forked and idle receiving only the metrics sampler's polls. The symptom is invisible in every
config file, because each file is individually right.

Current assignment — keep these unique when adding a pool:

| Pool | Authority |
|---|---|
| `[long]` | `fcgi://long-pool/...` |
| `[reporting]` | `fcgi://reporting-pool/...` |
| `[codeRunner]` | `fcgi://coderunner-pool/...` (solutions' front controller) |
| `[codeRunner]` | `fcgi://portal-coderunner-pool/...` (the portal's front controller — same socket, distinct worker on purpose) |
| `[www]` | `fcgi://localhost` (pathless, via `SetHandler`) |

Verify routing by counter, never by reading the config: sample `accepted conn` on **all three**
sockets around a burst of ~30 requests. A single request cannot distinguish a real hit from the
status query's own increment — that mistake cost most of 2026-08-17.

Silence means the repo still reflects production. Anything else means one side changed without
the other — fix it in the same sitting, or this folder reverts to being a confidently wrong
runbook. That is exactly what happened to its predecessor.

## The portal's Code Runner paths on the `[codeRunner]` pool

Code Runner is moving from solutions to the portal, where the executions are queued
(`CodeRunnerQueueService` in portal-backend) and dispatched at most four at a time. The queue is
the policy; `[codeRunner]` is the ceiling under it, and `apache/portal-coderunner.conf` is what
puts the portal's executions in that pool instead of `[www]`. It is a separate file because
`businessmap-sa-coderunner.conf` hardcodes solutions' `index.php` and must not be included from
`portal.conf`. This is **not** the solutions → portal proxy; that comes later and is not in this
folder yet.

```bash
sudo install -m 644 deploy/apache/portal-coderunner.conf /etc/apache2/conf-available/portal-coderunner.conf
sudo a2enconf portal-coderunner
sudo apachectl configtest && sudo systemctl reload apache2
```

`a2enconf` makes it global, like the other pool files here. The regex is anchored on
`codeRunner`, which no other vhost serves at those paths, so it captures nothing elsewhere; to
scope it anyway, `Include` it from inside `portal.conf`'s `<VirtualHost>` instead.

Prove which pool answered — the pool is `pm = ondemand`, so a child exists only once it has served:

```bash
curl -s -o /dev/null -w '%{http_code}\n' -X POST https://portal.businessmap.io/codeRunner \
     -H 'apikey: x' -H 'subdomain: x' -d 'source=print_value(1);&business_rule_id=probe'
ps -o pid,cmd -C php-fpm8.3 | grep 'pool codeRunner'
```

A refusal from the portal's `run()` is the expected body; the point is the child. If the body is
the SPA's HTML, the location walk saw a path the regex does not cover — add that exact path
rather than widening to a prefix, so `/api/v1/codeRunnerLogs` stays in `[www]`.

## Why two Apache files

`php-fpm-handler.conf` is **shared by three vhosts**:

| Vhost | DocumentRoot |
|---|---|
| `solutions.businessmap.io` (`000-default.conf`) | `/var/www/production/public` |
| `portal.businessmap.io` (`portal.conf`) | `/var/www/portal/backend/public` via `Alias /api/` etc. |
| `devportal.businessmap.io` (`devportal.conf`) | `/var/www/devportal/backend/public` |

It holds only the `SetHandler` that sends `.php` to the fast pool, plus `ProxyTimeout`. All
three applications therefore share the `[www]` pool's worker count, `memory_limit`,
`max_execution_time` and `request_terminate_timeout` — a spike or a wedged endpoint in one
degrades the others. `php8.4-fpm` does not exist on this host, so portal and devportal running
on 8.3 is deliberate.

`businessmap-sa-long.conf` is **Included only by `000-default.conf`**, because its
`ProxyPassMatch` target hardcodes `/var/www/production/public/index.php`. If it were Included
by the portal or devportal vhosts, any URL there matching its regex would silently execute
businessmap-sa's front controller — wrong app, wrong routes, wrong database. Note both of those
vhosts `Alias /internal/`, and the regex contains `internal/premiumHealthIndexes`.

`n8n-proxy.conf` includes neither; it is a pure reverse proxy with no PHP.

## Two traps this config has already fallen into

**Anchored alternatives in the long-pool regex.** The trailing `(/|$)` requires each
alternative to end on a full path segment. A bare `external/dw` therefore matched *nothing* —
the real routes are `/external/dwIngestRunner`, `dwDrainQueued` and `dwIdleCleanup`, so the
character after `dw` was never `/` or end-of-string, and every 7200-second DW ingest ran in the
fast pool under its 300-second kill. `external/dw\w*` covers those and any future sibling.

**Prefix, not controller, decides the pool.** A controller reachable under two prefixes must
list both. `internal/premiumHealthIndexes` and `reportingApi/v1/premiumHealthIndexes` are the
same controller; so are `proxyDataWarehouse` and `reportingApi/v1/proxyDataWarehouse`.

## Related

- [`coderunner-extension/`](coderunner-extension/) — a VS Code extension that gives
  `.coderunner` scripts JavaScript-style syntax highlighting, with the packaged `.vsix` and
  its install steps. Neither a mirror nor a deployed file: developer tooling, kept here so a
  checkout carries it. Read by no server.
- [`monitoring/`](monitoring/) — pool occupancy and memory sampler plus two analysis
  scripts, and the reasons the raw numbers mislead (summed RSS, counters that reset on
  reload, positional access-log parsing). Unlike the rest of this folder these are
  *deployed* files, not mirrors; install steps are in its README.
- [`FPM-LIMITS-AND-CONCURRENCY.md`](FPM-LIMITS-AND-CONCURRENCY.md) — how the limit chain
  behaves: what queues, what fails, and why `pm.max_children` is a concurrency limit rather
  than a connection cap. Written for the Power BI / Looker Studio traffic hitting the Data
  Warehouse gateway, which is where this tier's load actually comes from.
- The load balancer is an ALB with `idle_timeout = 600s` (AWS maximum is 4000s), so **no
  synchronous request can outlive 600 seconds** regardless of the FPM ceilings. Long jobs work
  only because they are dispatched fire-and-forget: the caller abandons after a 2-second cURL
  timeout while `ignore_user_abort(true)` keeps the worker going.
- The ALB health check targets `/healthcheck.php`, which runs through FPM. It previously
  targeted a static HTML file that Apache served without touching PHP — so a dead FPM left the
  instance in service while every real request failed.

The mod_php → PHP-FPM cutover runbook that used to live here has been removed; the migration is
complete and every item on its review checklist is resolved. See git history if you need it.
