; ============================================================================ ; Long pool - requests that legitimately run for minutes to hours. ; ; Routed here by conf-available/businessmap-sa-long.conf, Included ONLY by the ; solutions vhost (000-default.conf). Members as of 2026-08-17: ; healthIndexes, internal/premiumHealthIndexes set_time_limit(10800) ; reportingApi/v1/{playgroundReportingProxy, ; premiumHealthIndexes, userAssignment, proxyDataWarehouse} ; userAssignmentReportBB, boardAssignmentReport, ; generateExecutiveReportRD, hubspotIntegration set_time_limit(10800) ; proxyDataWarehouse (NOT /dashboard/*) DW tool endpoints ; external/dw* DW ingest, 7200s ; dw/v1 Power BI parquet download ; boardAllocationReport set_time_limit(600) ; ; Removed 2026-08-14: subscriptions, moveCards, planner2Bmap, inviteUsers. ; Removed 2026-08-17: reportingProxy, which now has its own pool - see ; pool.d/reporting.conf for why (incompatible timeout requirements, plus isolation ; from the DW pipeline). playgroundReportingProxy stays here on purpose. ; The proxyDataWarehouse dashboard stays in the fast pool - it is UI, and ; dashboardRefresh() only fires an async cURL at /external/dwIngestRunner. ; ; Fire-and-forget callers (Utils::executeAsyncPostCurl, 2s curl timeout) abandon ; the request while ignore_user_abort(true) keeps the worker running - so ; request_terminate_timeout, not the client, governs how long a job may run. ; ============================================================================ [long] ; Must match the owner of writable/ - same as the www pool. ; Confirmed on this host: stat -c '%U:%G' /var/www/production/writable/logs ; -> www-data:www-data user = www-data group = www-data ; Separate socket from the fast pool. listen = /run/php/php8.3-fpm-long.sock listen.owner = www-data listen.group = www-data listen.mode = 0660 ; Static: all children are forked at startup and never reaped, so this pool holds ; max_children processes permanently (~14MB each when fresh). pm = static ; 20, down from 80, once reportingProxy left for its own pool. Sized from concurrency ; of the REMAINING routes over 2026-08-03..17: ; peak seen on any day, whole fortnight: 11 (2026-08-13) ; peak on the days after the pool split: 5-9 ; p95 concurrency, every single day: 2-4 ; So 20 is roughly 2x the highest value ever recorded here. 80 was never sized from ; anything; it was raised on 2026-08-14 while this pool still carried reportingProxy, ; which accounted for ~39k of the ~52k requests/14 days. ; ; Caveat: those peaks were reconstructed by pool-traffic-report.pl before the interval ; bug found on 2026-08-17 (it used [end - duration, end]; Apache's %t is the RECEIPT ; time, so [t, t + duration] is correct). The error INFLATED peaks, so the real figures ; for this pool are at most 11 and 20 keeps its ~2x margin either way. Recompute before ; using them for anything tighter. ; ; The intervals do include the fire-and-forget DW jobs: Apache logs them with the real ; duration even though the caller abandoned the connection (a dwIngestRunner run ; appears at 1315s), so they are not missing from the peak. ; ; Do not cut this further without re-measuring. Paired with the 11000s terminate ; below, ~20 stuck async jobs would hold every worker for three hours, and ; playgroundReportingProxy stays in this pool running the same code that produced ; the 120-way fan-out in [reporting]. ; Cut 20 -> 12 on 2026-08-18, as the counterweight to raising memory_limit to 512M below. ; What bounds this pool is the PRODUCT, not either number: 20 x 512M is a 10.2GB theoretical ; ceiling on a 7.7GB box with ~5GB MemAvailable, where 12 x 512M is 6.1GB. Neither is ; reachable in practice - DW ingests are serialised per (subdomain, api-key) by dw_ingest_locks ; and measured concurrency here is 1-5 with a peak of 9 - but the ceiling is what decides ; whether a bad day ends in a caught exception or in the kernel OOM killer choosing a victim, ; and it does not have to choose a PHP worker. MySQL and Apache are on this host too. ; ; The cost is real and is the wedging margin the paragraph above warns about: 12 is still above ; the highest concurrency ever recorded for these routes (11, on 2026-08-13), but the cushion ; against the 11000s request_terminate_timeout drops from ~2x to ~1.1x. Twelve stuck async jobs ; now hold the whole pool for three hours where it used to take twenty. That trade is deliberate: ; a wedged [long] stalls DW ingests and shows up in dw_runs.status, while host memory exhaustion ; is not contained and not attributable. ; ; If this pool ever needs to grow back, raise max_children and memory_limit together as one ; decision, and watch memav/swout in the pool metrics rather than either figure alone. pm.max_children = 12 ; Raised from 1 on 2026-08-14. At 1 every request paid a fork plus a full ; CodeIgniter bootstrap - correct when this pool served only rare multi-hour jobs, ; wasteful once real traffic routed here. Kept at 200 rather than lowered with the ; child count: this pool's jobs are long and few, so recycling pressure is low. ; Workers measured at 52.5MB PSS against ~14MB fresh, so still watch for growth. pm.max_requests = 200 ; Per-worker request URI is NOT usable here: ProxyPassMatch rewrites the URI to ; /index.php before FPM sees it, so ?full reports /index.php for everything. ; Pool-level counters (listen queue, max active) are still meaningful. pm.status_path = /fpm-status ; --- Timeouts ------------------------------------------------------------ ; Wall-clock hard kill and the real governor for the async jobs under FPM. Set ; just above the longest legitimate job (3 h) plus margin. ; NOTE: the ALB in front has idle_timeout=600s and a 4000s maximum, so no ; synchronous request can ever reach this ceiling - only fire-and-forget jobs, ; whose caller has already walked away, can use it. request_terminate_timeout = 11000s ; CPU-time ceiling. On Linux this counts CPU only - sleep() and network waits do ; not accrue. php_admin_value, so set_time_limit() in code CANNOT change it; 10800 ; is above every set_time_limit() value in the codebase, so nothing is truncated. php_admin_value[max_execution_time] = 10800 ; Per-request memory ceiling. php_admin_value, so the ini_set('memory_limit','256M') calls in ; DwIngestRunner/DwDrainQueued/DwIngest cannot change it in either direction - they are now dead ; code that reads as if it set the limit. Previously unset, which meant the global php.ini default ; of 128M, the ceiling behind 105 OOM fatals in ReportingProxy on 2026-08-13. ; ; Raised 256M -> 512M on 2026-08-18. The trigger was dados-okr-memo-dev (initiativeStatus, run ; 1382) failing with ReportPayloadTooLargeException: an 11.3MB payload needing ~68MB to decode ; against 66.2MB free. Note what that arithmetic says - ~190MB of the 256MB was ALREADY allocated ; before the page was fetched, so the decode was not the consumer; accumulated $seenKeys and prior ; windows' enrichment were. This pool is also the right place for the headroom now that ; reportingProxy has left: act is 1.0-1.4 with peaks of 1-5 post-split. ; ; This is NOT what fixed that failure, and should not be read as the remedy. IngestRunner now ; treats an undecodable payload the same way it treats BM's tree cap - halve the per-window budget ; and retry the same cursor - and the FIRST halving takes that window to ~5.6MB needing ~34MB, ; which already fitted in the 66.2MB free at 256M. The run self-heals at any limit. 512M exists ; because INIT_CURSOR_MIN_BUDGET floors how far a window can shrink, so headroom still has to come ; from somewhere for the reports that hit the floor. ; ; Paired with pm.max_children = 12 above. Do not raise one without re-deciding the other. php_admin_value[memory_limit] = 512M ; Must stay off: CI4 throws if it is enabled (app/Config/Events.php) and the DW ; gateway streams pre-gzipped bytes. php_admin_value[zlib.output_compression] = Off catch_workers_output = yes ; The www pool has had a slowlog since the cutover; this one did not. 300s so it ; catches genuinely stuck work rather than normal long jobs. This is the only ; instrument that can identify what the long pool is running, given the ; pm.status_path limitation noted above. request_slowlog_timeout = 300s slowlog = /var/log/php8.3-fpm/long-slow.log