Skip to content

Added bounded Mailgun Logs fetching - #30690

Closed
jonatansberg wants to merge 5 commits into
jonatan-ber-3936-mailgun-logs-adapterfrom
jonatan-ber-3936-bounded-mailgun-fetching
Closed

jonatansberg wants to merge 5 commits into
jonatan-ber-3936-mailgun-logs-adapterfrom
jonatan-ber-3936-bounded-mailgun-fetching

Conversation

@jonatansberg

@jonatansberg jonatansberg commented Sep 10, 2026 •

Copy link
Copy Markdown
Member

ref https://linear.app/ghost/issue/BER-3936/raise-mailgun-event-fetch-throughput-concurrent-paging-logs-api

Serial polling waits for each Logs page and then its processing before requesting another. This adds opt-in, one-page prefetch so network latency can overlap processing, with bounded buffering, provider backoff and explicit shutdown draining.

Serial:    read 1 → process 1 → read 2 → process 2
Prefetch:  read 1 → process 1 ─────────→ process 2
                 └ read 2 (one ahead) ┘

Shutdown:  stop new cycles → abort reads/waits → finish processing + final flush
                                             → retain interrupted window for replay

emailAnalytics.fetchPrefetch defaults to false and applies only when fetchSource is logs. Callbacks and sending domains remain serial. A known event cap prevents another speculative request; processing failures abort and settle any pending read. The page token, filters and fixed time window stay the same on retries.

HTTP 429 reads get up to three retries with exponential backoff and jitter. We honor the later of Retry-After and Mailgun's millisecond reset timestamp, including exhausted-quota headers on successful responses. Each page has a 30-second total backoff budget; a longer reset fails the fetch without retrying early. Boot shares the cooldown across newsletter, automation and gift readers and subsequent polling cycles. Only numeric scheduling hints cross the sanitized transport boundary. Processor callbacks are outside the retry boundary.

Analytics now owns explicit pre-stop and cleanup hooks because its background worker returns after emitting an event on the main thread. Shutdown prevents new cycles and immediate restarts, aborts Logs reads and quota waits, and waits for active processing and final aggregation. An interrupted fetch retains its original cursor and scheduled backfill. The Events fallback lets its active SDK request settle before stopping further processing. Serverless boot remains supported, and a drained wrapper can initialize again without duplicate subscriptions.

Stacked on #30686. This fetch stack is independent of member preparation and counter rollout. Opens-first scheduling is retained. Domain concurrency remains deferred: one-page overlap preserves the current shared processor state and cursor ownership without concurrent database callbacks. Suppression cleanup is outside this performance scope: its events run in a separate fetch loop that is not performance sensitive. Durability and correctness will be revisited when the ongoing durable background jobs work is available.

Synthetic benchmark, 10 pages × 100 events, one warm-up plus three recorded samples per mode:

Injected read / processing delay per page Serial median Prefetch median
10 / 0 ms 163 ms 161 ms
40 / 40 ms 886 ms 497 ms
100 / 20 ms 1,279 ms 1,083 ms
20 / 100 ms 1,276 ms 1,048 ms

The harness uses the actual driver and transport with intercepted HTTP, no database and no live provider calls. All 32 runs checked exactly-once ordered callbacks, bounded reads and no pending requests. These measurements isolate overlap under synthetic delays; they do not establish production throughput or provider quota.

Validation: 287 unit tests across analytics and the legacy Mailgun client, seven integration tests through a real serverless Ghost boot, Core/test typechecks and focused lint pass. Coverage includes prefetch bounds, unchanged retries, quota waits, cancellation, final-flush drainage, immediate restarts, retained schedules/cursors, and serverless/repeated boot. Five independent review passes covered each slice and the accumulated change; the serverless boot finding was independently verified and fixed with a regression test.

pnpm check still stops on 452 existing unrelated formatting files. The default remains Events with prefetch off. Live Logs read permission, region/domain visibility and matching-window parity remain required before opt-in, as documented in Mailgun ingestion.

@coderabbitai

coderabbitai Bot commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Comment @coderabbitai help to get the list of available commands.

@nx-cloud

nx-cloud Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

🤖 Nx Cloud AI Fix

Ensure the fix-ci command is configured to always run in your CI pipeline to get automatic fixes in future runs. For more information, please see https://nx.dev/ci/features/self-healing-ci


View your CI Pipeline Execution ↗ for commit ca5208f

Command Status Duration Result
nx run ghost:test:e2e ✅ Succeeded 2m 16s View ↗
nx run ghost:test:legacy ✅ Succeeded 3m 29s View ↗
nx run ghost:test:ci:e2e ✅ Succeeded 4m 7s View ↗
nx run-many -t test:unit -p ghost ✅ Succeeded 36s View ↗
nx run @tryghost/e2e:test:fixtures ✅ Succeeded 1s View ↗
nx run @tryghost/admin:build ✅ Succeeded 4s View ↗
nx run-many --target=build --projects=tag:publi... ✅ Succeeded 1s View ↗
nx run ghost-monorepo:lint:boundaries ✅ Succeeded 20s View ↗
Additional runs (2) ✅ Succeeded ... View ↗

💡 Verify your cache is correct by running tasks in a sandbox. Read docs ↗


☁️ Nx Cloud last updated this comment at 2026-09-15 07:35:31 UTC

@codecov

codecov Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 38.56502% with 137 lines in your changes missing coverage. Please review.
✅ Project coverage is 67.64%. Comparing base (a87c3a4) to head (15a66bc).

Files with missing lines Patch % Lines
...ver/services/email-analytics/mailgun-rate-limit.ts 32.46% 45 Missing and 7 partials ⚠️
...ver/services/email-analytics/fetch-mailgun-logs.ts 18.96% 39 Missing and 8 partials ⚠️
...email-analytics/email-analytics-service-wrapper.ts 47.36% 24 Missing and 6 partials ⚠️
...ervices/email-analytics/email-analytics-service.ts 68.75% 3 Missing and 2 partials ⚠️
...core/core/server/services/email-analytics/index.ts 80.00% 1 Missing and 1 partial ⚠️
...er/services/email-analytics/mailgun-logs-client.ts 80.00% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@                            Coverage Diff                            @@
##           jonatan-ber-3936-mailgun-logs-adapter   #30690      +/-   ##
=========================================================================
- Coverage                                  67.76%   67.64%   -0.13%     
=========================================================================
  Files                                       1673     1674       +1     
  Lines                                      60495    60702     +207     
  Branches                                   10469    10518      +49     
=========================================================================
+ Hits                                       40996    41062      +66     
- Misses                                     17174    17291     +117     
- Partials                                    2325     2349      +24     
Flag Coverage Δ
e2e-tests 70.36% <38.56%> (-0.17%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jonatansberg
jonatansberg force-pushed the jonatan-ber-3936-mailgun-logs-adapter branch from 60abc15 to 75501f6 Compare September 11, 2026 09:32
@jonatansberg
jonatansberg force-pushed the jonatan-ber-3936-bounded-mailgun-fetching branch 4 times, most recently from 5fd980b to a63a3c7 Compare September 11, 2026 13:53
@jonatansberg
jonatansberg force-pushed the jonatan-ber-3936-mailgun-logs-adapter branch 2 times, most recently from f74cfcd to a87c3a4 Compare September 11, 2026 14:20
@jonatansberg
jonatansberg force-pushed the jonatan-ber-3936-bounded-mailgun-fetching branch from a63a3c7 to 15a66bc Compare September 11, 2026 14:21
ref https://linear.app/ghost/issue/BER-3936/raise-mailgun-event-fetch-throughput-concurrent-paging-logs-api

Allow one provider read to overlap current-page processing without overlapping
callbacks or domains. Keep timestamp ties and caps intact, handle pending read
errors immediately, and abort and settle speculative work on processing failure.

Keep prefetch disabled by default and preserve the serial Events fallback.
ref https://linear.app/ghost/issue/BER-3936/raise-mailgun-event-fetch-throughput-concurrent-paging-logs-api

Honor provider quota hints while retrying only page reads, with bounded waits and cancellation for prefetched requests.
ref https://linear.app/ghost/issue/BER-3936/raise-mailgun-event-fetch-throughput-concurrent-paging-logs-api

Keep provider cooldowns across polling cycles and drain interrupted ingestion safely before shutdown, retaining cursors and scheduled backfills for replay.
ref https://linear.app/ghost/issue/BER-3936/raise-mailgun-event-fetch-throughput-concurrent-paging-logs-api

A provider cooldown that outlasted the wait budget failed the whole run
and rewound its cursor to the window start, so pages already processed
were fetched again next cycle. The Logs adapter now stops such a run as
throttled: it keeps what each domain covered and leaves the whole window
for domains it never read. Cooldown hints are bounded to one hour so one
bad header cannot park every reader until restart, readers waking from a
shared cooldown add jitter so they do not fire together, and a reset sent
in seconds is read instead of ignored.

A planned shutdown is no longer reported as a job failure or a failed
provider request, the running flag is cleared when a stop interrupts a
run between lanes, reinitializing while interrupted fetches drain warns
instead of failing boot, lifecycle tasks register once per server, and
a non-boolean prefetch setting is rejected instead of silently ignored.
ref https://linear.app/ghost/issue/BER-3936/raise-mailgun-event-fetch-throughput-concurrent-paging-logs-api

A run that keeps waiting within its per-page budget has no whole-run
deadline; the doc now says so and names the lag warning as the signal, so
the absence of a deadline reads as a decision rather than an omission.
@jonatansberg
jonatansberg force-pushed the jonatan-ber-3936-bounded-mailgun-fetching branch from 15a66bc to ca5208f Compare September 15, 2026 07:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant