Skip to content

Notifications and webhooks

Documentation home

Synapse can route tenant events to signed HTTP webhooks, Slack incoming webhooks, and email recipients. Delivery runs in synapse-worker; API requests and scans do not wait for a remote service.

Personal email contacts

When notifications and SMTP are enabled, every authenticated human user can manage their own email destinations at My profile (/profile). GET and POST /api/v1/me/contacts list/add contacts; POST /api/v1/me/contacts/{id}/verification queues a verification message; POST /api/v1/me/contacts/{id}/verify consumes the eight-digit code; DELETE /api/v1/me/contacts/{id} removes a manually managed contact. These endpoints derive tenant and user from authentication. The admin user roster and user pickers do not expose contact addresses or verification metadata.

Verification is a security message addressed only to the requesting user, outside tenant notification rules and future personal mute preferences. It requires the existing SMTP relay and a stable vault key on both API and worker. The API returns 503 when SMTP is unavailable; it never claims that a contact was verified or an email sent. The code expires in 10 minutes; only five attempts are allowed. Resends are at least 60 seconds apart, with at most five requests per hour per user. A replacement challenge, deletion, disable, or email change invalidates pending delivery. The challenge code is encrypted at rest, and the durable job contains only its challenge ID. A worker retry can produce a duplicate mail if the SMTP relay accepted the first one before its acknowledgment was recorded.

A verified email and email_verified: true claim from the signed OIDC ID token creates an identity-provider-managed contact for the account already resolved by issuer and subject. Missing, false, null, and string "true" do not confer verification. A later unverified claim revokes the prior provider-managed contact; changed email increments its version. Email is never used to merge accounts.

Enable the framework

Both synapse-api and synapse-worker need the same PostgreSQL database and the same 32-byte vault key:

SYNAPSE_NOTIFICATIONS_ENABLED=true
SYNAPSE_DB_DSN=postgres://...
SYNAPSE_VAULT_MASTER_KEY=<64 hexadecimal characters or base64 for 32 bytes>
SYNAPSE_WORKER_PROFILE=all

The feature refuses to start without a stable vault key. Full webhook and Slack URLs are encrypted because their paths may contain credentials. List and history responses show only redacted destinations.

Email uses one relay controlled by the deployment operator. Configure SYNAPSE_NOTIFICATION_SMTP_HOST, SYNAPSE_NOTIFICATION_SMTP_PORT, SYNAPSE_NOTIFICATION_SMTP_FROM, and optional username/password. STARTTLS is required by default. Tenant administrators choose recipients but cannot redirect SMTP traffic to another relay.

Use the all worker profile; the specialized lifecycle and integrations profiles do not consume notifications. The framework requires PostgreSQL and is disabled by default. The development memory store does not emulate durable delivery. SYNAPSE_NOTIFICATION_SMTP_REQUIRE_TLS=false is intended only for a controlled development relay. Production relays must support verified STARTTLS.

Configure routing

Open Settings → Alerting. Create a channel, test it, then create rules for one of these events:

  • vulnerability_action.created
  • scan.completed
  • quality_gate.failed
  • sla.approaching_deadline
  • fleet.agent.offline
  • incident.created
  • finding.ownership_changed (requires explicit team_ids or all_teams scope)

Each event type accepts only the rule filters its producer can satisfy, and a rule with any other filter is rejected when it is saved:

Event Filters
vulnerability_action.created severity floor, action types, engagements
scan.completed engagements
quality_gate.failed none
sla.approaching_deadline engagements, lead time (24 hours by default)
fleet.agent.offline none
incident.created severity floor, engagements (when the incident has one)
finding.ownership_changed engagements, teams (required)
notification.destination_changed none; operator-driven, and a rule cannot target it

The severity floor is inclusive. Quality gate and fleet events carry no engagement, so an engagement scope on them could never match. Rules of that shape saved before this check were disabled on upgrade with disabled_reason: engagement_filter_unsupported; their engagement list is kept so you can see what was intended. Remove the engagement scope and save the rule to enable it again.

An incident carries the engagement its fleet correlation was scoped to. Incidents recorded before correlation was scoped to an engagement may carry none, and an engagement-scoped incident rule does not see them. Leave the scope empty to receive every incident.

Engagement and team scope use searchable pickers over the existing engagement list and ownership team pages. Each choice keeps its stable ID beside the display name, including when two records share a name. A saved ID that the directory no longer returns stays on the rule, with a warning, until an administrator removes it. Removing it is the only way the rule becomes broader. Archived teams can remain selected and are marked archived. Personal recipient roles are not offered here: the event catalog that decides which roles an event supports is still open (#1339), and a rule must not grow a recipient filter ahead of that contract.

The engagement list endpoint returns the tenant's engagements in one response, and the picker requests that list once, then shows 25 matches at a time. Team search walks the existing cursor pages and does not download a user directory. The picker does not include email addresses or contact verification state.

Personal inbox

When notifications are enabled, each human user has an inbox at /inbox and a bell in the application header. GET /api/v1/me/inbox and GET /api/v1/me/inbox/unread are scoped to the signed-in user. Machine roles are denied. The bell polls at most every 30 seconds and pauses while the tab is hidden. A deployment without the inbox returns 404 and the bell stops asking.

Inbox rows are written in the same database transaction as the notification event. In-app delivery does not create a channel delivery or a job. Replaying an event after retention does not recreate a deleted row. Mark-all-read uses the server time of that request, so a message that arrives while the request is running stays unread.

Personal recipients come from structured IDs already on the event: the canonical finding assignee and active ownership team members. A legacy assignee label, an email address, or a display name is never resolved into a recipient. Mentioned users, approvers, and engagement leads stay unsupported until a producer records a verified identity. notification.destination_changed is mandatory in-app for enabled tenant admins. It is not a routing rule and cannot be muted. One notice is stored per channel revision: repeating that save is a no-op, and changing only the secret or the URL path is not a host change. Changing back to an earlier host writes a new notice. The payload contains the actor, the action, the channel class, and the scheme plus host. It does not contain a URL path, query, port secret, or credential.

PUT /api/v1/me/notification-preferences stores inherit, enabled, or disabled for the signed-in user. Mandatory in-app wins over an explicit mute, and an explicit mute wins over the default for every other choice. Personal email is sent only to the verified contact version captured when the event was projected. A later email change does not retarget a message that is still queued. Personal delivery is currently available for finding ownership changes, approaching SLAs, and destination-change notices. Other framework event types, Slack direct messages, and Teams personal delivery are shown as unavailable until they have a structured personal recipient and subject.

Events created before the framework first activates for a tenant are not replayed automatically.

Only tenant administrators (PermAdminister) can read or change these settings, test channels, or inspect history. Channel type is immutable. Editing a URL or HMAC key creates a new encrypted version; pending deliveries retain their original version. Leaving both fields blank retains the secret. Email recipients are snapshotted individually when an event is routed. Rule updates require the current revision and apply to events that have not yet been routed.

Source coverage and deduplication

Event Persistence boundary and stable source identity
Vulnerability action Existing risk notification outbox ID; severity comes from the transition's immutable after-assessment. Requires the existing vulnerability notifications flag and dry-run disabled. This covers risk actions, not every raw scanner finding.
Scan completed Successful persisted scan_jobs ID, including synchronous and asynchronous SCA scans with the job store configured. Failed/cancelled jobs and standalone CLI scans without persistence do not emit.
Quality gate failed Finalized project analysis ID with gate.Passed=false; an engagement scan without a project analysis does not emit a gate event.
SLA approaching deadline Current assessment ID, deadline and configured lead time. Only open/mitigating lifecycles with a future deadline and a non-exception tier qualify. Each lead time is delivered once.
Fleet agent offline Agent ID and its last successful heartbeat timestamp. The freshness policy is SYNAPSE_FLEET_STALE_AFTER; polling is once per minute. Never-seen agents and agents already offline at initial activation do not page. A new heartbeat starts a new episode.
Incident created Persisted correlated incident ID. Further detections attached to that incident do not create another notification.

Scan, gate and incident records are captured in the same PostgreSQL transaction as their authoritative write. The worker routes each captured record, creates deliveries and durable jobs, and marks the record processed in one transaction. Risk outbox handoff uses the same atomic boundary. Multiple matching rules collapse to one delivery per channel (per recipient for email), with matched rule revisions retained. No-match events are recorded and are not replayed when a rule is added later.

SLA and fleet relevance is checked again immediately before sending. Resolving an SLA, changing its assessment/deadline, entering an exception, passing the deadline, or receiving a new heartbeat cancels the old pending reminder.

Channel tests return 202 with a delivery ID. This means the test is durably queued; inspect Delivery history for the final result. Disabling or deleting a channel cancels pending deliveries. Existing in-flight requests cannot be recalled.

Webhook contract

Signed webhooks receive JSON using schema version 1 and these headers:

X-Synapse-Timestamp: <unix seconds>
X-Synapse-Signature: sha256=<HMAC-SHA256(timestamp + "." + exact body)>
X-Synapse-Event-ID: <stable event id>
X-Synapse-Delivery-ID: <stable destination delivery id>

Receivers should reject stale timestamps, compare signatures in constant time, and deduplicate by delivery ID. A worker can crash after the receiver accepts a request but before success is persisted, so HTTP delivery is at least once rather than exactly once. Redirects and private, loopback, link-local, metadata, and DNS-rebound destinations are blocked by the HTTP transport.

Slack uses a fixed Block Kit message and observes Slack's 429 Retry-After. Email creates one delivery per normalized recipient and uses a stable Message-ID. SMTP acceptance means the relay accepted the message; it does not prove inbox delivery.

Retry and cutover behavior

Network errors, HTTP 408/429/5xx, and SMTP 4xx responses retry with exponential backoff and a one-hour cap. Other HTTP 4xx and SMTP 5xx responses are terminal. Each attempt is recorded without response bodies or secret-bearing error strings.

The queue allows at most eight attempts, starting at ten seconds with deterministic ±10% jitter; valid Retry-After values are capped at one hour. Throttling reschedules without consuming the attempt budget. An unfinished started attempt indicates an unknown outcome after interruption. Acknowledged success is persisted before queue completion, so replay of that committed success does not send again. Audit intents are committed with results and retried independently of transport delivery.

Limits per tenant are 50 channels, 200 rules, and 10,000 pending/retrying deliveries before another event fan-out is admitted (one admitted fan-out can add up to 2,500). Sending is limited to ten attempts/second per tenant and one/second per channel, across workers. Channel tests are limited to ten/minute/channel. API bodies are limited to 32 KiB, event data to 16 KiB, email recipients to 50/channel, and history pages to 200. Queue saturation rolls back source handoff for a later poll.

When API metrics are enabled, synapse_notification_jobs{state="queued|claimed|failed|done"} exposes aggregate backlog and terminal jobs without tenant or destination labels. synapse_notification_queue_scrape_error indicates unavailable statistics.

The worker exports delivery health on its own, isolated /metrics listener when SYNAPSE_NOTIFICATIONS_ENABLED=true and SYNAPSE_METRICS_ENABLED=true. Use SYNAPSE_METRICS_ADDR=127.0.0.1:9091 for a worker running beside the API; the default 9090 port is already used by the API on the same host. The listener is unauthenticated: allow only a loopback or private scrape network. For inClusterBroker Helm installations, set worker.metrics.enabled=true, enable notifications via extraEnv, and scrape the worker-only ClusterIP Service from worker.metrics.monitoringNamespace. The chart creates a matching ingress NetworkPolicy; the default remains disabled. The integrations and lifecycle worker profiles do not process channel deliveries and do not serve these metrics. The API never duplicates worker delivery counters.

Worker metric names all start with synapse_notification_worker_: - sent_total and failed_total count committed delivery attempts. Failed includes attempts scheduled to retry, not just final delivery failures. - dead_lettered_total counts committed terminal transitions observed by the worker, separately from the unsuccessful attempt that caused them. - delivery_duration_seconds is a histogram of attempt processing time; it deliberately excludes the time a delivery waits in the durable queue. - oldest_pending_age_seconds shows the oldest pending/retrying delivery age across all tenants, computed from tenant-scoped PostgreSQL reads on each scrape. An empty family reports 0. On any database read error all age gauges are omitted and pending_scrape_error reports 1 (0 on a healthy scrape). - template_fallback_total counts committed attempts whose email/Slack rendering had to use built-in title or summary fallback content. Generic webhooks send the event JSON and do not render that content.

Every per-channel metric has only channel_type and provider labels. The fixed combinations are webhook/generic, slack/slack, email/smtp, and other/other for unknown types. No tenant ID, channel ID, recipient, host, URL, provider credentials or raw error text can become a metric label. Counters restart with each worker process; the pending-age gauge reads durable state at scrape time.

The older SYNAPSE_ALERT_WEBHOOK_URL incident path remains available. When it is configured, the worker suppresses the new incident.created producer so an incident is not sent through both paths. Remove the legacy URL after equivalent tenant rules and channels have been tested.

Use the same legacy URL configuration on API and worker during cutover. Let the worker drain incident source records while legacy delivery is active, then stop the API and worker together, remove the legacy URL from both, and restart. Retained unprocessed incident records from a period without a worker can otherwise be sent after cutover. Initial activation is persisted per tenant; disabling the feature pauses delivery rather than deleting retained records or resetting that cutoff.

Secret rotation and retention

Use channel editing to rotate destinations and signing keys. Keep the previous receiver key valid until its pending deliveries have completed. Channel deletion is soft deletion; immutable event, attempt and encrypted version history remains. This release has no automatic retention purge or manual redrive API. Monitor database size and retain the vault master key with database backups. Replacing the master key directly makes existing ciphertext unreadable; master-key rotation requires an operator-controlled offline decrypt/re-encrypt migration of every retained channel version using the same tenant/channel/version associated data.