Detect & alert

Get told the moment something looks wrong. Define a rule, and when it fires you get a tracked event you can acknowledge and resolve — and a notification wherever you want it.

Rules that watch your data

An AlertRule evaluates one sensor against a condition. Three evaluation strategies are built in — pick the one that matches the question you're asking:

StrategyFires when…Best for
Single valueone reading crosses a thresholdhard limits — "alarm over 8000 W"
Time windowan aggregate (avg / min / max / count) over a sliding window crosses a thresholdsustained conditions — "average above 60 for 5 minutes"
Statistical anomalya reading strays too far from the sensor's own recent normal (rolling mean ± Nσ)drift you never set a threshold for

Every rule has a severity (INFO, WARNING, or CRITICAL), a cooldown, and a minimum consecutive breaches count — see Debounce and cooldown.

Single value

The workhorse. Each incoming reading is compared against your threshold with an operator matched to the sensor's value type:

Value typeOperators
Numeric (DOUBLE, BIGINT)GT, GTE, LT, LTE, EQ, NEQ
BooleanIS_TRUE, IS_FALSE
TextEQ, NEQ, CONTAINS, NOT_CONTAINS, STARTS_WITH, MATCHES_REGEX

Create it with POST /api/v1/alert-rules:

{
  "name": "Motor temperature over 50 °C",
  "sensorId": "<sensor id>",
  "evaluationStrategy": "SINGLE_VALUE",
  "operator": "GT",
  "thresholdValue": 50,
  "severity": "WARNING",
  "cooldownDurationMs": 300000,
  "minConsecutiveBreaches": 1,
  "channelIds": ["<channel id>"]
}

Time window

Aggregates the readings inside a sliding window — AVG, MIN, MAX, or COUNT — and compares the result against the threshold. The window always ends "now": a rule with a 5-minute window asks "what was the average over the last 5 minutes?" each time a new reading arrives. Numeric sensors only. Same endpoint, two extra fields:

{
  "name": "Humidity average high (5 min)",
  "sensorId": "<sensor id>",
  "evaluationStrategy": "TIME_WINDOW",
  "windowAggregation": "AVG",
  "windowDurationMs": 300000,
  "operator": "GT",
  "thresholdValue": 60,
  "severity": "WARNING",
  "cooldownDurationMs": 300000,
  "minConsecutiveBreaches": 1,
  "channelIds": ["<channel id>"]
}

A time-window rule fires the moment the aggregate crosses the line — the alert's trigger value is the actual window aggregate at that instant (say, 62.7), not the single reading that tipped it.

Statistical anomaly

Learns each sensor's normal range from its recent history and fires on readings that are unusual for that sensor — no fixed number to guess. It has its own configuration (sensitivity in σ, baseline window, warm-up, direction) and its own page: Statistical anomaly detection.

Debounce and cooldown

Two controls, on every strategy, decide how twitchy a rule is:

  • Min consecutive breaches — how many breaching readings in a row before the rule fires. One non-breaching reading resets the count. This is the most effective control against one-off spikes (electrical transients, glitches). A rule with 3 stays silent through high, high, normal, high, high — and fires exactly on the third consecutive high.
  • Cooldown — the minimum time between fires of the same rule. While a rule is cooling down, further breaches are suppressed entirely (no new event, no notification).

Fired alerts are tracked events

When a rule fires, the platform records an Event with an EventLifecycle that operators work through: ACTIVE → ACKNOWLEDGED → RESOLVED. The Event log is generic — alert firings are one kind of event among others — so everything that happens to your fleet lands in one timeline.

Notifications go where you work

A NotificationChannel is a delivery target owned by the organization. Two channel types are built in — email and webhook — and a rule can fan out to any number of them. You can send a test message before relying on a channel.

Email

Delivers a rendered subject and body (HTML or plain text) to one or more recipients. Create it with POST /api/v1/notification-channels:

{
  "name": "Ops email",
  "channelType": "EMAIL",
  "enabled": true,
  "config": {
    "type": "EMAIL",
    "recipients": ["[email protected]"],
    "subjectTemplate": "[{{severity}}] {{ruleName}}",
    "bodyTemplate": "{{ruleName}} fired.\nValue: {{triggerValue}} (threshold {{thresholdValue}})\nAt: {{triggeredAt}}"
  }
}

Webhook

Calls your endpoint with a payload you define — the body template is the request body, so it can be shaped for whatever receives it (a chat webhook, a ticketing system, your own service). Method defaults to POST; custom headers are supported and their values are templated too, so a header can carry the severity for routing. Same endpoint:

{
  "name": "Ops chat",
  "channelType": "WEBHOOK",
  "enabled": true,
  "config": {
    "type": "WEBHOOK",
    "url": "https://chat.example.com/hooks/T000/B000",
    "method": "POST",
    "headers": { "X-Severity": "{{severity}}" },
    "bodyTemplate": "{\"text\": \"{{ruleName}}: {{triggerValue}} ({{severity}})\"}"
  }
}

Template variables

Subject, body, and webhook header templates use {{variable}} placeholders. Unknown placeholders are left as-is.

VariableMeaning
{{alertId}}Unique id of the fired event (== alert id externally)
{{ruleId}}Configured alert-rule id that triggered
{{ruleName}}Human-readable rule name
{{sensorId}}Sensor whose reading crossed the rule threshold
{{severity}}Severity level: INFO / WARNING / CRITICAL
{{status}}Lifecycle status at fire time (always ACTIVE)
{{triggerValue}}The measured value that crossed the threshold
{{thresholdValue}}The configured threshold the value crossed
{{triggeredAt}}ISO-8601 instant the event was received
{{mean}}Rolling mean at fire time (statistical-anomaly rules)
{{stddev}}Rolling standard deviation at fire time (statistical-anomaly rules)
{{sigmaDistance}}Standard deviations the reading was from the mean (statistical-anomaly rules)

The last three are N/A for threshold rules — they carry the baseline context behind a statistical anomaly fire.

Delivery and retries

Delivery is tracked per alert and per channel. A failed send is retried up to 3 times with exponential backoff (roughly one and two minutes between attempts), and a background sweep rescues any dispatch that got lost in transit — so a flaky receiver delays a notification rather than losing it. Check what was delivered for an alert with GET /api/v1/alerts/{id}/notifications or on the alert's detail page.

Browser notifications

Separately from org-level channels, each user can opt into push notifications — on the desktop browser, and on iOS once the app is installed to the Home Screen:

  1. Open Notifications (/{slug}/notifications) and choose Enable browser notifications — the browser asks for permission once, and the device appears under Registered devices.
  2. Turn on Notify me of alerts for the organization and pick a minimum severity (INFO, WARNING, or CRITICAL).

From then on, every alert at or above that severity in that organization pushes to all of the user's registered devices. Clicking the notification opens the alert. Browser push is per-user and per-org — it doesn't need a NotificationChannel, and it doesn't affect what channels deliver.

In the app

Work fired alerts under Alerts (/{slug}/alerts) — acknowledge, resolve, and see which notifications went out. Manage delivery channels under Alerting (/{slug}/alerting), and personal browser push under Notifications (/{slug}/notifications).

In the API

  • GET|POST|PUT|DELETE /api/v1/alert-rules — define rules; POST …/{id}/enable and /disable
  • GET /api/v1/alerts — fired alerts; POST …/{id}/acknowledge and …/{id}/resolve
  • GET /api/v1/alerts/{id}/notifications — what was delivered for an alert
  • GET|POST|PUT|DELETE /api/v1/notification-channels — manage channels; POST …/test to try one
  • GET|POST|DELETE /api/v1/push-subscriptions — this user's registered push devices
  • GET|POST /api/v1/events — the generic event timeline

Was this page helpful?