Detect & alert
Get told the moment something looks wrong. Define a rule, and when it fires you get a tracked event you can acknowledge and resolve — and a notification wherever you want it.
Rules that watch your data
An AlertRule evaluates one sensor against a condition. Three evaluation strategies are built in — pick the one that matches the question you're asking:
| Strategy | Fires when… | Best for |
|---|---|---|
| Single value | one reading crosses a threshold | hard limits — "alarm over 8000 W" |
| Time window | an aggregate (avg / min / max / count) over a sliding window crosses a threshold | sustained conditions — "average above 60 for 5 minutes" |
| Statistical anomaly | a reading strays too far from the sensor's own recent normal (rolling mean ± Nσ) | drift you never set a threshold for |
Every rule has a severity (INFO, WARNING, or CRITICAL), a cooldown, and a minimum consecutive breaches count — see Debounce and cooldown.
Single value
The workhorse. Each incoming reading is compared against your threshold with an operator matched to the sensor's value type:
| Value type | Operators |
|---|---|
Numeric (DOUBLE, BIGINT) | GT, GTE, LT, LTE, EQ, NEQ |
| Boolean | IS_TRUE, IS_FALSE |
| Text | EQ, NEQ, CONTAINS, NOT_CONTAINS, STARTS_WITH, MATCHES_REGEX |
Create it with POST /api/v1/alert-rules:
{
"name": "Motor temperature over 50 °C",
"sensorId": "<sensor id>",
"evaluationStrategy": "SINGLE_VALUE",
"operator": "GT",
"thresholdValue": 50,
"severity": "WARNING",
"cooldownDurationMs": 300000,
"minConsecutiveBreaches": 1,
"channelIds": ["<channel id>"]
}
Time window
Aggregates the readings inside a sliding window — AVG, MIN, MAX, or COUNT — and compares the result against the threshold. The window always ends "now": a rule with a 5-minute window asks "what was the average over the last 5 minutes?" each time a new reading arrives. Numeric sensors only. Same endpoint, two extra fields:
{
"name": "Humidity average high (5 min)",
"sensorId": "<sensor id>",
"evaluationStrategy": "TIME_WINDOW",
"windowAggregation": "AVG",
"windowDurationMs": 300000,
"operator": "GT",
"thresholdValue": 60,
"severity": "WARNING",
"cooldownDurationMs": 300000,
"minConsecutiveBreaches": 1,
"channelIds": ["<channel id>"]
}
A time-window rule fires the moment the aggregate crosses the line — the alert's trigger value is the actual window aggregate at that instant (say, 62.7), not the single reading that tipped it.
Statistical anomaly
Learns each sensor's normal range from its recent history and fires on readings that are unusual for that sensor — no fixed number to guess. It has its own configuration (sensitivity in σ, baseline window, warm-up, direction) and its own page: Statistical anomaly detection.
Debounce and cooldown
Two controls, on every strategy, decide how twitchy a rule is:
- Min consecutive breaches — how many breaching readings in a row before the rule fires. One non-breaching reading resets the count. This is the most effective control against one-off spikes (electrical transients, glitches). A rule with
3stays silent through high, high, normal, high, high — and fires exactly on the third consecutive high. - Cooldown — the minimum time between fires of the same rule. While a rule is cooling down, further breaches are suppressed entirely (no new event, no notification).
Rules created in the app default to 3 consecutive breaches — a sensor reporting every 10 seconds fires about 30 seconds after a condition starts. Set it to 1 if you want first-breach firing.
Fired alerts are tracked events
When a rule fires, the platform records an Event with an EventLifecycle that operators work through: ACTIVE → ACKNOWLEDGED → RESOLVED. The Event log is generic — alert firings are one kind of event among others — so everything that happens to your fleet lands in one timeline.
Notifications go where you work
A NotificationChannel is a delivery target owned by the organization. Two channel types are built in — email and webhook — and a rule can fan out to any number of them. You can send a test message before relying on a channel.
Delivers a rendered subject and body (HTML or plain text) to one or more recipients. Create it with POST /api/v1/notification-channels:
{
"name": "Ops email",
"channelType": "EMAIL",
"enabled": true,
"config": {
"type": "EMAIL",
"recipients": ["[email protected]"],
"subjectTemplate": "[{{severity}}] {{ruleName}}",
"bodyTemplate": "{{ruleName}} fired.\nValue: {{triggerValue}} (threshold {{thresholdValue}})\nAt: {{triggeredAt}}"
}
}
Webhook
Calls your endpoint with a payload you define — the body template is the request body, so it can be shaped for whatever receives it (a chat webhook, a ticketing system, your own service). Method defaults to POST; custom headers are supported and their values are templated too, so a header can carry the severity for routing. Same endpoint:
{
"name": "Ops chat",
"channelType": "WEBHOOK",
"enabled": true,
"config": {
"type": "WEBHOOK",
"url": "https://chat.example.com/hooks/T000/B000",
"method": "POST",
"headers": { "X-Severity": "{{severity}}" },
"bodyTemplate": "{\"text\": \"{{ruleName}}: {{triggerValue}} ({{severity}})\"}"
}
}
Template variables
Subject, body, and webhook header templates use {{variable}} placeholders. Unknown placeholders are left as-is.
| Variable | Meaning |
|---|---|
{{alertId}} | Unique id of the fired event (== alert id externally) |
{{ruleId}} | Configured alert-rule id that triggered |
{{ruleName}} | Human-readable rule name |
{{sensorId}} | Sensor whose reading crossed the rule threshold |
{{severity}} | Severity level: INFO / WARNING / CRITICAL |
{{status}} | Lifecycle status at fire time (always ACTIVE) |
{{triggerValue}} | The measured value that crossed the threshold |
{{thresholdValue}} | The configured threshold the value crossed |
{{triggeredAt}} | ISO-8601 instant the event was received |
{{mean}} | Rolling mean at fire time (statistical-anomaly rules) |
{{stddev}} | Rolling standard deviation at fire time (statistical-anomaly rules) |
{{sigmaDistance}} | Standard deviations the reading was from the mean (statistical-anomaly rules) |
The last three are N/A for threshold rules — they carry the baseline context behind a statistical anomaly fire.
Delivery and retries
Delivery is tracked per alert and per channel. A failed send is retried up to 3 times with exponential backoff (roughly one and two minutes between attempts), and a background sweep rescues any dispatch that got lost in transit — so a flaky receiver delays a notification rather than losing it. Check what was delivered for an alert with GET /api/v1/alerts/{id}/notifications or on the alert's detail page.
Browser notifications
Separately from org-level channels, each user can opt into push notifications — on the desktop browser, and on iOS once the app is installed to the Home Screen:
- Open Notifications (
/{slug}/notifications) and choose Enable browser notifications — the browser asks for permission once, and the device appears under Registered devices. - Turn on Notify me of alerts for the organization and pick a minimum severity (
INFO,WARNING, orCRITICAL).
From then on, every alert at or above that severity in that organization pushes to all of the user's registered devices. Clicking the notification opens the alert. Browser push is per-user and per-org — it doesn't need a NotificationChannel, and it doesn't affect what channels deliver.
In the app
Work fired alerts under Alerts (/{slug}/alerts) — acknowledge, resolve, and see which notifications went out. Manage delivery channels under Alerting (/{slug}/alerting), and personal browser push under Notifications (/{slug}/notifications).
In the API
GET|POST|PUT|DELETE /api/v1/alert-rules— define rules;POST …/{id}/enableand/disableGET /api/v1/alerts— fired alerts;POST …/{id}/acknowledgeand…/{id}/resolveGET /api/v1/alerts/{id}/notifications— what was delivered for an alertGET|POST|PUT|DELETE /api/v1/notification-channels— manage channels;POST …/testto try oneGET|POST|DELETE /api/v1/push-subscriptions— this user's registered push devicesGET|POST /api/v1/events— the generic event timeline
Fired alerts are stored as Event + EventLifecycle records, not as standalone alert rows — so acknowledging and resolving is the same workflow for any kind of event.