From cdc9a38013bab7c743074ad0e9cde9432bceb686 Mon Sep 17 00:00:00 2001 From: CChislett Date: Wed, 30 Sep 2026 17:27:14 +0100 Subject: [PATCH] Apply Shane's review: imported rules, state filter, and corrections Adds an Imported Rules page, from Shane's write-up. Rules carried over from FusionReactor Alerts keep their original format, with the condition inside the query rather than in a Threshold step, and a prometheus_math Math step between Query and Threshold. The page covers recognizing one, why the chart draws a threshold of 0, the imported vs standard comparison, how to update one, and the two edge cases that change afterwards. Linked from the Alerting overview and Alert Rules. Corrections from the same review: - The state filter starts on Firing, Error, Pending and Recovering, so healthy rules are not listed by name until it is widened. A card's chips count everything it holds regardless of the filter, which is why a card can read Normal 11 while listing one Pending rule. The Status page previously described this as promotion rather than filtering - Recovering no longer refers to a Keep firing for period, and the flapping fix that recommended that field is removed. The field was never added to the rule form - The expanded rule view shows Namespace, not Data source - Updating an imported rule on a duplicate is a recommendation rather than the procedure - Contact points gain a note on OpsGenie's removal, pointing at Jira Service Management and Atlassian's migration guide Co-Authored-By: Claude Opus 5 (1M context) --- .../Features/New-alerting/contact-points.md | 3 + .../Features/New-alerting/imported-rules.md | 70 +++++++++++++++++++ .../Features/New-alerting/overview.md | 1 + .../Features/New-alerting/rules.md | 7 +- .../Features/New-alerting/status.md | 8 ++- .../Features/New-alerting/troubleshooting.md | 5 +- mkdocs.yml | 1 + 7 files changed, 87 insertions(+), 8 deletions(-) create mode 100644 docs/Data-insights/Features/New-alerting/imported-rules.md diff --git a/docs/Data-insights/Features/New-alerting/contact-points.md b/docs/Data-insights/Features/New-alerting/contact-points.md index ae0df8f..6ad446f 100644 --- a/docs/Data-insights/Features/New-alerting/contact-points.md +++ b/docs/Data-insights/Features/New-alerting/contact-points.md @@ -38,6 +38,9 @@ Use **Search contact points** to find one by name. Use the **type filter** dropd | **Email** | Send notifications by email | | **Kafka REST Proxy** | Publish notifications to a Kafka topic | +!!! note "Moving from OpsGenie" + OpsGenie is no longer offered as a contact point type, following Atlassian's decision to retire the product. If you previously sent notifications to OpsGenie, use **Jira Service Management** instead. Atlassian's [OpsGenie migration guide](https://www.atlassian.com/software/opsgenie/migration) covers moving your on-call setup across. + ## Adding a contact point !!! tip "You can also add one while creating a rule" diff --git a/docs/Data-insights/Features/New-alerting/imported-rules.md b/docs/Data-insights/Features/New-alerting/imported-rules.md new file mode 100644 index 0000000..62b537b --- /dev/null +++ b/docs/Data-insights/Features/New-alerting/imported-rules.md @@ -0,0 +1,70 @@ +# Imported Alert Rules + +Alert rules you created in FusionReactor Alerts were carried over to OpsPilot Alerting in their original format. They keep working exactly as they always did. + +You can leave them alone, or update them to the standard format used by rules created in OpsPilot. Nothing forces the change - this page explains how to tell the two apart, what differs, and how to convert one safely if you decide to. + +## Recognizing an imported rule + +Two things give an imported rule away: + +- **Its query ends in a comparison**, such as `up{job="api"} == 0` or `sum(rate(http_errors_total[5m])) > 5`. The condition that makes the rule fire is part of the query itself. +- **In [Advanced mode](rules.md#advanced-mode), it has a Math step named `prometheus_math`** between the Query step and the Threshold step. + +That Math step checks only whether the query returned anything at all. It returns `1` if it did and `0` if it didn't, and the Threshold step then fires on anything above `0`. + +!!! note "Why the chart shows a threshold of 0" + This is why an imported rule's graph draws its threshold line at `0`, even for a rule named something like *CPU Process Usage > 90%*. The real condition is inside the query - the Threshold step is only asking whether the query returned a result. + +## Why imported rules keep their original format + +Imported rules were not converted automatically, because the standard format behaves slightly differently. Converting them without asking could change when your alerts fire, so the decision is left to you, rule by rule. + +| | Imported format | Standard format | +|---|---|---| +| **Where the condition lives** | Inside the query (`... == 0`) | In the Threshold step | +| **What the query returns** | Only the values that already meet the condition | Every value, whether it meets the condition or not | +| **Chart on the alert page** | Shows data only while the rule is firing | Shows the full history, with the threshold drawn on it | +| **Alert instances** | Appear only while firing | Always listed, each in its current state | +| **No data setting** | Must stay at **Normal**, because no data is how the rule reports *all clear* | Your choice: **No Data**, **Alerting**, **Normal**, or **Keep last state** | + +## Updating an imported rule + +Updating a rule means moving the condition out of the query and into the Threshold step, then removing the `prometheus_math` Math step. + +!!! tip "Consider working on a duplicate" + You can edit the rule directly, but on a rule carried over from the old system it is safer to select **Duplicate** first and change the copy. The original keeps alerting while you confirm the copy behaves the way you want, and you remove the original only once you are sure. + +1. Open the rule - or select **Duplicate** and open the copy - and select **Advanced**. +2. In the **Query** step, delete the comparison from the end of the query - for example, change `up{job="api"} == 0` to `up{job="api"}`. Note the comparison and the number you removed. +3. Remove the **Math** step named `prometheus_math`. +4. In the **Threshold** step, set the input to the Query step. Then set the comparison and the number you noted in step 2. +5. Set **No data** and **On error** to what you want the rule to do in those cases. See [What changes after you update](#what-changes-after-you-update). +6. Check the preview chart. It should now show your full data, with the threshold drawn on it. +7. Select **Save rule**. + +If you worked on a duplicate, it is saved with *(copy)* at the end of its name. Let both rules run side by side and compare when each one fires - while both are running, you receive notifications from each. + +When the copy behaves the way you want, pause or delete the original rule, and rename the copy if you like. + +### Choosing a comparison + +The Threshold step offers **Is above**, **Is below**, **Is within range**, and **Is outside range**. If your rule used `==`, `!=`, `>=`, or `<=`, choose the closest option and adjust the number. + +!!! example + `up == 0` becomes **Is below** `1`, because `up` is only ever `0` or `1`. + +## What changes after you update + +An updated rule fires on the same condition, but two edge cases may behave differently than before. Check both before relying on the updated rule. + +**No data now means no data.** In the imported format, an empty result meant *inactive* - the rule would not fire. In the standard format, an empty result means the query found nothing at all, for example because a service stopped reporting. You choose what happens then with the **No data** setting: raise a No Data alert, fire the alert, resolve it, or keep its current state. + +**NaN and infinite values may not be caught.** The imported rule fired on any value the query returned, including NaN (not a number) and infinity. A Threshold step compares numbers, so a NaN value never crosses it. If your data can produce these values, handle them in the query itself. + +These differences are why the update is left to you. If a rule already does what you need, you can keep it in its imported format. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/New-alerting/overview.md b/docs/Data-insights/Features/New-alerting/overview.md index 9cbd4c0..14a8e2b 100644 --- a/docs/Data-insights/Features/New-alerting/overview.md +++ b/docs/Data-insights/Features/New-alerting/overview.md @@ -11,6 +11,7 @@ Navigate to **Alerting** in the left-hand menu to open it. | [Status](status.md) | See every alert rule's state at a glance, and find what needs attention first | | [Alert Rules](rules.md) | Build, manage, and investigate static alert rules | | [Annotation Templates](annotation-templates.md) | Write alert messages that carry live values from the query that fired | +| [Imported Rules](imported-rules.md) | Recognize rules carried over from FusionReactor Alerts, and update them if you want to | | [Recording Rules](recording-rules.md) | Pre-compute expensive queries and save the result as a new metric | | [Service Anomaly Detectors](service-anomaly-detectors.md) | Tune the detectors created automatically for each instrumented service | | [Custom Anomaly Detectors](custom-anomaly-detectors.md) | Run anomaly detection against your own PromQL series | diff --git a/docs/Data-insights/Features/New-alerting/rules.md b/docs/Data-insights/Features/New-alerting/rules.md index 300c965..484785e 100644 --- a/docs/Data-insights/Features/New-alerting/rules.md +++ b/docs/Data-insights/Features/New-alerting/rules.md @@ -7,6 +7,9 @@ Navigate to **Alerting > Alert Rules** to open it. !!! info "Rules vs detectors" **Rules** are static checks - they run on a fixed schedule against fixed thresholds, best for known conditions with clear boundaries (like system CPU or allocated memory). **[Detectors](service-anomaly-detectors.md)** use AI to learn normal behavior and flag anomalies automatically, so they adapt as your system changes. +!!! note "Rules carried over from FusionReactor Alerts" + Rules you created in FusionReactor Alerts were imported in their original format and work differently from rules built here - their condition sits inside the query rather than in a Threshold step. See [Imported Rules](imported-rules.md) to recognize one and to update it if you want to. + ## The rules list ![Screenshot](/Data-insights/Features/images/Alerting/rule-table.png) @@ -60,7 +63,7 @@ The **Dashboard** and **Runbook** buttons (top right of the expanded view) open | **Annotations** | The annotations from the rule (such as its description) | | **Expression** | The query and threshold condition as chained steps - for example, `A` `max_over_time(up[5m])` feeding a `C` condition `< 1` | | **Evaluation** | How often the rule is checked and the pending duration (such as, `every 60s · pending 5m`) | -| **Data source** | The data source the rule queries | +| **Namespace** | The namespace the rule is stored in | | **On no data** | What state the rule enters when the query returns no data | | **On query error** | What state the rule enters when the query fails | | **Labels** | The rule's labels as name/value chips (such as, `severity = warning`) - these are what [notification policies](notification-policy.md) route on. Only shown when the rule has labels | @@ -124,7 +127,7 @@ A row of summary cards sits below the header: | **Annotations** | The annotations from the rule (such as its description) | | **Expression** | The query and threshold condition as chained steps - a **Query** step (with its data source) feeding a **Threshold** step | | **Evaluation** | How often the rule is checked and the pending duration (such as, `every 60s · pending 5m`) | -| **Data source** | The data source the rule queries | +| **Namespace** | The namespace the rule is stored in | | **On no data** | What state the rule enters when the query returns no data | | **On query error** | What state the rule enters when the query fails | | **Notifies** | The contact points configured to receive notifications | diff --git a/docs/Data-insights/Features/New-alerting/status.md b/docs/Data-insights/Features/New-alerting/status.md index ee2f6d6..f5162f1 100644 --- a/docs/Data-insights/Features/New-alerting/status.md +++ b/docs/Data-insights/Features/New-alerting/status.md @@ -19,7 +19,7 @@ The band always covers every rule on the account. Filtering the page to a single | **Firing** | The alert condition is met and the rule is actively firing | | **Error** | The rule's query failed to evaluate | | **Pending** | The condition has been met, but not yet for long enough to fire | -| **Recovering** | The condition is no longer met, but the rule is still held by its **Keep firing for** period before it returns to Normal. See [Alert is flapping](troubleshooting.md#alert-is-flapping) | +| **Recovering** | The condition is no longer met and the rule is on its way back to Normal | | **Normal** | The rule is evaluating and its condition is not currently met | | **No Data** | The query returned no data, so there was nothing to evaluate | | **Paused** | The rule is paused and not being evaluated | @@ -35,7 +35,7 @@ Controls across the top shape how the page is laid out: | **List / Grid** | Switch between the list view and a compact grid view | | **Source / Namespace** | Group the cards by data source or by namespace | | **All sources** / **All namespaces** | Narrow the page to particular sources or namespaces - the label follows the grouping toggle. Open it for a searchable list, with **Select all** to take everything and **Clear all** to start again. Selected entries are ticked in the list, and the button itself becomes a chip naming your selection, with an **✕** to remove it | -| **State filter** | A multi-select to show only rules in the chosen states, with a search box and **Select all**. It reads **All states** when nothing is excluded, and **N selected** once you narrow it. Click **✕** to clear it | +| **State filter** | A multi-select controlling which rules are listed by name, with a search box and **Select all**. It starts on **Firing**, **Error**, **Pending** and **Recovering** - the states that need you - and reads **4 selected**. Widen it to **All states** to list healthy rules too. Click **✕** to clear it | | **Expand all** | Expand every group to show all of its rules; it toggles to **Shrink all** to collapse them | | **Hide filtered-out cards** | Hide the groups and rules that don't match the current filters | | **Refresh interval** | How often the page auto-refreshes (for example, **30s**) | @@ -46,7 +46,9 @@ Controls across the top shape how the page is laid out: Rules are organized into cards - by **namespace** or by **source**, depending on the toggle. Each group card shows its name, a health bar, and a chip with a count for each state present. -A card holding rules that need attention is **highlighted** and moved to the front of the page, and those rules are listed on the card without you expanding it. Healthy rules stay folded away behind **Show all N rules**, so what needs looking at is what you see first. +**The chips count everything the card holds; the state filter decides what is listed by name.** Because that filter starts on Firing, Error, Pending and Recovering, a card shows only the rules wanting attention, while its chips still account for the healthy ones. A card reading *Normal 11* with a single Pending rule listed is working as intended - widen the filter to **All states** to see the rest by name. + +A card holding rules that need attention is also **highlighted** and moved to the front of the page, so what needs looking at is what you see first. Each rule shows: diff --git a/docs/Data-insights/Features/New-alerting/troubleshooting.md b/docs/Data-insights/Features/New-alerting/troubleshooting.md index 9fcbcf4..18461ff 100644 --- a/docs/Data-insights/Features/New-alerting/troubleshooting.md +++ b/docs/Data-insights/Features/New-alerting/troubleshooting.md @@ -100,9 +100,8 @@ This page covers common issues when setting up and operating OpsPilot alerting, **Fix:** 1. **Increase the Pending period** - A pending period of `5m` or `10m` requires the condition to be continuously true before firing, smoothing out brief spikes. -2. **Use "Keep firing for"** - This holds the alert in a firing state for a period after the condition resolves, preventing rapid recovered/re-fired cycles. -3. **Adjust the threshold** - If the metric hovers just at the threshold, add a buffer (such as, changing `> 80` to `> 85`). -4. **Increase the Repeat interval** on the [notification policy](notification-policy.md) - This reduces re-notification frequency without changing how the rule evaluates. +2. **Adjust the threshold** - If the metric hovers just at the threshold, add a buffer (such as, changing `> 80` to `> 85`). +3. **Increase the Repeat interval** on the [notification policy](notification-policy.md) - This reduces re-notification frequency without changing how the rule evaluates. --- diff --git a/mkdocs.yml b/mkdocs.yml index 0ad022e..f3ec264 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -132,6 +132,7 @@ nav: - Alert Rules: - Overview: Data-insights/Features/New-alerting/rules.md - Annotation Templates: Data-insights/Features/New-alerting/annotation-templates.md + - Imported Rules: Data-insights/Features/New-alerting/imported-rules.md - Recording Rules: Data-insights/Features/New-alerting/recording-rules.md - Service Anomaly Detectors: Data-insights/Features/New-alerting/service-anomaly-detectors.md - Custom Anomaly Detectors: Data-insights/Features/New-alerting/custom-anomaly-detectors.md