diff --git a/docs/Data-insights/Features/New-alerting/contact-points.md b/docs/Data-insights/Features/New-alerting/contact-points.md index 4fb95c4..ae0df8f 100644 --- a/docs/Data-insights/Features/New-alerting/contact-points.md +++ b/docs/Data-insights/Features/New-alerting/contact-points.md @@ -32,7 +32,7 @@ Use **Search contact points** to find one by name. Use the **type filter** dropd | **Telegram** | Send notifications via Telegram | | **Google Chat** | Post notifications to Google Chat | | **PagerDuty** | Create incidents in PagerDuty | -| **OpsGenie** | Create alerts in OpsGenie | +| **Jira Service Management** | Raise issues in Jira Service Management | | **Pushover** | Send push notifications via Pushover | | **Webhook** | POST a JSON payload to any URL | | **Email** | Send notifications by email | @@ -43,13 +43,15 @@ Use **Search contact points** to find one by name. Use the **type filter** dropd !!! tip "You can also add one while creating a rule" Contact points are not only made here. In the rule editor, **Then notify** lets you search your existing contact points, or click **Create new contact point** to make one without leaving the rule. See [Rules](rules.md). +A third route is the **Wizard** on the [Status](status.md) page. Choosing **Contact Point** asks **How should notifications be delivered?** and offers the notification channels as a grid of cards. Picking one takes you into its configuration. As elsewhere in the wizard, **Skip to form** leaves for the full form, **Ask OpsPilot** suggests a channel, **←** goes back and **✕** closes without creating anything. + 1. Click **+ New contact point** to open the integration picker -2. Select an integration type from the grid. Use the category tabs to filter: +2. Select an integration type from the grid. Use the category tabs to filter - each shows how many types it holds: | Category | Integrations | |---|---| | **Chat** | Slack, Discord, Microsoft Teams, Telegram, Google Chat | -| **On-call** | PagerDuty, OpsGenie, Pushover | +| **On-call** | PagerDuty, Jira Service Management, Pushover | | **Webhook** | Webhook | | **Other** | Email, Kafka REST Proxy | diff --git a/docs/Data-insights/Features/New-alerting/custom-anomaly-detectors.md b/docs/Data-insights/Features/New-alerting/custom-anomaly-detectors.md index 5040ba9..667d1cc 100644 --- a/docs/Data-insights/Features/New-alerting/custom-anomaly-detectors.md +++ b/docs/Data-insights/Features/New-alerting/custom-anomaly-detectors.md @@ -16,6 +16,26 @@ Each detector row shows its **State**, **Name**, **Threshold**, and **Last evalu Click **+ New custom detector** (top right of the page), or use the **Wizard** (**New custom detector**) on the [Status](status.md) or [Rules](rules.md) page. Fill in the fields below, then click **Create detector** to save. +### The guided wizard + +Starting from the **Wizard** asks one question at a time instead of presenting the whole form. After choosing **Anomaly Detector** and then **Custom Detector** (see [Creating a detector from the wizard](status.md#creating-a-detector-from-the-wizard)), it asks which metric to watch, and then how sensitive detection should be. + +**How sensitive should detection be?** offers three presets. Higher sensitivity catches more anomalies but may alert more often: + +| Preset | Anomaly threshold | Description | +|---|---|---| +| **Low** | 95% | Fewer alerts, only strong anomalies | +| **Medium** | 90% | Balanced sensitivity | +| **High** | 80% | More alerts, catches subtle changes | + +The percentage is the **anomaly threshold** described under [When to fire](#when-to-fire) - the score at or above which the detector fires. The relationship is inverted, so a *lower* sensitivity sets a *higher* threshold: at **Low**, the model has to be 95% confident before anything fires. + +Whichever preset you pick, you can change the threshold afterwards on the detector itself, so this is a starting point rather than a commitment. + +**Who gets notified?** is the last step. Pick one or more [contact points](contact-points.md) with **+ Add contact point**, or click **Skip** to create the detector without notifications and add them later. Click **Done** to finish. + +As in the [alert rule wizard](rules.md#the-guided-wizard), each step offers **←** to go back, **Skip to form** to leave the wizard for the full form, **Ask OpsPilot** for a recommendation, and **✕** to close without creating anything. **Skip to form** is offered on every step but the last, where **Skip** and **Done** take its place. + ### Signal Defines the PromQL series the detector watches. The query is validated against the Prometheus datasource on save. diff --git a/docs/Data-insights/Features/New-alerting/rules.md b/docs/Data-insights/Features/New-alerting/rules.md index b85f47d..300c965 100644 --- a/docs/Data-insights/Features/New-alerting/rules.md +++ b/docs/Data-insights/Features/New-alerting/rules.md @@ -69,11 +69,30 @@ The **Dashboard** and **Runbook** buttons (top right of the expanded view) open ![Screenshot](/Data-insights/Features/images/Alerting/rule-expanded.png) +### The rule panel + +Clicking a rule on the [Status](status.md) page opens a panel beside the list - a quick look at that rule without leaving the page. Click the **✕** to close it. + +The header shows the rule name, its state and how long it has held it, and the namespace and data source it belongs to (such as, *FusionReactor Alerts / Metrics*), with four actions: + +| Action | Description | +|---|---| +| **Silence** | Create a [silence](silences.md) for this rule | +| **View rule** | Open the full [rule detail view](#rule-detail-view) | +| **Edit rule** | Open the rule editor | +| **Active** | A toggle to pause and resume evaluation | + +Below the header sit two summary cards, **Duration** and **State**, then: + +**Metric** - a graph of the query with the threshold drawn on it. Use the time range picker, its step arrows and the zoom buttons to adjust the window, or click **Open in Explore** to investigate the metric in Explore. Three checkboxes below the graph toggle the **Threshold**, **State transitions** and **Pending window** overlays. + +**State history** - the rule's state changes, newest first. The header gives the period and transition count (such as, *last 24h · 20 transitions*), and **See all →** opens the full history. Each row reads as the new state *from* the previous one - for example, **Pending** from **Normal** - with when it happened, how long the previous state was held, and how long ago that was. + ### Rule detail view ![Screenshot](/Data-insights/Features/images/Alerting/high-cpu-rule.png) -Open the full rule detail view by clicking the **eye** icon (**View rule**) in the rules list, or by clicking a rule on the [Status](status.md) page. The header shows the rule name and current state, with these actions in the top right: +Open the full rule detail view by clicking the **eye** icon (**View rule**) in the rules list, or **View rule** in the [rule panel](#the-rule-panel). The header shows the rule name and current state, with these actions in the top right: | Action | Description | |---|---| @@ -138,6 +157,60 @@ A good alert rule has three things: a query that targets the right signal, a thr Click **+ New rule** (top right) to open the rule editor. (To create an anomaly detector instead, see [Service](service-anomaly-detectors.md) or [Custom Anomaly Detectors](custom-anomaly-detectors.md), or use the **Wizard** on the [Status](status.md) page.) +### The guided wizard + +Starting a rule from the **Wizard** on the [Status](status.md) page walks you through the decisions one at a time, rather than presenting the whole form at once. A row of dots below the heading tracks your progress through the steps. + +You are never locked into the wizard. Each step offers: + +| Control | What it does | +|---|---| +| **←** | Go back to the previous step | +| **Skip to form** | Leave the wizard and go straight to the main rule form | +| **Not sure? Ask OpsPilot** | Get a recommendation for the step you're on | +| **✕** | Close the wizard without creating anything | + +**Skip to form** is offered on every step but the last, where **Skip** and **Done** take its place. + +The steps are: + +**What are you monitoring?** - pick the data source the alert will query. The list holds every data source configured on your account, including those added by [integrations](../integrations.md), so an AWS installation appears here alongside your metrics and logs sources. + +**How do you want to build the query?** - choose how to express the condition: + +| Option | Description | +|---|---| +| **Guided builder** | Build your query step by step | +| **Write PromQL** | Write the query expression directly | + +This choice is not binding - you can switch between the two later in the form. + +**What should we watch?** - pick the metric, or write the query, that the alert will evaluate. What this step shows depends on the choice you made at the previous one: + +- **Guided builder** gives you a **Metric** dropdown. Once you choose a metric, a preview graph appears below it showing that metric's recent behavior, so you can confirm you have the right signal before going further. +- **Write PromQL** gives you a **PromQL expression** box to type the query into directly. + +Either way, click **Next** to continue. The wizard is the same length whichever you pick. + +**When should it fire?** - set the threshold that triggers notifications: + +| Field | Description | +|---|---| +| **Alert when value is** | **Above** or **Below** the threshold | +| **Threshold** | The value to compare against (such as, `80`) | +| **Wait before alerting** | How long the condition must hold before the alert fires - **1m**, **5m**, or **10m** | + +**Wait before alerting** is the pending period under a plainer name. Leaving it at anything above **1m** is what stops a brief spike from paging someone. + +**Who gets notified?** - pick one or more [contact points](contact-points.md) with **+ Add contact point**. This step is optional: click **Skip** to create the rule without notifications and add them later, or **Done** to finish. + +A rule with no contact point still evaluates and still shows its state on [Status](status.md) - it just won't notify anyone. Its expanded view reads *None configured* under **Notifies**. + +!!! note + The earlier steps advance as soon as you pick an option. From **What should we watch?** onward, you make a choice and then click **Next**. + +### Rule editor modes + The rule editor has two modes, toggled in the top right: | Mode | Description | @@ -180,17 +253,26 @@ Build the alert from a chain of **Queries & expressions**. Click **Add query** ( A new rule starts with a default **Query → Reduce → Threshold** chain. The **Alert condition** dropdown at the top selects which step's firing state determines whether the rule alerts. +Each step is labeled with a chip naming its reference ID and type, colored by category, and steps that take input from another show which one they follow (such as, *← $reduce*). The step serving as the alert condition is outlined and carries an **Alert condition** badge, so you can see at a glance which one decides the outcome. + +Reorder steps with the arrows to the left of each one, and remove a step with the **✕** on its right. + +Below the chain, **Evaluation preview** runs the pipeline through the alert condition step and shows the result as it would be evaluated. If a step cannot run, the preview reports the failure and names the step responsible - an empty or malformed query, for example, reads *Couldn't evaluate this rule*. Use it to catch mistakes before saving rather than after the rule goes live. + #### Evaluation The **Evaluation** section controls how the rule runs: | Setting | Description | |---|---| -| **Evaluate every** | How often the rule is checked (such as, `1m`) | +| **Group** | The evaluation group the rule belongs to (such as, `default`) | +| **Every** | How often the rule is checked (such as, `1m`). The schedule belongs to the group rather than to the individual rule | | **Pending for** | How long the condition must be continuously met before the alert fires (such as, `5m`). Prevents notifications for temporary spikes | | **No data** | The state the rule enters when the query returns no data - **No Data**, **Alerting**, **Normal**, or **Keep last state** | | **On error** | The state the rule enters when the query fails - **Error**, **Alerting**, **Normal**, or **Keep last state** | +Tick **Choose the group and schedule myself** to set the group and its interval by hand. Rules in a group all evaluate together on one schedule, so changing it changes every rule in that group - not only the one you are editing. + #### Namespace Expand **Namespace** and choose the **namespace** where the rule is stored. Namespaces keep rules organized and control access. diff --git a/docs/Data-insights/Features/New-alerting/service-anomaly-detectors.md b/docs/Data-insights/Features/New-alerting/service-anomaly-detectors.md index 1ba16bd..e3a4c73 100644 --- a/docs/Data-insights/Features/New-alerting/service-anomaly-detectors.md +++ b/docs/Data-insights/Features/New-alerting/service-anomaly-detectors.md @@ -14,7 +14,7 @@ Each instrumented service automatically gets three detectors, shown as **R**, ** - **Rate** (R) - request rate anomalies - **Error** (E) - error rate anomalies -- **Latency** (D) - response time anomalies (also shown as *Duration* on the Status page) +- **Duration** (D) - response time anomalies (also referred to as *latency*) The badges are filled while the service's detectors are running, and grayed out while they are paused. **Last evaluation** reads a timestamp (such as, *29 Sept, 13:01*) for a running service, and **—** for a paused one. @@ -24,6 +24,10 @@ Each row in the service list shows the service name, its **Detectors** (R/E/D), Click **Scan for services** to detect your instrumented services and auto-create detectors for them. If no service detectors exist yet, this is the first step. +You can also reach it from the **Wizard** on the [Status](status.md) page, by choosing **Anomaly Detector** and then **Service Scan**. That route ends on a confirmation step explaining what the scan will do, with **Scan now** to run it and **Back** to return to the previous step. + +Either way, the scan creates rate, error, and duration detectors for every instrumented service it finds. Sensitivity and notifications are then tuned per detector afterwards, in [Detector settings](#detector-settings). + ## State counters | State | Description | diff --git a/docs/Data-insights/Features/New-alerting/status.md b/docs/Data-insights/Features/New-alerting/status.md index c45868d..ee2f6d6 100644 --- a/docs/Data-insights/Features/New-alerting/status.md +++ b/docs/Data-insights/Features/New-alerting/status.md @@ -12,13 +12,20 @@ A summary band at the top leads with what matters: how many rules **need attenti Below it, a health bar shows the split across states, with a chip and count for each. Only states that have rules in them appear - an account with nothing firing or pending shows just **Normal** and **Paused**. +The band always covers every rule on the account. Filtering the page to a single namespace narrows the cards below, but the band still counts the lot - so a page showing one card can still read *1 of 70 rules*. + | State | Meaning | |---|---| | **Firing** | The alert condition is met and the rule is actively firing | +| **Error** | The rule's query failed to evaluate | | **Pending** | The condition has been met, but not yet for long enough to fire | +| **Recovering** | The condition is no longer met, but the rule is still held by its **Keep firing for** period before it returns to Normal. See [Alert is flapping](troubleshooting.md#alert-is-flapping) | | **Normal** | The rule is evaluating and its condition is not currently met | +| **No Data** | The query returned no data, so there was nothing to evaluate | | **Paused** | The rule is paused and not being evaluated | +These are the same seven states offered by the state filter in the toolbar. + ## Viewing and filtering Controls across the top shape how the page is laid out: @@ -27,8 +34,8 @@ Controls across the top shape how the page is laid out: |---|---| | **List / Grid** | Switch between the list view and a compact grid view | | **Source / Namespace** | Group the cards by data source or by namespace | -| **All sources** | Filter to a specific source | -| **State filter** | A multi-select (for example, **4 selected**) to show only rules in the chosen states. Click **✕** to clear it | +| **All sources** / **All namespaces** | Narrow the page to particular sources or namespaces - the label follows the grouping toggle. Open it for a searchable list, with **Select all** to take everything and **Clear all** to start again. Selected entries are ticked in the list, and the button itself becomes a chip naming your selection, with an **✕** to remove it | +| **State filter** | A multi-select to show only rules in the chosen states, with a search box and **Select all**. It reads **All states** when nothing is excluded, and **N selected** once you narrow it. Click **✕** to clear it | | **Expand all** | Expand every group to show all of its rules; it toggles to **Shrink all** to collapse them | | **Hide filtered-out cards** | Hide the groups and rules that don't match the current filters | | **Refresh interval** | How often the page auto-refreshes (for example, **30s**) | @@ -47,9 +54,11 @@ Each rule shows: - Its current **state** and how long it has been in it - for example, *Firing for 9m* - An **Active** toggle to enable or disable the rule - it reads **Paused** when the rule is off - A **mute** icon to silence its notifications -- An **eye** icon (**View rule**) to open the rule's [detail view](rules.md#rule-detail-view) - clicking the rule itself opens the same view +- An **eye** icon (**View rule**) to open the rule's [detail view](rules.md#rule-detail-view) + +Clicking the rule itself opens a [panel](rules.md#the-rule-panel) beside the list, with its metric graph, state history, and buttons to silence, edit, or open the full rule. -When a group has more rules than fit, click **Show all N rules** to expand it, and **Show less** to collapse it again. +When a group holds more rules than its card shows, click **Show all N rules** to reveal the rest. Once they are all visible the card's footer reads **All N rules shown**. Shift-click rules to select more than one at a time, then silence the whole selection in one go. See [Silences](silences.md). @@ -57,13 +66,30 @@ The **List** view stacks the groups and shows their rule cards inline, while the ## Creating from Status -The **+ Wizard** button in the top right lets you create alerting resources without leaving the Status page. Click it to start a new alert rule, or use its dropdown for more options: +The **+ Wizard** button in the top right lets you create alerting resources without leaving the Status page. There are two ways into it. + +Click **+ Wizard** itself to open the **What would you like to create?** dialog, which describes each starting point: + +| Option | Description | +|---|---| +| **Alert Rule** | Monitor a metric against a threshold you define. See [the guided wizard](rules.md#the-guided-wizard) | +| **Anomaly Detector** | Automatically detect unusual metric behavior | +| **Contact Point** | Set up where notifications get sent. See [Contact Points](contact-points.md) | + +If you are not sure which you need, click **Ask OpsPilot** at the bottom of the dialog for a recommendation. Click the **✕** in the top right to close the dialog without creating anything. + +### Creating a detector from the wizard + +Choosing **Anomaly Detector** asks a second question - **What kind of detector?** - because there are two ways to get one: | Option | Description | |---|---| -| **New rule** | Create a new alert rule. See [Rules](rules.md) | -| **New contact point** | Add a new contact point. See [Contact Points](contact-points.md) | -| **New custom detector** | Create a custom anomaly detector. See [Custom Anomaly Detectors](custom-anomaly-detectors.md) | +| **Custom Detector** | Pick a metric and set sensitivity yourself. See [Custom Anomaly Detectors](custom-anomaly-detectors.md) | +| **Service Scan** | Auto-detect services from your FusionReactor agents and create rate, error, and duration detectors for each. See [Service Anomaly Detectors](service-anomaly-detectors.md) | + +**Service Scan** is the same action as **Scan for services** on the [Service Anomaly Detectors](service-anomaly-detectors.md) page - it creates the R/E/D detectors for every service it finds, rather than one detector at a time. + +The **⌄** dropdown beside the button skips the dialog and goes straight to the same three: **New rule**, **New contact point**, and **New custom detector**. !!! question "Need more help?" Contact support in the chat bubble and let us know how we can assist.