Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions docs/Data-insights/Features/New-alerting/contact-points.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ Use **Search contact points** to find one by name. Use the **type filter** dropd
| **Telegram** | Send notifications via Telegram |
| **Google Chat** | Post notifications to Google Chat |
| **PagerDuty** | Create incidents in PagerDuty |
| **OpsGenie** | Create alerts in OpsGenie |
| **Jira Service Management** | Raise issues in Jira Service Management |
| **Pushover** | Send push notifications via Pushover |
| **Webhook** | POST a JSON payload to any URL |
| **Email** | Send notifications by email |
Expand All @@ -43,13 +43,15 @@ Use **Search contact points** to find one by name. Use the **type filter** dropd
!!! tip "You can also add one while creating a rule"
Contact points are not only made here. In the rule editor, **Then notify** lets you search your existing contact points, or click **Create new contact point** to make one without leaving the rule. See [Rules](rules.md).

A third route is the **Wizard** on the [Status](status.md) page. Choosing **Contact Point** asks **How should notifications be delivered?** and offers the notification channels as a grid of cards. Picking one takes you into its configuration. As elsewhere in the wizard, **Skip to form** leaves for the full form, **Ask OpsPilot** suggests a channel, **←** goes back and **✕** closes without creating anything.

1. Click **+ New contact point** to open the integration picker
2. Select an integration type from the grid. Use the category tabs to filter:
2. Select an integration type from the grid. Use the category tabs to filter - each shows how many types it holds:

| Category | Integrations |
|---|---|
| **Chat** | Slack, Discord, Microsoft Teams, Telegram, Google Chat |
| **On-call** | PagerDuty, OpsGenie, Pushover |
| **On-call** | PagerDuty, Jira Service Management, Pushover |
| **Webhook** | Webhook |
| **Other** | Email, Kafka REST Proxy |

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,26 @@ Each detector row shows its **State**, **Name**, **Threshold**, and **Last evalu

Click **+ New custom detector** (top right of the page), or use the **Wizard** (**New custom detector**) on the [Status](status.md) or [Rules](rules.md) page. Fill in the fields below, then click **Create detector** to save.

### The guided wizard

Starting from the **Wizard** asks one question at a time instead of presenting the whole form. After choosing **Anomaly Detector** and then **Custom Detector** (see [Creating a detector from the wizard](status.md#creating-a-detector-from-the-wizard)), it asks which metric to watch, and then how sensitive detection should be.

**How sensitive should detection be?** offers three presets. Higher sensitivity catches more anomalies but may alert more often:

| Preset | Anomaly threshold | Description |
|---|---|---|
| **Low** | 95% | Fewer alerts, only strong anomalies |
| **Medium** | 90% | Balanced sensitivity |
| **High** | 80% | More alerts, catches subtle changes |

The percentage is the **anomaly threshold** described under [When to fire](#when-to-fire) - the score at or above which the detector fires. The relationship is inverted, so a *lower* sensitivity sets a *higher* threshold: at **Low**, the model has to be 95% confident before anything fires.

Whichever preset you pick, you can change the threshold afterwards on the detector itself, so this is a starting point rather than a commitment.

**Who gets notified?** is the last step. Pick one or more [contact points](contact-points.md) with **+ Add contact point**, or click **Skip** to create the detector without notifications and add them later. Click **Done** to finish.

As in the [alert rule wizard](rules.md#the-guided-wizard), each step offers **←** to go back, **Skip to form** to leave the wizard for the full form, **Ask OpsPilot** for a recommendation, and **✕** to close without creating anything. **Skip to form** is offered on every step but the last, where **Skip** and **Done** take its place.

### Signal

Defines the PromQL series the detector watches. The query is validated against the Prometheus datasource on save.
Expand Down
86 changes: 84 additions & 2 deletions docs/Data-insights/Features/New-alerting/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,11 +69,30 @@ The **Dashboard** and **Runbook** buttons (top right of the expanded view) open

![Screenshot](/Data-insights/Features/images/Alerting/rule-expanded.png)

### The rule panel

Clicking a rule on the [Status](status.md) page opens a panel beside the list - a quick look at that rule without leaving the page. Click the **✕** to close it.

The header shows the rule name, its state and how long it has held it, and the namespace and data source it belongs to (such as, *FusionReactor Alerts / Metrics*), with four actions:

| Action | Description |
|---|---|
| **Silence** | Create a [silence](silences.md) for this rule |
| **View rule** | Open the full [rule detail view](#rule-detail-view) |
| **Edit rule** | Open the rule editor |
| **Active** | A toggle to pause and resume evaluation |

Below the header sit two summary cards, **Duration** and **State**, then:

**Metric** - a graph of the query with the threshold drawn on it. Use the time range picker, its step arrows and the zoom buttons to adjust the window, or click **Open in Explore** to investigate the metric in Explore. Three checkboxes below the graph toggle the **Threshold**, **State transitions** and **Pending window** overlays.

**State history** - the rule's state changes, newest first. The header gives the period and transition count (such as, *last 24h · 20 transitions*), and **See all →** opens the full history. Each row reads as the new state *from* the previous one - for example, **Pending** from **Normal** - with when it happened, how long the previous state was held, and how long ago that was.

### Rule detail view

![Screenshot](/Data-insights/Features/images/Alerting/high-cpu-rule.png)

Open the full rule detail view by clicking the **eye** icon (**View rule**) in the rules list, or by clicking a rule on the [Status](status.md) page. The header shows the rule name and current state, with these actions in the top right:
Open the full rule detail view by clicking the **eye** icon (**View rule**) in the rules list, or **View rule** in the [rule panel](#the-rule-panel). The header shows the rule name and current state, with these actions in the top right:

| Action | Description |
|---|---|
Expand Down Expand Up @@ -138,6 +157,60 @@ A good alert rule has three things: a query that targets the right signal, a thr

Click **+ New rule** (top right) to open the rule editor. (To create an anomaly detector instead, see [Service](service-anomaly-detectors.md) or [Custom Anomaly Detectors](custom-anomaly-detectors.md), or use the **Wizard** on the [Status](status.md) page.)

### The guided wizard

Starting a rule from the **Wizard** on the [Status](status.md) page walks you through the decisions one at a time, rather than presenting the whole form at once. A row of dots below the heading tracks your progress through the steps.

You are never locked into the wizard. Each step offers:

| Control | What it does |
|---|---|
| **←** | Go back to the previous step |
| **Skip to form** | Leave the wizard and go straight to the main rule form |
| **Not sure? Ask OpsPilot** | Get a recommendation for the step you're on |
| **✕** | Close the wizard without creating anything |

**Skip to form** is offered on every step but the last, where **Skip** and **Done** take its place.

The steps are:

**What are you monitoring?** - pick the data source the alert will query. The list holds every data source configured on your account, including those added by [integrations](../integrations.md), so an AWS installation appears here alongside your metrics and logs sources.

**How do you want to build the query?** - choose how to express the condition:

| Option | Description |
|---|---|
| **Guided builder** | Build your query step by step |
| **Write PromQL** | Write the query expression directly |

This choice is not binding - you can switch between the two later in the form.

**What should we watch?** - pick the metric, or write the query, that the alert will evaluate. What this step shows depends on the choice you made at the previous one:

- **Guided builder** gives you a **Metric** dropdown. Once you choose a metric, a preview graph appears below it showing that metric's recent behavior, so you can confirm you have the right signal before going further.
- **Write PromQL** gives you a **PromQL expression** box to type the query into directly.

Either way, click **Next** to continue. The wizard is the same length whichever you pick.

**When should it fire?** - set the threshold that triggers notifications:

| Field | Description |
|---|---|
| **Alert when value is** | **Above** or **Below** the threshold |
| **Threshold** | The value to compare against (such as, `80`) |
| **Wait before alerting** | How long the condition must hold before the alert fires - **1m**, **5m**, or **10m** |

**Wait before alerting** is the pending period under a plainer name. Leaving it at anything above **1m** is what stops a brief spike from paging someone.

**Who gets notified?** - pick one or more [contact points](contact-points.md) with **+ Add contact point**. This step is optional: click **Skip** to create the rule without notifications and add them later, or **Done** to finish.

A rule with no contact point still evaluates and still shows its state on [Status](status.md) - it just won't notify anyone. Its expanded view reads *None configured* under **Notifies**.

!!! note
The earlier steps advance as soon as you pick an option. From **What should we watch?** onward, you make a choice and then click **Next**.

### Rule editor modes

The rule editor has two modes, toggled in the top right:

| Mode | Description |
Expand Down Expand Up @@ -180,17 +253,26 @@ Build the alert from a chain of **Queries & expressions**. Click **Add query** (

A new rule starts with a default **Query → Reduce → Threshold** chain. The **Alert condition** dropdown at the top selects which step's firing state determines whether the rule alerts.

Each step is labeled with a chip naming its reference ID and type, colored by category, and steps that take input from another show which one they follow (such as, *← $reduce*). The step serving as the alert condition is outlined and carries an **Alert condition** badge, so you can see at a glance which one decides the outcome.

Reorder steps with the arrows to the left of each one, and remove a step with the **✕** on its right.

Below the chain, **Evaluation preview** runs the pipeline through the alert condition step and shows the result as it would be evaluated. If a step cannot run, the preview reports the failure and names the step responsible - an empty or malformed query, for example, reads *Couldn't evaluate this rule*. Use it to catch mistakes before saving rather than after the rule goes live.

#### Evaluation

The **Evaluation** section controls how the rule runs:

| Setting | Description |
|---|---|
| **Evaluate every** | How often the rule is checked (such as, `1m`) |
| **Group** | The evaluation group the rule belongs to (such as, `default`) |
| **Every** | How often the rule is checked (such as, `1m`). The schedule belongs to the group rather than to the individual rule |
| **Pending for** | How long the condition must be continuously met before the alert fires (such as, `5m`). Prevents notifications for temporary spikes |
| **No data** | The state the rule enters when the query returns no data - **No Data**, **Alerting**, **Normal**, or **Keep last state** |
| **On error** | The state the rule enters when the query fails - **Error**, **Alerting**, **Normal**, or **Keep last state** |

Tick **Choose the group and schedule myself** to set the group and its interval by hand. Rules in a group all evaluate together on one schedule, so changing it changes every rule in that group - not only the one you are editing.

#### Namespace

Expand **Namespace** and choose the **namespace** where the rule is stored. Namespaces keep rules organized and control access.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ Each instrumented service automatically gets three detectors, shown as **R**, **

- **Rate** (R) - request rate anomalies
- **Error** (E) - error rate anomalies
- **Latency** (D) - response time anomalies (also shown as *Duration* on the Status page)
- **Duration** (D) - response time anomalies (also referred to as *latency*)

The badges are filled while the service's detectors are running, and grayed out while they are paused. **Last evaluation** reads a timestamp (such as, *29 Sept, 13:01*) for a running service, and **—** for a paused one.

Expand All @@ -24,6 +24,10 @@ Each row in the service list shows the service name, its **Detectors** (R/E/D),

Click **Scan for services** to detect your instrumented services and auto-create detectors for them. If no service detectors exist yet, this is the first step.

You can also reach it from the **Wizard** on the [Status](status.md) page, by choosing **Anomaly Detector** and then **Service Scan**. That route ends on a confirmation step explaining what the scan will do, with **Scan now** to run it and **Back** to return to the previous step.

Either way, the scan creates rate, error, and duration detectors for every instrumented service it finds. Sensitivity and notifications are then tuned per detector afterwards, in [Detector settings](#detector-settings).

## State counters

| State | Description |
Expand Down
42 changes: 34 additions & 8 deletions docs/Data-insights/Features/New-alerting/status.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,13 +12,20 @@ A summary band at the top leads with what matters: how many rules **need attenti

Below it, a health bar shows the split across states, with a chip and count for each. Only states that have rules in them appear - an account with nothing firing or pending shows just **Normal** and **Paused**.

The band always covers every rule on the account. Filtering the page to a single namespace narrows the cards below, but the band still counts the lot - so a page showing one card can still read *1 of 70 rules*.

| State | Meaning |
|---|---|
| **Firing** | The alert condition is met and the rule is actively firing |
| **Error** | The rule's query failed to evaluate |
| **Pending** | The condition has been met, but not yet for long enough to fire |
| **Recovering** | The condition is no longer met, but the rule is still held by its **Keep firing for** period before it returns to Normal. See [Alert is flapping](troubleshooting.md#alert-is-flapping) |
| **Normal** | The rule is evaluating and its condition is not currently met |
| **No Data** | The query returned no data, so there was nothing to evaluate |
| **Paused** | The rule is paused and not being evaluated |

These are the same seven states offered by the state filter in the toolbar.

## Viewing and filtering

Controls across the top shape how the page is laid out:
Expand All @@ -27,8 +34,8 @@ Controls across the top shape how the page is laid out:
|---|---|
| **List / Grid** | Switch between the list view and a compact grid view |
| **Source / Namespace** | Group the cards by data source or by namespace |
| **All sources** | Filter to a specific source |
| **State filter** | A multi-select (for example, **4 selected**) to show only rules in the chosen states. Click **✕** to clear it |
| **All sources** / **All namespaces** | Narrow the page to particular sources or namespaces - the label follows the grouping toggle. Open it for a searchable list, with **Select all** to take everything and **Clear all** to start again. Selected entries are ticked in the list, and the button itself becomes a chip naming your selection, with an **✕** to remove it |
| **State filter** | A multi-select to show only rules in the chosen states, with a search box and **Select all**. It reads **All states** when nothing is excluded, and **N selected** once you narrow it. Click **✕** to clear it |
| **Expand all** | Expand every group to show all of its rules; it toggles to **Shrink all** to collapse them |
| **Hide filtered-out cards** | Hide the groups and rules that don't match the current filters |
| **Refresh interval** | How often the page auto-refreshes (for example, **30s**) |
Expand All @@ -47,23 +54,42 @@ Each rule shows:
- Its current **state** and how long it has been in it - for example, *Firing for 9m*
- An **Active** toggle to enable or disable the rule - it reads **Paused** when the rule is off
- A **mute** icon to silence its notifications
- An **eye** icon (**View rule**) to open the rule's [detail view](rules.md#rule-detail-view) - clicking the rule itself opens the same view
- An **eye** icon (**View rule**) to open the rule's [detail view](rules.md#rule-detail-view)

Clicking the rule itself opens a [panel](rules.md#the-rule-panel) beside the list, with its metric graph, state history, and buttons to silence, edit, or open the full rule.

When a group has more rules than fit, click **Show all N rules** to expand it, and **Show less** to collapse it again.
When a group holds more rules than its card shows, click **Show all N rules** to reveal the rest. Once they are all visible the card's footer reads **All N rules shown**.

Shift-click rules to select more than one at a time, then silence the whole selection in one go. See [Silences](silences.md).

The **List** view stacks the groups and shows their rule cards inline, while the **Grid** view lays the groups out as compact cards in a multi-column grid - each summarizing its state at a glance, and expandable to reveal the rules inside.

## Creating from Status

The **+ Wizard** button in the top right lets you create alerting resources without leaving the Status page. Click it to start a new alert rule, or use its dropdown for more options:
The **+ Wizard** button in the top right lets you create alerting resources without leaving the Status page. There are two ways into it.

Click **+ Wizard** itself to open the **What would you like to create?** dialog, which describes each starting point:

| Option | Description |
|---|---|
| **Alert Rule** | Monitor a metric against a threshold you define. See [the guided wizard](rules.md#the-guided-wizard) |
| **Anomaly Detector** | Automatically detect unusual metric behavior |
| **Contact Point** | Set up where notifications get sent. See [Contact Points](contact-points.md) |

If you are not sure which you need, click **Ask OpsPilot** at the bottom of the dialog for a recommendation. Click the **✕** in the top right to close the dialog without creating anything.

### Creating a detector from the wizard

Choosing **Anomaly Detector** asks a second question - **What kind of detector?** - because there are two ways to get one:

| Option | Description |
|---|---|
| **New rule** | Create a new alert rule. See [Rules](rules.md) |
| **New contact point** | Add a new contact point. See [Contact Points](contact-points.md) |
| **New custom detector** | Create a custom anomaly detector. See [Custom Anomaly Detectors](custom-anomaly-detectors.md) |
| **Custom Detector** | Pick a metric and set sensitivity yourself. See [Custom Anomaly Detectors](custom-anomaly-detectors.md) |
| **Service Scan** | Auto-detect services from your FusionReactor agents and create rate, error, and duration detectors for each. See [Service Anomaly Detectors](service-anomaly-detectors.md) |

**Service Scan** is the same action as **Scan for services** on the [Service Anomaly Detectors](service-anomaly-detectors.md) page - it creates the R/E/D detectors for every service it finds, rather than one detector at a time.

The **⌄** dropdown beside the button skips the dialog and goes straight to the same three: **New rule**, **New contact point**, and **New custom detector**.

!!! question "Need more help?"
Contact support in the chat bubble and let us know how we can assist.
Loading