diff --git a/docs/Data-insights/Features/Integrations/Chat/opspilot-mcp.md b/docs/Data-insights/Features/Integrations/Chat/opspilot-mcp.md new file mode 100644 index 0000000..444ae49 --- /dev/null +++ b/docs/Data-insights/Features/Integrations/Chat/opspilot-mcp.md @@ -0,0 +1,93 @@ +# OpsPilot MCP + +Let AI assistants query OpsPilot over the Model Context Protocol (MCP). + +OpsPilot MCP brings OpsPilot to where your team already works. An AI assistant there can work with your telemetry and situations the way the OpsPilot coworker does, so you can ask without switching tools. + +Navigate to **Integrations** from the left-hand sidebar, then select **OpsPilot MCP**. + +--- + +## What assistants can do + +Once connected, an assistant can read your **dashboards**, **metrics**, **logs**, **traces**, and the **situations** your coworker is tracking - and, depending on the permission tier, act on them. + +Each person connects their own assistant and signs in as themselves, so there is no key to paste and everyone sees what they can already see in OpsPilot. + +OpsPilot MCP declares no **capabilities**, so it provisions no dashboards or alerts of its own. It reads and acts on what is already in your account. + +--- + +## Permission tiers + +| Tier | Description | +|---|---| +| **Read + Act** | The default. Assistants may read and act on what they find. Requires nothing - your plan grants it | +| **Read-only** | Assistants may only read. Requires nothing, and holds the whole organisation to reading | + +The tier set on the install is the **ceiling for the whole organisation**, whichever address a person connects with. Read-only is enforced by OpsPilot rather than trusted to the client, so an assistant on a read-only connection is never offered the tools that change anything. + +--- + +## Installing + +Click **Install** on the **OpsPilot MCP** card in the [integration catalog](../../integrations.md), or from its detail view. The dialog confirms which account the install will serve. Click **Install** to confirm. + +Installing makes OpsPilot available. Each person still points their own assistant at it. + +--- + +## Connect your assistant + +Open **Connect** on the integration. Choose the address you want, then send it to your client or copy it. + +### Which address + +| Address | What it gives | +|---|---| +| **Full access** | The default. Read and act | +| **Read-only** | Never offers the tools that change anything | + +The choice applies to the **connection**, not the account - so somebody who also adds the full address gets the full surface. To hold everyone to reading whatever they connect with, set the install's permission tier to **Read-only**. Where both apply, the stricter wins. + +### Sending it to your client + +**Connect** offers a button per client: + +| Client | How | +|---|---| +| **VS Code** and **VS Code Insiders** | Click the button to hand the address straight to the editor | +| **Claude Code** | Copy the `claude mcp add` command and run it | +| **Claude desktop and web** | Click **Open**, or copy the address and add it under **Customize → Connectors → Add custom connector** | + +!!! note "The address lives on the Connect page" + OpsPilot runs in more than one environment, so a written-down address names only one of them. Always take the address from **Connect** rather than copying one from documentation. + +--- + +## Setting it up for everyone + +If you are an **Owner** of a Claude Team or Enterprise organisation, add it once for everybody: **Organization settings** → **Connectors** → **Add**, and paste the address from **Connect**. + +That covers desktop, web, mobile and Claude Code together. It makes OpsPilot available; each person still enables it and signs in as themselves. + +--- + +## Check it works + +Ask your assistant: + +> *what situations are active right now?* + +A list means you are connected. + +--- + +## Troubleshooting + +**A VS Code button does nothing.** Those buttons hand the address to the editor through a link the operating system routes, so nothing happens when that editor is not installed on the machine you are reading this on. Copy the address instead and add it from inside the client. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/Chat/slack.md b/docs/Data-insights/Features/Integrations/Chat/slack.md index 1f3ddb9..8ad3970 100644 --- a/docs/Data-insights/Features/Integrations/Chat/slack.md +++ b/docs/Data-insights/Features/Integrations/Chat/slack.md @@ -1,25 +1,110 @@ # Slack Integration -Connect OpsPilot to Slack to receive alert notifications and incident updates directly in your Slack channels. +Talk to OpsPilot from Slack - mention it in a channel or DM it directly. + +The Slack integration brings OpsPilot to where your team already talks. Invite the bot to a channel and switch posting on, and the alerts, situations, and digests the coworker surfaces land there as they happen. Mention it there or message it directly and you get the same coworker you use in the app, so you can ask about what it just posted without switching tools. + +DMs are personal: each person links their own account. Navigate to **Integrations** from the left-hand sidebar, then select **Slack**. --- -## Setup +## Permission tiers + +| Tier | Description | +|---|---| +| **Read-only** | The default. Requires nothing to supply - it is granted with your plan. OpsPilot answers questions in Slack from what the organisation can already see | +| **Read + Act** | Requires Read, plus a plan that includes assistant actions. Lets Slack users act on the organisation's behalf - acknowledging alerts, silencing, and triggering runbooks - subject to their own account link | + +--- + +## Connecting your workspace 1. In OpsPilot, go to **Integrations** and click **Slack**. -2. Click **Connect** and follow the OAuth flow to authorise OpsPilot in your Slack workspace. -3. Select the default channel you want notifications sent to. -4. Click **Save**. +2. Click **Add to Slack**. A Slack window opens asking you to sign in to your workspace. +3. Enter your workspace's Slack URL (for example, `your-workspace.slack.com`) and click **Continue**. If you don't know it, use **Find your workspaces**. +4. Sign in to the workspace. The method depends on how your workspace is configured - many require a Google account or another single sign-on provider on your organisation's domain. +5. On the **Allow the "OpsPilot" app to access Slack** screen, choose the **Workspace** to install into, review the permissions OpsPilot is asking for, and click **Allow**. + +!!! note "Permissions OpsPilot requests" + Slack asks you to approve what OpsPilot can do in your workspace: + + | | Permission | + |---|---| + | **View** | Content and info about channels and conversations | + | **View** | Content and info about your workspace | + | **Act** | Perform actions in channels and conversations | + + Expand **More permissions** to see the full list, or click **Manage permissions** to change them. Slack shares the permissions you grant with OpsPilot. + +--- + +## Inviting OpsPilot to a channel + +Connecting the workspace posts nothing anywhere. OpsPilot waits to be invited, so you choose which channels it appears in. + +In the channel you want situations in, run: + +``` +/invite @OpsPilot +``` + +OpsPilot introduces itself and waits. Situation posting is off until someone turns it on - tap **Enable situation posting** on the welcome card, or run: + +``` +/opspilot unmute +``` + +Situations at **warning** severity or above then land in that channel. + +--- + +## Connecting your own account + +Direct messages are personal, so each person links their OpsPilot account once before DMing the bot. Message it and it hands you the link. + +Channel mentions need no linking - only DMs. + +--- + +## What OpsPilot reads + +OpsPilot only reads messages that mention it. Background reading is off in every channel until someone turns it on: + +``` +/opspilot listen passive +``` + +`/opspilot status` shows where a channel stands, and `/opspilot listen off` stops it again. + +--- + +## Checking it works + +In a channel OpsPilot is in, ask *"what situations are active right now?"*. An answer means the workspace is connected and the bot can reach your data. + +--- + +## Slash commands + +| Command | Description | +|---|---| +| `/invite @OpsPilot` | Add OpsPilot to the channel | +| `/opspilot unmute` | Turn situation posting on for the channel | +| `/opspilot mute` | Stop the posts again, without removing the bot | +| `/opspilot route` | Narrow what is posted by service, severity, or category | +| `/opspilot listen passive` | Turn on background reading of the channel | +| `/opspilot listen off` | Turn background reading off again | +| `/opspilot status` | Show where the channel stands | --- -## Using Slack with Alerts +## Using Slack with alerts -Once connected, you can select Slack as a contact point when configuring alert notification policies. +You can also select Slack as a contact point when configuring alert notification policies. -Navigate to **Alerts > Contact Points** and choose Slack as the delivery method. You can specify a channel per contact point to route different alerts to different channels. +Navigate to **Alerting > Notifications**, open the **Contact Points** tab, and choose Slack as the integration type. See [Contact Points](../../New-alerting/contact-points.md) for the full setup. --- diff --git a/docs/Data-insights/Features/Integrations/Cloud/aws.md b/docs/Data-insights/Features/Integrations/Cloud/aws.md new file mode 100644 index 0000000..25e4bee --- /dev/null +++ b/docs/Data-insights/Features/Integrations/Cloud/aws.md @@ -0,0 +1,65 @@ +# AWS + +Connect AWS for EC2, RDS, and other CloudWatch-held metrics. + +The AWS integration puts your AWS account's own monitoring data next to everything else you watch. CloudWatch metrics and logs arrive with dashboards for the services you are already paying for, so you can see how an account is behaving without building the views yourself or leaving to go and look. + +Navigate to **Integrations** from the left-hand sidebar, then select **AWS**. + +--- + +## Permission tiers + +Unlike most integrations, AWS needs credentials you supply - an IAM role or key with the right CloudWatch permissions. + +| Tier | Requires | +|---|---| +| **Read-only** | The default. An AWS IAM role or key with `CloudWatch:Get*`, `List*`, and `Describe*` permissions | +| **Read + Write** | The Read permissions, plus `CloudWatch:PutMetricAlarm` and related write actions | + +--- + +## What it provisions + +AWS declares four capabilities - **Data Sources**, **Dashboards**, **Recording Rules**, and **Alerts**. The dashboards, alert rules, and recording rules are provisioned as soon as the install succeeds. + +The current version installs **16 dashboards**, covering the services CloudWatch reports on. + +--- + +## Working across regions + +The dashboards follow the region configured on the installation. If you connect several AWS regions, install the AWS integration **once per region** - each instance gets its own data source, its own dashboards, and its own folder. + +Two dashboards are the exception, because AWS only reports their metrics to `us-east-1`. Both pin that region themselves, so they work whatever region an installation uses. + +--- + +## Two dashboards need extra setup + +Two of the sixteen dashboards read metrics that AWS does not publish by default. + +### AWS Billing + +The **AWS Billing** dashboard reads `EstimatedCharges` from the `AWS/Billing` namespace. AWS publishes that metric only if billing alerts are switched on, and only into `us-east-1`. The dashboard queries that region regardless of the region you configured, so no change to the installation is needed. + +To switch it on, in the AWS console: + +1. Sign in to the **management account** of your organisation. The metric is published for the payer account, not for member accounts. +2. Open **Billing and Cost Management** → **Billing preferences**. +3. Enable **Receive CloudWatch billing alerts**, and save. + +!!! note + The metric starts being published from that point on and is not backfilled. The dashboard stays empty for the first few hours, and shows no history from before you enabled it. + +### AWS CloudFront + +CloudFront reports its metrics only to `us-east-1` as well. The dashboard pins that region itself, so it works whatever region the installation uses - but the credentials you supplied must be allowed to read CloudWatch in `us-east-1`. + +!!! warning + An IAM policy scoped to a single region will refuse this. If your CloudFront dashboard is empty, check that the credentials can read CloudWatch in `us-east-1`. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/Data/opentrace.md b/docs/Data-insights/Features/Integrations/Data/opentrace.md new file mode 100644 index 0000000..6469ba0 --- /dev/null +++ b/docs/Data-insights/Features/Integrations/Data/opentrace.md @@ -0,0 +1,46 @@ +# OpenTrace + +Connect OpenTrace so OpsPilot can read your code alongside your telemetry. + +OpsPilot knows what your systems are doing. OpenTrace knows what your code is: what calls what, who owns a file, and what changed in which pull request. Connect your OpenTrace account and OpsPilot can use both when it answers you, so a question about a service that started erroring can reach the change that touched it. + +Navigate to **Integrations** from the left-hand sidebar, then select **OpenTrace**. + +--- + +## What it adds + +With OpenTrace connected, OpsPilot can draw on your code as well as your telemetry when it answers you: + +- **What calls what** - the relationships between your services and components +- **Who owns a file** - who to ask, or who to tell +- **What changed in which pull request** - the change history behind the code + +That means a question about a service that has started erroring can reach the change that touched it, rather than stopping at the symptom. + +--- + +## Your own account + +You connect your own OpenTrace account, and OpsPilot uses it only in **your** conversations. Each person connects their own account - it isn't shared across the organisation. + +Connect yours from **Integrations** → **User MCPs**, where OpenTrace appears under **Data** with a **Connect** button. The agent uses that connection only in chat, never in scheduled work and never on anyone else's behalf. + +--- + +## Permission tiers + +OpenTrace runs **Read-only**, which is the default and the only tier. It requires nothing to enable. + +OpenTrace declares no **capabilities**, so it doesn't provision dashboards or alerts of its own - it adds context to the answers OpsPilot gives you. + +--- + +## Installing + +Click **Install** on the **OpenTrace** card in the [integration catalog](../../integrations.md), or from the integration's detail view. The install dialog confirms which account the install will serve and shows the permission tier. Click **Install** to confirm. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/Infrastructure/docker.md b/docs/Data-insights/Features/Integrations/Infrastructure/docker.md new file mode 100644 index 0000000..f5b31e4 --- /dev/null +++ b/docs/Data-insights/Features/Integrations/Infrastructure/docker.md @@ -0,0 +1,114 @@ +# Docker + +Monitor Docker containers - CPU, memory against limit, network and disk IO. + +Your containers show up alongside the rest of your telemetry on a dashboard that is already built, so there are no panels to design. Memory is measured against the limit you set rather than as a raw number, which is what tells you a container is heading for a restart before it gets one. + +!!! info "What this integration does" + It provisions a **dashboard**. It does not collect anything itself and never talks to your Docker daemon - the metrics come from **Grafana Alloy**, which has a cAdvisor collector built in. Until Alloy is collecting them, the dashboard is empty. + +Navigate to **Integrations** from the left-hand sidebar, then select **Docker**. + +--- + +## Permission tiers + +Docker runs **Read-only**, which is the default and requires nothing - the metrics endpoint and signing key come from the service environment. + +--- + +## 1. Configure Alloy + +If you have no collector yet, the [Grafana Alloy guide](/Monitor-your-data/OpenTelemetry/Shipping/Collector/) covers installing one and getting an API key. That guide configures Alloy for OTLP; container metrics are Prometheus-format, so the components below are different, but the install and the API key are the same. + +Add this to your `config.alloy`: + +```river +prometheus.exporter.cadvisor "docker" { + docker_host = "unix:///var/run/docker.sock" + docker_only = true + storage_duration = "5m" +} + +prometheus.scrape "docker" { + targets = prometheus.exporter.cadvisor.docker.targets + + job_name = "docker" + scrape_interval = "15s" + + forward_to = [prometheus.remote_write.opspilot.receiver] +} + +prometheus.remote_write "opspilot" { + endpoint { + url = "https://api.fusionreactor.io/v1/metrics" + + headers = { + "authorization" = sys.env("OPSPILOT_API_KEY"), + } + } + + external_labels = { + host = "", + } +} +``` + +Three values matter more than they look: + +| Value | Why it matters | +|---|---| +| `docker_only = true` | Keeps the collector to containers. Without it, it reports every cgroup on the host - systemd units, user sessions, desktop services - which on a typical machine is fifty times the series for no benefit, and none of them carry a container name for the dashboard to group by | +| `job_name = "docker"` | Required. Every panel filters on it, because your metrics land in a store shared with everything else you send to OpsPilot. Alloy's default job label is the component's own id, which would change if you renamed the component, so it is set explicitly. Change it and the dashboard goes blank | +| `host` | Yours to choose, and it populates the dashboard's host selector. Give each Docker host a distinct name if you run more than one - otherwise they produce indistinguishable series and the panels cannot tell them apart | + +--- + +## 2. Run Alloy with access to Docker + +The collector reads container metadata from the Docker socket and container resource usage from the host's cgroups, so it needs both. If Alloy runs as a container itself, it needs these mounts: + +```yaml +services: + alloy: + image: grafana/alloy:latest + privileged: true + volumes: + - /:/rootfs:ro + - /var/run:/var/run:ro + - /sys:/sys:ro + - /var/lib/docker:/var/lib/docker:ro + - ./config.alloy:/etc/alloy/config.alloy +``` + +!!! warning "Linux hosts only" + Docker Desktop for macOS and Windows runs containers inside a Linux VM, which stops the collector resolving container metadata: you get cgroup ids with no container names, and the dashboard has nothing readable to group by. + +--- + +## 3. Confirm the data arrived + +In OpsPilot, open **Explore**, select the **Metrics** data source, and run: + +```promql +container_cpu_usage_seconds_total{job="docker"} +``` + +You should get one series per container, each carrying a `name` and an `image`. If that returns data, the dashboard will too. + +--- + +## Troubleshooting + +**Nothing in Explore.** Check the Alloy UI on port 12345 first - a failing component shows there with the error. If the scrape is healthy, check the `job` label really is `docker` by querying `container_last_seen` with no selector and reading the labels back. + +**Series with an `id` but no `name`.** The collector can see the cgroups but not the Docker daemon. Check the socket mount, and check you are on a Linux host rather than Docker Desktop. + +**Memory against limit is empty.** No container has a memory limit set. Docker reports an unlimited container's limit as zero, and the panel excludes those rather than dividing by them. Set a limit on a container with `--memory`, or `mem_limit` in Compose, and it will appear. + +**Disk IO is empty.** Container filesystem metrics depend on the storage driver. `overlay2` reports them; some others do not. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/Infrastructure/proxmox-ve.md b/docs/Data-insights/Features/Integrations/Infrastructure/proxmox-ve.md new file mode 100644 index 0000000..90642eb --- /dev/null +++ b/docs/Data-insights/Features/Integrations/Infrastructure/proxmox-ve.md @@ -0,0 +1,149 @@ +# Proxmox VE + +Monitor Proxmox VE clusters, nodes, guests, and storage pools. + +Your whole cluster shows up alongside the rest of your telemetry - every node, guest and storage pool from a single scrape, on dashboards that are already built. You can see how the cluster is behaving without opening Proxmox to look, and next to the applications running on it. + +!!! info "What this integration does" + It provisions **dashboards**. It does not collect anything itself and never talks to your Proxmox cluster - the metrics come from **pve-exporter** running somewhere you control, scraped by **Grafana Alloy** and forwarded to OpsPilot. Until both are running, the dashboards are empty. + +Navigate to **Integrations** from the left-hand sidebar, then select **Proxmox VE**. + +--- + +## Permission tiers + +Proxmox VE runs **Read-only**, which is the default and requires nothing - the metrics endpoint and signing key come from the service environment. + +--- + +## 1. Create a Proxmox API token + +The exporter needs read access to the cluster. In the Proxmox web UI: + +1. **Datacenter → Permissions → Users**, add a user - `prometheus` is the convention. Set the realm to **Proxmox VE authentication server**, which is what makes it `prometheus@pve`. The PAM realm would give you `prometheus@pam` and the token below would not match. +2. **Datacenter → Permissions → API Tokens**, add a token for that user. Clear **Privilege Separation** so the token inherits the user's permissions. Copy the secret now - Proxmox shows it once. +3. **Datacenter → Permissions**, add a permission: path `/`, the user you created, role `PVEAuditor`, propagate on. + +`PVEAuditor` is read-only. The exporter never needs more than that. + +The token and permission can be created from a shell on any node once the user exists: + +```bash +pveum user token add prometheus@pve monitoring -privsep 0 +pveum acl modify / --user prometheus@pve --role PVEAuditor +``` + +The first command prints the token value. Copy it now. + +--- + +## 2. Run pve-exporter + +It needs to reach the Proxmox API on port 8006, so run it on a host that can - a management VM, a container host, or one of the nodes. + +```yaml +services: + pve-exporter: + image: prompve/prometheus-pve-exporter:latest + restart: unless-stopped + ports: + - "9221:9221" + environment: + PVE_USER: prometheus@pve + PVE_TOKEN_NAME: monitoring + PVE_TOKEN_VALUE: ${PVE_TOKEN_VALUE} + PVE_VERIFY_SSL: "true" +``` + +!!! warning "About PVE_VERIFY_SSL" + This is the exporter's own default, set here for visibility rather than to change anything. Proxmox ships with a self-signed certificate, so if you have not replaced it the exporter refuses to connect, because it cannot verify what it is talking to. + + Either add your cluster's CA to the exporter's trust store, or set `PVE_VERIFY_SSL: "false"` - which skips the check entirely, and means an attacker positioned between the exporter and the cluster could capture the API token. On a trusted network that is usually an acceptable trade, but make it deliberately rather than by default. + +Check it works, substituting one of your node's hostnames: + +```bash +curl "http://localhost:9221/pve?module=default&target=pve-node1" +``` + +You should get a few hundred lines beginning `pve_`. Note the path is `/pve`, not `/metrics`, and that the target is passed as a query parameter - the exporter queries the cluster API rather than reading the local machine. + +--- + +## 3. Configure Alloy + +If you have no collector yet, the [Grafana Alloy guide](/Monitor-your-data/OpenTelemetry/Shipping/Collector/) covers installing one and getting an API key. That guide configures Alloy for OTLP; Proxmox metrics are Prometheus-format, so the components below are different, but the install and the API key are the same. + +Add this to your `config.alloy`: + +```river +prometheus.scrape "proxmox_ve" { + targets = [{ + __address__ = ":9221", + __metrics_path__ = "/pve", + "__param_module" = "default", + "__param_target" = "", + }] + + job_name = "proxmox-ve" + scrape_interval = "30s" + scrape_timeout = "10s" + + forward_to = [prometheus.remote_write.opspilot.receiver] +} + +prometheus.remote_write "opspilot" { + endpoint { + url = "https://api.fusionreactor.io/v1/metrics" + + headers = { + "authorization" = sys.env("OPSPILOT_API_KEY"), + } + } + + external_labels = { + cluster = "", + } +} +``` + +Two values matter more than they look: + +| Value | Why it matters | +|---|---| +| `job_name = "proxmox-ve"` | Required. Every panel filters on it, because your metrics land in a store shared with everything else you send to OpsPilot. Alloy's default job label is the component's own id, which would change if you renamed the component, so it is set explicitly. Change it and the dashboards go blank | +| `cluster` | Yours to choose, and it populates the dashboard's cluster selector. Give each cluster a distinct name if you run more than one - otherwise they produce identical series and the panels cannot tell them apart | + +`__param_target` only needs one node. The exporter asks that node's API about the whole cluster, so you get every node, guest, and storage pool from a single scrape. If you run the exporter directly on a Proxmox node rather than a separate host, `__param_target` can be dropped entirely - it defaults to localhost. + +Restart Alloy and check `http://:12345` - the `prometheus.scrape` component should show the target as up. + +--- + +## 4. Confirm the data arrived + +In OpsPilot, open **Explore**, select the **Metrics** data source, and run: + +```promql +pve_up{job="proxmox-ve"} +``` + +You should get one series per node, guest, and storage pool. If that returns data, the dashboards will too. + +--- + +## Troubleshooting + +**Nothing in Explore.** Check the Alloy UI first - a failing scrape shows there with the error. If the target is up, check the `job` label really is `proxmox-ve` by querying `pve_up` with no selector and reading the labels back. + +**Authentication failures from the exporter.** Privilege separation left on is the usual cause. A token with it enabled has no permissions of its own, regardless of what the user can do. + +**Guests show `n/a` for disk usage.** Expected. Proxmox reports filesystem usage for LXC containers only - it cannot see inside a QEMU guest's disk. The allocated size is still shown. + +**Two clusters overlapping.** Both are sending the same `cluster` label. Give them distinct names in `external_labels`. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/Infrastructure/unix.md b/docs/Data-insights/Features/Integrations/Infrastructure/unix.md new file mode 100644 index 0000000..d6c2bc1 --- /dev/null +++ b/docs/Data-insights/Features/Integrations/Infrastructure/unix.md @@ -0,0 +1,106 @@ +# Unix + +Monitor Unix and Linux hosts - CPU, memory, filesystems, and processes. + +Your hosts show up alongside the rest of your telemetry on a dashboard that is already built, with a selector to move between them. When an application slows down, you can see whether the machine underneath it is the reason. + +!!! info "What this integration does" + It provisions a **dashboard**. It does not collect anything itself and never connects to your hosts - the metrics come from **Grafana Alloy**, which has a node_exporter collector built in. Until Alloy is collecting them, the dashboard is empty. + +Navigate to **Integrations** from the left-hand sidebar, then select **Unix**. + +--- + +## Permission tiers + +Unix runs **Read-only**, which is the default and requires nothing - the metrics endpoint and signing key come from the service environment. + +--- + +## 1. Configure Alloy + +If you have no collector yet, the [Grafana Alloy guide](/Monitor-your-data/OpenTelemetry/Shipping/Collector/) covers installing one and getting an API key. That guide configures Alloy for OTLP; host metrics are Prometheus-format, so the components below are different, but the install and the API key are the same. + +Add this to your `config.alloy` on **each host you want to monitor**: + +```river +prometheus.exporter.unix "host" { } + +prometheus.scrape "host" { + targets = prometheus.exporter.unix.host.targets + + job_name = "unix" + scrape_interval = "15s" + + forward_to = [prometheus.remote_write.opspilot.receiver] +} + +prometheus.remote_write "opspilot" { + endpoint { + url = "https://api.fusionreactor.io/v1/metrics" + + headers = { + "authorization" = sys.env("OPSPILOT_API_KEY"), + } + } +} +``` + +`job_name` can be `unix` or `node` - the dashboard accepts either. `node` is the long-standing convention for node_exporter, so if you are already scraping hosts under that name there is nothing to change. What matters is that it is one of the two: every panel filters on it, because your metrics land in a store shared with everything else you send to OpsPilot. + +Each host appears in the dashboard's host selector under its `instance` label, which Alloy sets from the scrape target. + +--- + +## 2. Run Alloy with access to the host + +The collector reads from `/proc` and `/sys`, so if Alloy runs as a container it needs them mounted and the root path pointed at them: + +```yaml +services: + alloy: + image: grafana/alloy:latest + pid: host + volumes: + - /:/rootfs:ro + - ./config.alloy:/etc/alloy/config.alloy +``` + +and in the exporter block: + +```river +prometheus.exporter.unix "host" { + rootfs_path = "/rootfs" +} +``` + +Without that, the collector reports the container's own view: every mount shows the size of the underlying filesystem rather than its own, and the hostname is a container id. Running Alloy directly on the host avoids the question entirely. + +--- + +## 3. Confirm the data arrived + +In OpsPilot, open **Explore**, select the **Metrics** data source, and run: + +```promql +node_uname_info{job="unix"} +``` + +You should get one series per host, carrying its `nodename` and `release`. If that returns data, the dashboard will too. + +--- + +## Troubleshooting + +**Nothing in Explore.** Check the Alloy UI on port 12345 first - a failing component shows there with the error. If the scrape is healthy, check the `job` label really is `unix` or `node` by querying `node_uname_info` with no selector and reading the labels back. + +**Every filesystem reports the same size.** Alloy is running in a container without the root filesystem mounted, so the collector sees the container's view rather than the host's. See step 2. + +**A filesystem you expected is missing.** The panel excludes pseudo-filesystems - tmpfs, overlay, squashfs and the rest - because they report the memory or image backing them rather than disk. A real mount with an unusual filesystem type may be caught by that too; the exclusion list is in the panel's query. + +**The host selector shows a container id.** `instance` comes from the scrape target, which is the Alloy container when Alloy runs in Docker. Set it explicitly with a relabel rule if you want the hostname instead. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/Infrastructure/windows.md b/docs/Data-insights/Features/Integrations/Infrastructure/windows.md new file mode 100644 index 0000000..0601e64 --- /dev/null +++ b/docs/Data-insights/Features/Integrations/Infrastructure/windows.md @@ -0,0 +1,170 @@ +# Windows + +Monitor Windows hosts - CPU, memory, disks, and network. + +Your Windows hosts show up alongside the rest of your telemetry on a dashboard that is already built, including which services set to start automatically are not running. When an application slows down or stops responding, you can see whether the machine or a service underneath it is the reason. + +!!! info "What this integration does" + It provisions a **dashboard**. It does not collect anything itself and never connects to your hosts - the metrics come from **Grafana Alloy**, which has a windows_exporter collector built in. Until Alloy is collecting them, the dashboard is empty. + +Navigate to **Integrations** from the left-hand sidebar, then select **Windows**. + +--- + +## Permission tiers + +Windows runs **Read-only**, which is the default and requires nothing - the metrics endpoint and signing key come from the service environment. + +--- + +## 1. Install Alloy on the host + +Alloy runs as a Windows service. Download the Windows installer from [Grafana's releases](https://github.com/grafana/alloy/releases) and run it, or install with winget: + +```powershell +winget install Grafana.Alloy +``` + +The installer puts the configuration at `C:\Program Files\GrafanaLabs\Alloy\config.alloy`. + +--- + +## 2. Configure it + +If you have no API key yet, the [Grafana Alloy guide](/Monitor-your-data/OpenTelemetry/Shipping/Collector/) covers getting one. That guide configures Alloy for OTLP; host metrics are Prometheus-format, so the components below are different, but the API key is the same. + +```river +prometheus.exporter.windows "host" { + enabled_collectors = [ + "cpu", + "memory", + "logical_disk", + "net", + "os", + "system", + "service", + ] +} + +prometheus.scrape "host" { + targets = prometheus.exporter.windows.host.targets + + job_name = "windows" + scrape_interval = "15s" + + forward_to = [prometheus.relabel.trim_services.receiver] +} + +// The service collector is by far the largest thing in that list - see below. +// These rules keep the two label values this dashboard reads and drop the rest. +prometheus.relabel "trim_services" { + forward_to = [prometheus.remote_write.opspilot.receiver] + + // Neither of these is read by the dashboard. + rule { + source_labels = ["__name__"] + regex = "windows_service_(info|process)" + action = "drop" + } + + // windows_service_state ships one series per service per state. Only + // "stopped" is read. + rule { + source_labels = ["__name__", "state"] + separator = ";" + regex = "windows_service_state;(continue pending|pause pending|paused|running|start pending|stop pending|unknown)" + action = "drop" + } + + // windows_service_start_mode ships one per service per mode. Only "auto" is + // read. + rule { + source_labels = ["__name__", "start_mode"] + separator = ";" + regex = "windows_service_start_mode;(boot|disabled|manual|system)" + action = "drop" + } +} + +prometheus.remote_write "opspilot" { + endpoint { + url = "https://api.fusionreactor.io/v1/metrics" + + headers = { + "authorization" = sys.env("OPSPILOT_API_KEY"), + } + } +} +``` + +Then restart the service: + +```powershell +Restart-Service Alloy +``` + +### Four things that matter more than they look + +**`memory` is not a default collector.** Alloy's defaults are `cpu`, `logical_disk`, `net`, `os`, `service` and `system` - no memory. Leave it out and the memory panels are permanently empty with nothing to explain why, so it is listed explicitly above. + +**`job_name = "windows"` is required.** Every panel filters on it, because your metrics land in a store shared with everything else you send to OpsPilot. Alloy's default job label is the component's own id, which would change if you renamed the component, so it is set explicitly. Change it and the dashboard goes blank. + +**windows_exporter 0.25 or newer.** The memory collector's `windows_memory_physical_total_bytes` and `windows_memory_physical_free_bytes` are relatively recent, and `windows_os_info` - which the host selector is built from - has to be present, or the dashboard shows nothing at all rather than partially working. Any Alloy release from 2024 onwards embeds a new enough exporter. + +**The service collector is the expensive one.** It emits four metrics per service, two of them multiplied out across every possible value: `windows_service_state` once per state (eight of them) and `windows_service_start_mode` once per start mode (five). On a Windows Server with around 250 services that is roughly **3,750 active series**, against a couple of hundred for every other collector in this list combined - and your metrics are metered. This dashboard reads exactly two of those values, `state="stopped"` and `start_mode="auto"`, so the relabel rules above drop the rest and take it to about **500**. + +--- + +## Narrowing it further + +If you only care about specific services rather than every automatic one, filter at collection instead and skip the relabel entirely: + +```river +prometheus.exporter.windows "host" { + enabled_collectors = [...] + + service { + include = "MSSQLSERVER|W3SVC|YourAppService" + } +} +``` + +That is cheaper again, at the cost of the dashboard only knowing about the services you named. + +!!! warning "Keep the collector list short" + Several collectors fail, or take the whole Alloy process down, when the thing they measure is not installed - `mscluster`, `vmware`, `hyperv`, `ad`, `dns`, `msmq` and `nps` among them. The list above is what this dashboard reads and nothing more. Add others deliberately, one at a time, and check Alloy is still running afterwards. + +Each host appears in the dashboard's host selector under its `instance` label, which Alloy sets from the scrape target - the machine's own hostname. + +If you already run windows_exporter as a standalone service and would rather point Alloy at that than enable the built-in collector, that works: scrape `http://localhost:9182/metrics` instead. The only difference is that `instance` then reads `HOSTNAME:9182` rather than the bare hostname, so that is what the host selector and every legend will show. + +--- + +## 3. Confirm the data arrived + +In OpsPilot, open **Explore**, select the **Metrics** data source, and run: + +```promql +windows_os_info{job="windows"} +``` + +You should get one series per host. If that returns data, the dashboard will too. + +--- + +## Troubleshooting + +**Nothing in Explore.** Check the Alloy UI at `http://localhost:12345` on the host first - a failing component shows there with the error. If the scrape is healthy, check the `job` label really is `windows` by querying `windows_os_info` with no selector and reading the labels back. + +**The memory panels are empty but everything else works.** The memory collector is not enabled. See step 2. + +**The service panels are empty.** The service collector is not enabled, or the Alloy service lacks the rights to enumerate services. It runs as LocalSystem by default, which has them. + +**Stopped services shows a count but the table below is empty.** The table also filters on start mode, so it lists only services set to start automatically - a stopped service set to Manual is stopped by design and is left out deliberately. + +**A disk you expected is missing.** Volumes reporting a size of zero are excluded, because dividing by that gives nothing useful. An empty card reader or an unmounted optical drive reports zero. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/SDKs-view.png b/docs/Data-insights/Features/Integrations/SDKs-view.png new file mode 100644 index 0000000..2b89c74 Binary files /dev/null and b/docs/Data-insights/Features/Integrations/SDKs-view.png differ diff --git a/docs/Data-insights/Features/Integrations/SDKs/sdk-integrations.md b/docs/Data-insights/Features/Integrations/SDKs/sdk-integrations.md new file mode 100644 index 0000000..e37c3d4 --- /dev/null +++ b/docs/Data-insights/Features/Integrations/SDKs/sdk-integrations.md @@ -0,0 +1,63 @@ +# SDK Integrations + +The SDK integrations provision runtime dashboards and alert rules for applications instrumented with OpenTelemetry. Install one, instrument your application, and you have runtime monitoring in place without building it yourself. + +Navigate to **Integrations** from the left-hand sidebar and open the **SDKs** tab. + +--- + +## What they provision + +Each SDK integration installs one runtime dashboard per upstream metric-set version, plus a handful of runtime alert rules. + +| SDK | Dashboards | Alert rules | Alert group | +|---|---|---|---| +| Java | 26 | 4 | `java_runtime_alerts` | +| .NET | 8 | 4 | `dotnet_runtime_alerts` | +| Go | 6 | 4 | `go_runtime_alerts` | +| Node.js | 9 | 3 | `nodejs_runtime_alerts` | +| Python | 5 | 5 | `python_runtime_alerts` | + +The dashboard counts differ because each covers that language's full range of upstream metric-set versions. Java's 26 span thirteen Java-agent eras from v0.11.0 and thirteen semantic-convention sets from 1.9.0. + +!!! note "Ruby" + The Ruby OpenTelemetry SDK does not officially support metrics, so a Ruby integration cannot provide runtime dashboards. Ruby applications still send traces and span metrics, which Coworker can analyse. + +--- + +## Alerts ship paused + +The alert rules are installed **paused**. You opt in per rule rather than being alerted on everything from the moment you install, and once you enable a rule that choice persists across integration upgrades. + +Each SDK's rules sit in their own alert group, named in the table above, so you can find them together in Alerting. + +Every threshold is either scale-free or derived from the runtime itself, so the same rules apply unchanged whatever the size of your service. Go's four rules, for instance, cover heap usage against `GOMEMLIMIT`, goroutine count, GC pause overhead, and scheduler latency. + +--- + +## Dashboards and data sources + +Provisioned dashboards are tagged `integration` along with a tag for the language, such as `go`, so you can find everything an integration added. + +The SDK integrations bind to your account's existing shared metrics data source rather than creating their own, so you keep a single Prometheus source for all metrics. That source is adopted, not owned - uninstalling the integration leaves it in place. + +--- + +## Installing + +SDK integrations run **Read-only**, which is the default and requires nothing: the metrics endpoint and signing key come from the service environment. There is no configuration to complete during the install itself. + +Once installed, the **Installation guide** tab on the integration's detail view walks through instrumenting your application and points you to the dashboard to open. See [Integrations](../../integrations.md) for the install flow and how to manage an installation afterwards. + +### Instrumenting your application + +These dashboards are built on **runtime** metrics - the heap, garbage collector, threads and scheduler your language exposes about itself. Your application has to emit them, which takes the runtime instrumentation for that language: the Go integration wants `go.opentelemetry.io/contrib/instrumentation/runtime` registered at startup, and each of the others has its equivalent. + +The per-language steps live in the **Installation guide** tab on each integration, which is written against the metric sets these dashboards expect. Follow that first. + +The OpenTelemetry [instrumentation guides](/Monitor-your-data/OpenTelemetry/Instrumentation/Overview/) cover general instrumentation - exporting metrics, traces and logs to OpsPilot - and are worth reading alongside, but most of them do not set up runtime instrumentation, so they will not fill these dashboards on their own. + +--- + +!!! question "Need more help?" + Contact support in the chat bubble and let us know how we can assist. diff --git a/docs/Data-insights/Features/Integrations/Ticketing/jira.md b/docs/Data-insights/Features/Integrations/Ticketing/jira.md index 58c04e0..a9ecbc4 100644 --- a/docs/Data-insights/Features/Integrations/Ticketing/jira.md +++ b/docs/Data-insights/Features/Integrations/Ticketing/jira.md @@ -2,13 +2,18 @@ Connect OpsPilot to Jira to automatically create issues from alerts and incidents. -Navigate to **Integrations** from the left-hand sidebar, then select **Jira**. +!!! note "Legacy integration" + This page covers the **legacy** Jira integration, which is still available and still works. To find it, navigate to [Integrations](../../integrations.md) and switch on **Legacy**, at the right-hand end of the category tabs. + + A new Jira integration appears in the main integration catalog as **Coming soon**, and will replace this one. + +Navigate to **Integrations** from the left-hand sidebar, switch on **Legacy**, then select **Jira**. --- ## Setup -1. In OpsPilot, go to **Integrations** and click **Jira**. +1. In OpsPilot, go to **Integrations**, switch on **Legacy**, and click **Jira**. 2. Enter your Jira instance URL (e.g., `https://yourorg.atlassian.net`). 3. Provide your Jira **email address** and an **API token**. You can generate an API token from your [Atlassian account settings](https://id.atlassian.com/manage-profile/security/api-tokens). 4. Select the default **project** where issues should be created. diff --git a/docs/Data-insights/Features/Integrations/chat-view.png b/docs/Data-insights/Features/Integrations/chat-view.png new file mode 100644 index 0000000..472f8de Binary files /dev/null and b/docs/Data-insights/Features/Integrations/chat-view.png differ diff --git a/docs/Data-insights/Features/Integrations/cloud-view.png b/docs/Data-insights/Features/Integrations/cloud-view.png new file mode 100644 index 0000000..c4ded9e Binary files /dev/null and b/docs/Data-insights/Features/Integrations/cloud-view.png differ diff --git a/docs/Data-insights/Features/Integrations/data-view.png b/docs/Data-insights/Features/Integrations/data-view.png new file mode 100644 index 0000000..32b0931 Binary files /dev/null and b/docs/Data-insights/Features/Integrations/data-view.png differ diff --git a/docs/Data-insights/Features/Integrations/infrastructure-view.png b/docs/Data-insights/Features/Integrations/infrastructure-view.png new file mode 100644 index 0000000..9595832 Binary files /dev/null and b/docs/Data-insights/Features/Integrations/infrastructure-view.png differ diff --git a/docs/Data-insights/Features/Integrations/networking-view.png b/docs/Data-insights/Features/Integrations/networking-view.png new file mode 100644 index 0000000..237af57 Binary files /dev/null and b/docs/Data-insights/Features/Integrations/networking-view.png differ diff --git a/docs/Data-insights/Features/Integrations/observability-view.png b/docs/Data-insights/Features/Integrations/observability-view.png new file mode 100644 index 0000000..b08725c Binary files /dev/null and b/docs/Data-insights/Features/Integrations/observability-view.png differ diff --git a/docs/Data-insights/Features/Integrations/ticketing-view.png b/docs/Data-insights/Features/Integrations/ticketing-view.png new file mode 100644 index 0000000..36fc202 Binary files /dev/null and b/docs/Data-insights/Features/Integrations/ticketing-view.png differ diff --git a/docs/Data-insights/Features/integrations.md b/docs/Data-insights/Features/integrations.md index 40d324e..22dd15b 100644 --- a/docs/Data-insights/Features/integrations.md +++ b/docs/Data-insights/Features/integrations.md @@ -1,128 +1,289 @@ # Integrations Hub -![Integrations hub](../../images/integrations.png) +=== "Integrations Hub" -The **Integrations** page lets you connect external tools and services to your OpsPilot workspace. + ![Integrations hub](../../images/integrations.png) -Navigate to **Integrations** from the left-hand sidebar to browse and manage all available integrations. +=== "Chat" + ![Chat integrations](Integrations/chat-view.png) +=== "Cloud" + + ![Cloud integrations](Integrations/cloud-view.png) + +=== "Data" + + ![Data integrations](Integrations/data-view.png) + +=== "Infrastructure" + + ![Infrastructure integrations](Integrations/infrastructure-view.png) + +=== "Networking" + + ![Networking integrations](Integrations/networking-view.png) + +=== "Observability" + + ![Observability integrations](Integrations/observability-view.png) + +=== "SDKs" + + ![SDK integrations](Integrations/SDKs-view.png) + +=== "Ticketing" + + ![Ticketing integrations](Integrations/ticketing-view.png) + +Integrations bring your data into OpsPilot from wherever it lives - chat tools, cloud providers, language SDKs, Kubernetes, databases, and more - so you can work with all of it in one place. Most install in a single click and arrive with dashboards and alerts already set up for you. + +Navigate to **Integrations** from the left-hand sidebar to browse the integration catalog and manage what you've installed. --- ## Browsing integrations -Integrations are grouped into categories. Use the filter tabs at the top of the page to narrow the list: +Integrations are grouped into categories, each showing how many it holds. Use the filter tabs at the top of the page to narrow the list: | Tab | Description | |---|---| -| **All** | Shows every available integration | +| **All** | Shows every integration | | **Chat** | Messaging and notification tools | | **Cloud** | Cloud platform providers | | **Data** | Databases and data streaming services | | **Infrastructure** | Infrastructure and orchestration tools | | **Networking** | Service mesh and proxy tools | | **Observability** | Third-party monitoring and observability platforms | -| **SDKs** | Language SDKs and FusionReactor agent | +| **SDKs** | Language SDKs | | **Ticketing** | Issue tracking and project management tools | -You can also use the **Search integrations** bar to find a specific integration by name. +Use the **Search integrations** bar to find one by name, and the status dropdown beside it to filter by whether an integration is installed. The **Legacy** toggle, at the right-hand end of the category tabs, switches to the legacy integrations - earlier ones that are still available and still work. The legacy view lists those integrations on their own, without the category tabs, search, or state counters. ---- +Each card shows the integration's name, its category, a short description, and its current status - **Coming soon** for one not yet released, or **Installed** for one already connected. -## Available integrations +## User MCPs -### Chat +**User MCPs**, in the toolbar beside the Legacy toggle, is a separate page for MCP integrations. These work differently from the rest of the integration catalog. Rather than being set up once for the organisation, **each person connects their own account**, and OpsPilot uses that connection only when answering that person in chat: -| Integration | Status | -|---|---| -| Slack | Available | -| Discord | Coming Soon | -| MS Teams | Coming Soon | +> These connect your own account, not your organisation's. The agent uses them only in chat, never in scheduled work or on anyone else's behalf. -### SDKs +## Installing an integration -| Integration | Status | -|---|---| -| [FusionReactor](https://docs.fusionreactor.io/Getting-started/install-fr/) | Installed | -| [.NET](/Monitor-your-data/OpenTelemetry/Instrumentation/DotNet/) | Coming Soon | -| Browser | Coming Soon | -| [C++](/Monitor-your-data/OpenTelemetry/Instrumentation/Cpp/) | Coming Soon | -| [Erlang](/Monitor-your-data/OpenTelemetry/Instrumentation/Erlang/) | Coming Soon | -| [Go](/Monitor-your-data/OpenTelemetry/Instrumentation/Go/) | Coming Soon | -| [Java](/Monitor-your-data/OpenTelemetry/Instrumentation/Java/) | Coming Soon | -| [Node.js](/Monitor-your-data/OpenTelemetry/Instrumentation/node/) | Coming Soon | -| [PHP](/Monitor-your-data/OpenTelemetry/Instrumentation/PHP/) | Coming Soon | -| [Python](/Monitor-your-data/OpenTelemetry/Instrumentation/Python/) | Coming Soon | -| [Ruby](/Monitor-your-data/OpenTelemetry/Instrumentation/Ruby/) | Coming Soon | -| [Rust](/Monitor-your-data/OpenTelemetry/Instrumentation/Rust/) | Coming Soon | -| [Swift](/Monitor-your-data/OpenTelemetry/Instrumentation/Swift/) | Coming Soon | - -### Ticketing - -| Integration | Status | -|---|---| -| Jira | Available | -| Linear | Coming Soon | -| Notion | Coming Soon | +For most integrations, onboarding is a single click. Click **Install** on an integration's card or from its detail view to open the install dialog, which shows: -### Cloud +- The **Permission tier** to install with, and what that tier requires +- Whether any configuration is needed - *This integration needs no configuration* means there is nothing else to set up +- Which account the install will serve, where that applies -| Integration | Status | -|---|---| -| AWS | Coming Soon | -| Azure | Coming Soon | -| Google Cloud | Coming Soon | +Click **Install** to confirm, or **Cancel** to back out. Confirming takes you straight to [Manage integration](#managing-an-installed-integration), where you can check the installation's health and change its permission tier. -### Data +Once connected, the integration's card in the integration catalog shows an **Installed** badge and its button changes to **Uninstall**. On the integration's own detail view, the **Install** button is replaced - by **Installed**, or by the actions that integration offers once it is in place, such as **Connect** and **Add instance**. -| Integration | Status | -|---|---| -| Altinity ClickHouse Operator | Coming Soon | -| Kafka | Coming Soon | -| MongoDB | Coming Soon | -| MySQL | Coming Soon | -| PostgreSQL | Coming Soon | -| RabbitMQ | Coming Soon | -| Redis | Coming Soon | -| Strimzi Kafka | Coming Soon | -| TigerData | Coming Soon | - -### Infrastructure - -| Integration | Status | +Some integrations need more than a permission tier - credentials, endpoints, or other configuration. Where that applies, the steps are built into the UI, so you can work through them without leaving OpsPilot. [Slack](Integrations/Chat/slack.md) has its own **Add to Slack** button, which starts the Slack authorisation flow, and [AWS](Integrations/Cloud/aws.md) needs an IAM role or key with the right CloudWatch permissions. + +### What you get + +Installing an integration does more than connect a data source. OpsPilot automatically loads a set of default **dashboards** and **alerts** onto your account for that integration, so you have useful monitoring in place with nothing to build by hand. + +How much you get for that click varies. Some integrations are self-contained - install one and it starts working. Others provision dashboards only, and depend on a collector you run and configure yourself, so their dashboards stay empty until you have set that up. [Docker](Integrations/Infrastructure/docker.md), [Proxmox VE](Integrations/Infrastructure/proxmox-ve.md), [Unix](Integrations/Infrastructure/unix.md) and [Windows](Integrations/Infrastructure/windows.md) are of the second kind: each needs Grafana Alloy collecting and forwarding the metrics before anything appears. Their **Installation guide** tab carries the steps. + +The **Capabilities** panel on an integration's detail view names what it provisions. AWS, for example, declares **Data Sources**, **Dashboards**, **Recording Rules**, and **Alerts**. Some integrations declare none. + +What arrives, and whether it is active straight away, varies by integration. The SDK integrations provision a runtime dashboard for each upstream metric-set version, plus a handful of runtime alert rules - Java installs 26 dashboards, Node.js nine, .NET eight, Go six, and Python five. Those alert rules ship **paused**, so you opt in per rule rather than being alerted on everything from day one, and once you enable a rule that choice persists across upgrades. Every threshold is either scale-free or derived from the runtime itself, so the rules apply unchanged whatever the size of your service. + +Each SDK's rules sit in their own alert group, named after the language - `java_runtime_alerts`, for example. Check the **Versions** tab on an integration's detail view for what its current version installs. + +Provisioned dashboards are tagged `integration` along with a tag for the integration itself, so you can find everything one of them added. Where an integration needs a shared resource such as a metrics data source, it adopts the account's existing one rather than creating its own - uninstalling the integration leaves that source in place. + +### The integration detail view + +Click any integration to open its detail view - this is where you'll find how to use it. The header shows the integration's name, its **category**, its **version**, and the **Install** button, with the same short description that appears on its card. + +Below the header, the detail view is made up of these panels. Not every integration shows all of them - some have no **Overview**, for example: + +| Panel | Description | |---|---| -| ArgoCD | Coming Soon | -| Host Metrics | Coming Soon | -| KEDA | Coming Soon | -| Kubernetes | Coming Soon | -| Terraform | Coming Soon | +| **Overview** | What the integration does and why you'd use it | +| **Permission tiers** | The access levels the integration can run with, what each one requires, and which is applied by **default**. Tiers vary by integration - AWS offers **Read-only** and **Read + Write**, each needing different AWS IAM permissions, while OpsPilot MCP offers **Read-only** and **Read + Act** | +| **Capabilities** | Badges naming what the integration provisions, such as **Data Sources**, **Dashboards**, **Recording Rules**, and **Alerts**. Some integrations declare none | -### Networking +Below those panels sit three tabs: -| Integration | Status | +| Tab | Description | |---|---| -| Cilium | Coming Soon | -| Istio | Coming Soon | -| Linkerd | Coming Soon | -| NGINX | Coming Soon | -| Traefik | Coming Soon | +| **Installation guide** | Any setup needed beyond installing the integration. For the SDK integrations this is a full walkthrough of instrumenting your application - the Go guide covers adding the `opentelemetry-go-contrib` runtime instrumentation and registering it at startup, then points you to the dashboard to open. Where nothing further is needed, the tab reads *This integration needs no setup beyond installing it* | +| **Versions** | Each released version, what upgrading from the previous one involves (**Initial version**, **Automatic**, or **Manual** where the upgrade takes action), and what changed in each. The current version is marked **latest** | +| **Licenses** | Any third-party work the integration includes, and the terms it carries. Where there is none, the tab reads *This integration includes no third-party work* | + +### Your installation -### Observability +Once you have installed an integration, a **Your installation** panel appears at the top of its detail view, above **Permission tiers**, listing what you have installed: -| Integration | Status | +| Column | Description | |---|---| -| AppDynamics | Coming Soon | -| Dash0 | Coming Soon | -| Datadog | Coming Soon | -| Grafana | Coming Soon | -| Loki | Coming Soon | -| Mimir | Coming Soon | -| New Relic | Coming Soon | -| Sentry | Coming Soon | -| Tempo | Coming Soon | +| **Name** | The name of your installation - the language for an SDK (such as, `dotnet`), or an identifier where the integration has no natural name | +| **Version** | The version you have installed | +| **Tier** | The active permission tier (such as, `read` or `act`) | +| **Health** | The installation's current health, such as **Healthy** | + +Click the row to open **Manage integration**. Some installations also carry a **Manage** button, and the **...** menu at the right-hand end offers **Manage** and **Uninstall**. + +An integration that supports more than one installation shows **Add instance** in the header, so it can be installed again - [AWS](Integrations/Cloud/aws.md), for example, is installed once per region. + +### Managing an installed integration + +**Manage integration** opens as soon as you confirm an install, and you can return to it later from the **Your installation** panel. It shows the installation's name and health, with the integration it was installed from, its category, and the installed version beneath, and an **Uninstall** button in the top right. + +The **Permission tier** panel shows the **Active tier**, which you can change after installing: + +- **Narrowing** the tier, granting the integration less access, applies straight away +- **Widening** it needs fresh credentials, so it means reinstalling the integration + +The **Capabilities** panel shows what the integration provisions, the same as on its detail view. + +--- + +## Available now + +
+ +- :material-robot-outline: **[OpsPilot MCP](Integrations/Chat/opspilot-mcp.md)** - Chat + + --- + + Let AI assistants query OpsPilot over the Model Context Protocol. + +- :material-chat-outline: **[Slack](Integrations/Chat/slack.md)** - Chat + + --- + + Talk to OpsPilot from Slack - mention it in a channel or DM it directly. + +- :material-cloud-outline: **[AWS](Integrations/Cloud/aws.md)** - Cloud + + --- + + Connect AWS for EC2, RDS, and other CloudWatch-held metrics. + +- :material-code-tags: **[OpenTrace](Integrations/Data/opentrace.md)** - Data + + --- + + Connect OpenTrace so OpsPilot can read your code alongside your telemetry. + +- :material-server-network: **[Docker](Integrations/Infrastructure/docker.md)** - Infrastructure + + --- + + Monitor Docker containers - CPU, memory against limit, network and disk IO. + +- :material-server-network: **[Proxmox VE](Integrations/Infrastructure/proxmox-ve.md)** - Infrastructure + + --- + + Monitor Proxmox VE clusters, nodes, guests, and storage pools. + +- :material-server-network: **[Unix](Integrations/Infrastructure/unix.md)** - Infrastructure + + --- + + Monitor Unix and Linux hosts - CPU, memory, filesystems, and processes. + +- :material-server-network: **[Windows](Integrations/Infrastructure/windows.md)** - Infrastructure + + --- + + Monitor Windows hosts - CPU, memory, disks, and network. + +
+ +### SDKs + +Each SDK instruments your applications with OpenTelemetry for metrics, traces, and logs, and provisions runtime dashboards and alert rules for them. See [SDK Integrations](Integrations/SDKs/sdk-integrations.md) for what each one installs. + +
+ +- :material-code-braces: **Go** + + --- + + [OpenTelemetry instrumentation](/Monitor-your-data/OpenTelemetry/Instrumentation/Go/) + +- :material-dot-net: **.NET** + + --- + + [OpenTelemetry instrumentation](/Monitor-your-data/OpenTelemetry/Instrumentation/DotNet/) + +- :material-code-braces: **Java** + + --- + + Instruments Java and JVM applications. [OpenTelemetry instrumentation](/Monitor-your-data/OpenTelemetry/Instrumentation/Java/) + +- :material-nodejs: **Node.js** + + --- + + [OpenTelemetry instrumentation](/Monitor-your-data/OpenTelemetry/Instrumentation/node/) + +- :material-language-python: **Python** + + --- + + [OpenTelemetry instrumentation](/Monitor-your-data/OpenTelemetry/Instrumentation/Python/) + +
+ +### FusionReactor + +The [FusionReactor agent](Integrations/SDKs/fusionreactor.md) is **managed automatically** - it cannot be modified or removed. It provides: + +- Discovery of services +- Application performance monitoring, to identify bottlenecks and optimize response times +- Centralized log collection, search, and analysis across all your applications +- Real-time code-level profiling to detect resource consumption patterns +- Visibility into servers, Kubernetes, and system-level resource utilization +- Intelligent anomaly detection and alerting for rapid incident response + +--- + +## Coming soon + +=== "Chat" + + Discord · MS Teams + +=== "Cloud" + + Azure · Google Cloud + +=== "Data" + + Altinity ClickHouse Operator · Kafka · MongoDB · MySQL · PostgreSQL · RabbitMQ · Redis · Strimzi Kafka · TigerData + +=== "Infrastructure" + + ArgoCD · Host Metrics · iDRAC · KEDA · Kubernetes · Terraform · TrueNAS SCALE + +=== "Networking" + + Cilium · Istio · Linkerd · NGINX · Traefik + +=== "Observability" + + AppDynamics · Dash0 · Datadog · Grafana · Loki · Mimir · New Relic · Sentry · Tempo + +=== "SDKs" + + Browser · C++ · Erlang · PHP · Ruby · Rust · Swift + +=== "Ticketing" + + Jira · Linear · Notion --- -!!! question "Need more help?" - Contact support in the chat bubble and let us know how we can assist. +!!! question "Don't see the integration you need?" + If the integration you're looking for isn't in the list, contact support in the chat bubble and let us know which one you need. diff --git a/docs/stylesheets/extra.css b/docs/stylesheets/extra.css index 8edd5b2..6245ed5 100644 --- a/docs/stylesheets/extra.css +++ b/docs/stylesheets/extra.css @@ -128,7 +128,7 @@ .md-typeset .grid.cards>ol>li, .md-typeset .grid.cards>ul>li, .md-typeset .grid>.card { - border: 1px solid #8b8686; /* Darker border */ + border: 1px solid var(--card-border); transition: border .75s,box-shadow .75s } @@ -136,18 +136,24 @@ /* Smaller grid cards custom colours */ :root { --card-bg-light: white; - --card-bg-dark: white; + --card-bg-dark: rgba(255, 255, 255, .04); + --card-border-light: #d5d2d2; + --card-border-dark: rgba(255, 255, 255, .14); --card-border-focus-light: black; - --card-border-focus-dark: #212441; + --card-border-focus-dark: #FF6600; } [data-md-color-scheme="default"] { --card-bg: var(--card-bg-light); + --card-border: var(--card-border-light); + --card-border-focus: var(--card-border-focus-light); } [data-md-color-scheme="slate"] { --card-bg: var(--card-bg-dark); + --card-border: var(--card-border-dark); + --card-border-focus: var(--card-border-focus-dark); } diff --git a/mkdocs.yml b/mkdocs.yml index a24eb73..227ee66 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -85,8 +85,19 @@ nav: - INTEGRATIONS: - Overview: Data-insights/Features/integrations.md - Chat: + - OpsPilot MCP: Data-insights/Features/Integrations/Chat/opspilot-mcp.md - Slack: Data-insights/Features/Integrations/Chat/slack.md + - Cloud: + - AWS: Data-insights/Features/Integrations/Cloud/aws.md + - Data: + - OpenTrace: Data-insights/Features/Integrations/Data/opentrace.md + - Infrastructure: + - Docker: Data-insights/Features/Integrations/Infrastructure/docker.md + - Proxmox VE: Data-insights/Features/Integrations/Infrastructure/proxmox-ve.md + - Unix: Data-insights/Features/Integrations/Infrastructure/unix.md + - Windows: Data-insights/Features/Integrations/Infrastructure/windows.md - SDKs: + - SDK Integrations: Data-insights/Features/Integrations/SDKs/sdk-integrations.md - FusionReactor: Data-insights/Features/Integrations/SDKs/fusionreactor.md - Ticketing: - Jira: Data-insights/Features/Integrations/Ticketing/jira.md @@ -335,6 +346,8 @@ markdown_extensions: - pymdownx.mark - attr_list - md_in_html + - pymdownx.tabbed: + alternate_style: true extra_css: