Browser

observability-aiops

Try it

Query Prometheus, Alertmanager, Grafana, and Loki with built-in RCA and governed writes.

What it does

39 MCP tools for self-hosted observability stacks — covers Prometheus HTTP API + PromQL queries, scrape-target health and rule evaluation, firing/pending alerts and Alertmanager silences, Grafana dashboards/datasources, and Loki log reads. Includes five analysis tools: firing-alert RCA, scrape-health classification, alert-noise/flap detection, log-error-burst RCA, and log-volume/cardinality checks. All write operations (silences, annotations, dashboard edits, config reloads) are guarded with a local audit log, undo-token recording, and risk-tier labels. Tokens are stored encrypted with Fernet + scrypt. Loki is read-only.

When to use it

  • Get a firing-alert count and root-cause the noise
  • Find why scrape targets are down and classify the error
  • Tame flapping alerts with deduplication recommendations
  • Investigate a Loki error burst and correlate logs to a firing alert

The skill document

Observability AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by the Prometheus or Grafana projects, Grafana Labs, or the CNCF. Prometheus, Alertmanager and Grafana are trademarks of their respective owners. Source at github.com/AIops-tools/Observability-AIops under the MIT license.

Governed self-hosted observability operations — 39 MCP tools across Prometheus (HTTP API + PromQL), Alertmanager (alerts + silences), Grafana (dashboards, datasources, folders), and Grafana Loki (bounded LogQL log reads + log RCA, read-only), every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.observability-aiops/, token/runaway budget guard, undo-token recording, and descriptive risk-tier labels. One config can span the whole stack. Bearer tokens are stored encrypted (~/.observability-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.

This is the self-hosted-observability complement to enterprise monitoring suites: it speaks the open Prometheus/Grafana APIs an SRE actually runs.

Standalone: the governance harness is bundled in the package (observability_aiops.governance) — no external skill-family dependency. Beyond the mock suite, the Prometheus/Alertmanager/Grafana surfaces have been exercised against a live Prometheus 3.x + Alertmanager + Grafana 13 stack; the Loki surface has not yet been exercised live (see docs/VERIFICATION.md).

What This Skill Does

GroupPlatformToolsCountR/W
MetricsPrometheusinstant_query, range_query, label_values, series_metadata4read
Targets & statusPrometheuslist_targets, target_scrape_health, dropped_targets, prometheus_config_status, prometheus_tsdb_status5read
RulesPrometheuslist_rules, rule_health2read
AlertsPrometheus/Alertmanagerfiring_alerts, pending_alerts, alertmanager_alerts, list_silences4read
GrafanaGrafanalist_dashboards, get_dashboard, list_datasources, datasource_health, list_folders5read
LokiLokiloki_labels, loki_label_values, loki_query, loki_tail_errors4read
Overview + analysesallobservability_overview + firing_alert_rca, target_scrape_health_analysis, alert_noise_and_flap_analysis4read
Log analyses + cross-signalLoki (+ Prometheus)log_error_burst_rca, log_volume_analysis, alert_log_context3read
WritesAlertmanager/Grafana/Prometheuscreate_silence, expire_silence (med) · create_annotation (med) · update_dashboard (med) · delete_dashboard (high) · reload_prometheus_config (med)6write

The three metric flagship analyses are transparent heuristics that report their numbers: firing_alert_rca joins each firing alert to its rule expression and maps it to a cause + action; target_scrape_health_analysis ranks down/erroring scrape targets and classifies each lastError; alert_noise_and_flap_analysis finds noisy/duplicate alerts and recommends a dedup/rollup. The two log analyses mirror this: log_error_burst_rca compares per-stream error counts against a baseline window and classifies each burst (new signature / volume spike / single-instance); log_volume_analysis ranks the highest-volume streams and warns on high-cardinality (high-churn) labels. alert_log_context bridges the two signals — it maps a firing Prometheus alert's labels to a Loki stream selector and pulls the correlated logs. Loki is read-only (no safe write surface).

Quick Install

uv tool install observability-aiops
observability-aiops init       # wizard: pick platform (prometheus/grafana) + encrypted token
observability-aiops doctor

When to Use This Skill

  • Get a snapshot (overview / observability_overview): firing-alert count, scrape targets up/down, rules erroring (Prometheus) or dashboard/datasource counts (Grafana)
  • Run PromQL (instant_query / range_query), enumerate label_values or series_metadata
  • Check scrape health (target_scrape_health, dropped_targets) and rule health (rule_health, list_rules)
  • Triage alerts: firing_alerts / pending_alerts, the Alertmanager view (alertmanager_alerts, list_silences), then firing_alert_rca to root-cause
  • Reduce alert noise (alert_noise_and_flap_analysis) → group_by / inhibition / longer for
  • Grafana: list_dashboards, get_dashboard, list_datasources, datasource_health, list_folders
  • Loki logs: enumerate loki_labels / loki_label_values, run a bounded loki_query (LogQL, stream selector required), loki_tail_errors for a selector; then log_error_burst_rca to root-cause an error burst and log_volume_analysis for volume/cardinality; alert_log_context to pull the logs behind a firing alert
  • Governed writes: silence an alert (create_silence, time-boxed), annotate an event (create_annotation), update/delete a dashboard (dry_run first for either), or hot-reload Prometheus (reload_prometheus_config)

Do NOT use when the target is not a Prometheus/Grafana observability stack — route hypervisor, storage, backup, container-orchestrator, network-device-config, or OT/industrial work to the appropriate other AIops-tools skill. Hosted/SaaS monitoring suites (Datadog, New Relic, enterprise NMS) are out of scope.

If the user wants…Use
Prometheus / Alertmanager / Grafana observability opsobservability-aiops (this skill)
A different platform (hypervisor, storage, backup, orchestrator, network config, OT edge)the appropriate other AIops-tools skill
Hosted/SaaS monitoring (Datadog, New Relic, enterprise NMS)out of scope for this tool

Common Workflows

The CLI covers the reads and the three RCAs (alert, query, logs, overview); the guarded writes (silences, annotations, dashboards, config reload) are MCP tools — those steps name the tool rather than a CLI command.

"Pager went off" — root-cause the firing alerts and time-box the noise

  1. observability-aiops overview → one-shot stack picture: firing counts, target health, rule health — is this one alert or the whole stack?
  2. observability-aiops alert firing → what is firing right now, grouped by severity
  3. observability-aiops alert rca → each firing alert joined to its rule expression with a likely cause and a recommended action (advisory heuristic — verify it, do not act on it blind)
  4. observability-aiops query instant '' → evaluate the alert's own expression yourself and confirm the RCA's reading of it
  5. observability-aiops query range '' --start --end --step 60s → see when it crossed the threshold, which usually names the change that caused it
  6. Time-box the noise while you fix the cause: the create_silence MCP tool on a specific matcher (a positive duration is required — silences cannot be open-ended), then observability-aiops alert silences to confirm it landed
  7. Failure branch: if the silence was too broad, expire_silence ends it immediately, or observability-aiops undo apply replays the recorded inverse (create_silence's undo is expire). If alert rca returns nothing while alerts are visibly firing, the alerts are coming from Alertmanager without a matching Prometheus rule — check alertmanager_alerts and list_rules rather than assuming the RCA is broken.

Investigate a scrape gap ("metrics went missing")

  1. observability-aiops overview → up/down target counts at a glance
  2. target_scrape_health → the unhealthy targets with their raw lastError
  3. target_scrape_health_analysis → down targets ranked, each lastError classified (connection refused / timeout / auth / DNS / TLS) with a concrete fix
  4. dropped_targets → if a target is missing entirely rather than down, it was relabeled away; this is where that shows up
  5. observability-aiops query instant 'up{job=""}' → confirm the gap in the metric itself, not just in the target page
  6. After fixing scrape config, reload_prometheus_config (a governed write) → then re-run target_scrape_health to confirm the target came back
  7. Failure branch: if reload_prometheus_config succeeds but the target is still down, the config on disk was not what you thought — check prometheus_config_status for what Prometheus actually loaded. A reload with a broken config is rejected by Prometheus and leaves the old config running, so a failed reload is not an outage.

Tame a noisy / flapping alert

  1. observability-aiops alert firing → the volume of what is firing
  2. alert_noise_and_flap_analysis → alertnames with many instances or exact duplicates, each with a group_by / inhibition / longer-for recommendation
  3. list_rules and rule_health → read the offending rule's current for duration and confirm it is evaluating cleanly
  4. observability-aiops query range '' --start --end --step 60s → see the flapping in the data and pick a for window that actually covers it
  5. create_silence for a time-boxed quiet period while the rule change ships; observability-aiops alert silences to confirm
  6. Failure branch: silencing is a stopgap, not a fix — if the silence expires and the flapping returns, the rule threshold or for window is still wrong. Use observability-aiops undo list to see exactly which silences this tool created, so no stale silence quietly hides a real outage.

Root-cause a log error burst (Loki, read-only)

  1. alert_log_context → the firing alert's labels mapped to a Loki stream selector plus the correlated error logs (or start from a selector directly)
  2. observability-aiops logs errors '{app="api"}' --hours 2 --limit 200 → tail the error-level lines for that stream
  3. log_error_burst_rca → per-stream error counts against a baseline window, each burst classified (new signature / volume spike / single instance)
  4. observability-aiops logs query '{app="api"} |= "timeout"' --hours 2 → confirm the specific signature the RCA named
  5. log_volume_analysis → the highest-volume streams and any high-cardinality label driving a stream/index explosion
  6. Failure branch: Loki here is read-only and bounded — queries require a stream selector and are capped by lookback and line count. A query rejected for a missing selector is the guard working, not a bug: narrow it with observability-aiops logs labels first. There is no write surface for Loki, so remediation happens in the emitting service, not through this tool.

Safely change or retire a Grafana dashboard (reversible)

  1. list_dashboards / list_folders → locate the dashboard and its folder
  2. get_dashboard → confirm this is the right dashboard before touching it
  3. update_dashboard with dry_run=True → preview; then for real — it fetches and stashes the prior model and records a restore undo
  4. To retire one: delete_dashboard with dry_run=True first. Delete is high risk — the prior model is captured before the delete so the undo can recreate it; set OBSERVABILITY_AUDIT_APPROVED_BY (and OBSERVABILITY_AUDIT_RATIONALE) if you want that recorded on the audit row
  5. create_annotation → mark the change on the timeline so the next responder can correlate a metric shift with this edit
  6. Failure branch: wrong dashboard or a bad edit — observability-aiops undo list then observability-aiops undo apply restores the captured prior model (or recreates a deleted dashboard from it). If the write fails outright, that is the connecting account's permissions (this tool does not gate it) — check the token's role before assuming observability-aiops doctor connectivity is at fault.

Governance & Safety

The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (give it a Grafana token with only Viewer scope, and a Prometheus/Alertmanager reached without the admin/write API — writes then fail at the server). There is no read-only switch, policy file, or approval gate.

  • Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to ~/.observability-aiops/audit.db (relocatable via OBSERVABILITY_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.
  • OBSERVABILITY_AUDIT_APPROVED_BY / OBSERVABILITY_AUDIT_RATIONALE are optional annotations recorded on the audit row (who/why); they are never required and never block.
  • Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with OBSERVABILITY_RUNAWAY_MAX=0.
  • Writes support --dry-run / dry_run=True and double confirmation at the CLI.
  • Silences are time-boxed (require a positive duration). Reversible writes capture the real fetched before-state and record an inverse descriptor (create_silence→expire, update/delete dashboard→restore/recreate).

References

  • references/capabilities.md — full tool + platform + API-path reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity
  • references/agent-guardrails.md — running this with a smaller / local model: what the harness enforces for you, and a ready-made system prompt for the rest

Questions people ask

What observability platforms does this support?
Prometheus (HTTP API + PromQL), Alertmanager, Grafana (dashboards, datasources, folders), and Grafana Loki for logs. All reads only — there is no write surface for Loki.
How are writes protected?
Every write tool carries a risk-tier label (med/high), writes to a local audit log under ~/.observability-aiops/, and records an undo token. Dashboard deletes capture the prior model before removal so it can be restored.
Can I run PromQL queries?
Yes — instant_query and range_query cover both instant and range vectors. label_values and series_metadata let you enumerate available label names and series without writing PromQL.
What does the RCA analysis actually do?
firing_alert_rca joins each firing alert to its rule expression and maps it to a likely cause and recommended action. It is an advisory heuristic — verify before acting. The alert-noise analysis finds duplicates and suggests group_by, inhibition, or longer-for windows.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

Measure whether a transformation changed your growth engine or just added a one-time bump.

by deciqai1 installs2 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

End-of-day options analytics ranked against each ticker's own history: IV rank, put/call percentile, skew, max pain, and unusually active contracts.

by thesentitrader2 installs2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs