Memory

ceph-aiops

Try it

Decode Ceph HEALTH_WARN/ERR into plain-language causes and actions; inspect and govern OSD, PG, pool, RBD, CephFS, and RGW operations with full audit trail.

What it does

37 MCP tools for operating and diagnosing Ceph clusters via the ceph-mgr Dashboard REST API. The flagship cluster_health tool translates raw HEALTH_WARN/ERR check codes into what they mean, why they happened, and what to do next. Read tools cover OSD tree/df/perf, PG summary/stuck/scrub, pool usable capacity, RBD images and snapshots, CephFS/MDS, RGW, monitors, slow ops, and capacity forecast. Write tools handle cluster flags, OSD reweight/mark-in/mark-out/purge, scrub triggers, pool size/quota/pg_num/autoscale, pool and RBD lifecycle — each wrapped with audit logging, dry-run, undo tokens, and a runaway circuit breaker. Works against vanilla ceph-mgr (cephadm, hypervisor-bundled, or MicroC…

When to use it

  • Ceph cluster shows HEALTH_WARN or HEALTH_ERR and you need to decode what it means and what to do
  • Inspect OSD utilization, performance, or tree to find the most-full or slowest drives
  • Investigate stuck PGs, overdue scrubs, slow operations, or capacity forecast
  • Drain, reweight, or purge an OSD with reversible operations and audit trail

The skill document

Ceph AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by the Ceph project or any storage vendor. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Ceph-AIops under the MIT license.

Governed Ceph operations via the ceph-mgr Dashboard REST API — 37 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.ceph-aiops/, token/runaway budget guard, undo-token recording, and descriptive risk tiers. The Dashboard password is stored encrypted (~/.ceph-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk. The flagship cluster_health turns raw HEALTH_WARN/ERR check codes into plain-language cause + suggested action.

Standalone: the governance harness is bundled in the package (ceph_aiops.governance) — ceph-aiops has no external skill-family dependency. Works against vanilla ceph-mgr (cephadm / hypervisor-bundled / MicroCeph); no croit, no Kubernetes.

What This Skill Does

GroupToolsCountRead or Write
Healthcluster_health (flagship RCA), cluster_status22 read
OSDosd_tree, osd_df, osd_perf33 read
cluster_flag_set, osd_reweight, osd_mark_in, osd_mark_out, osd_purge55 write
PGpg_summary, pg_dump_stuck, scrub_status33 read
trigger_scrub, trigger_deep_scrub22 write
Poolpool_ls, pool_df22 read
set_pool_quota, set_pool_pg_num, set_pool_autoscale, pool_create, set_pool_size, pool_delete66 write
RBDrbd_ls11 read
rbd_image_create, rbd_snapshot_create, rbd_image_delete, rbd_snapshot_delete44 write
CephFS / RGWcephfs_status, rgw_status22 read
Cluster-opsmon_status, mgr_status, slow_ops, capacity_forecast44 read
throttle_recovery11 write
Undoundo_list, undo_apply22 undo

Totals: 37 tools — 17 read, 18 write, 2 undo. The MCP server exposes all 37; the CLI is a convenience subset.

Quick Install

uv tool install ceph-aiops
ceph-aiops init       # interactive wizard: mgr host/port/username + encrypted Dashboard password
ceph-aiops doctor

When to Use This Skill

  • Decode a HEALTH_WARN/ERR state (cluster_health / health detail) — cause + action per active check (PG_DEGRADED, OSD_NEARFULL, SLOW_OPS, MON_DOWN, LARGE_OMAP_OBJECTS, …)
  • One-shot triage (overview): HEALTH status + active checks + OSD up/in counts
  • Inspect OSDs (osd_tree / osd_df most-full first / osd_perf slowest first), PGs (pg_summary / pg_dump_stuck / scrub_status), pools (pool_ls / pool_df usable capacity)
  • Investigate slow requests (slow_ops), MDS trimming lag (cephfs_status), RGW large-omap (rgw_status), mon quorum (mon_status), and days-to-nearfull (capacity_forecast)
  • Safely drain + purge an OSD, change pool size/quota, or throttle a slow rebalance (governed writes with dry-run + undo)

Do NOT use when the target is not Ceph (a hypervisor, another storage appliance, a backup product, a container cluster, or a network device). Route those to the appropriate other AIops-tools skill.

If the user wants…Use
Ceph: HEALTH_WARN RCA, OSD/PG/pool/RBD/CephFS/RGW, rebalance, slow opsceph-aiops (this skill)
Any non-Ceph target (hypervisor, other storage, backup, cluster, network)the appropriate other AIops-tools skill

Common Workflows

1. "The cluster went HEALTH_WARN overnight" — decode it (read-only)

  1. ceph-aiops doctor → confirm the mgr Dashboard is reachable and the JWT login works before trusting anything else
  2. ceph-aiops overview → HEALTH status, the list of active check codes, and OSD up/in counts in one shot
  3. ceph-aiops health detail (MCP: cluster_health) → each active check translated into what it means, the likely cause, and a suggested action
  4. Drill into the implicated resource: PG_DEGRADED → pg_dump_stuck; OSD_NEARFULL → ceph-aiops osd df (most-full first); SLOW_OPS → slow_ops; MON_DOWN → mon_status; LARGE_OMAP_OBJECTS → rgw_status
  5. Failure branch: if doctor fails on auth, the Dashboard password is wrong or the store is locked — re-run ceph-aiops secret set (or export CEPH_AIOPS_MASTER_PASSWORD for non-interactive use). If doctor fails on reachability, the mgr dashboard module is likely not enabled; no read is issued against an unauthenticated session.

2. Retire a failing OSD: drain, mark out, purge (governed)

  1. ceph-aiops health detail → confirm the OSD is genuinely the problem (e.g. OSD_SLOW_PING_TIME, repeated SLOW_OPS on one id) rather than a cluster-wide symptom
  2. ceph-aiops osd df → confirm the id, and that the remaining OSDs have room to absorb its data before you drain anything
  3. ceph-aiops osd reweight 0.0 → start a gradual drain; reversible, the prior CRUSH weight is captured as the undo descriptor
  4. ceph-aiops osd out --dry-run, then re-run without --dry-run → high risk, double confirmation, needs CEPH_AUDIT_APPROVED_BY
  5. Wait for ceph-aiops health status / pg_summary to show all PGs active+clean — do not purge while backfill is running
  6. ceph-aiops osd purge --dry-run, then re-run without --dry-run → high, irreversible
  7. Failure branch: if client I/O tanks during the drain, stop and reverse — ceph-aiops undo list then ceph-aiops undo apply restores the prior weight (and osd_mark_in reverses the mark-out). Purge has no undo, which is exactly why it comes last and after active+clean.

3. Recovery is starving client I/O

  1. ceph-aiops health detail → confirm the cluster is actually backfilling/recovering (PG_DEGRADED, PG_BACKFILL_FULL) rather than hitting a different bottleneck
  2. slow_ops → check whether client requests are genuinely being blocked, and by which OSDs
  3. throttle_recovery(max_backfills=1, recovery_max_active=1) → med risk, reversible; the prior osd_max_backfills / osd_recovery_max_active are captured as the undo descriptor
  4. Re-check slow_ops and pg_summary — recovery is slower but client latency should recover
  5. Once the cluster is quiet, raise the values back (or ceph-aiops undo apply to restore the exact prior settings)
  6. Failure branch: if throttling does not help, the bottleneck is not recovery — go back to osd_perf (slowest OSDs first) and mon_status; do not keep lowering the throttle, you will only extend the degraded window.

4. A pool is running out of usable capacity

  1. ceph-aiops overview → look for POOL_NEARFULL / OSD_NEARFULL among the active checks
  2. pool_df → per-pool usage with usable capacity = raw ÷ size (a size=3 pool reports a third of raw — this is where most "but the disks aren't full" confusion comes from)
  3. capacity_forecast → days-to-nearfull at the current fill rate, so you know whether this is a this-week problem or a this-quarter one
  4. Buy time reversibly first: set_pool_quota (med, undo → prior quota) or set_pool_autoscale (med, undo) to let pg_num track the new size
  5. Only if a replica change is genuinely the answer: set_pool_size --dry-run then the real call — high risk, because lowering size reduces durability and any change forces cluster-wide data movement
  6. Failure branch: if the resulting rebalance saturates the cluster, apply workflow 3 (throttle_recovery) rather than reverting the size mid-flight; if the size change itself was wrong, ceph-aiops undo apply replays the recorded prior value — but expect a second full rebalance.

Governance & Safety

The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (a ceph-mgr Dashboard account with a read-only role — writes then fail at the mgr). There is no read-only switch, policy file, or approval gate.

  • Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to ~/.ceph-aiops/audit.db (relocatable via CEPH_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.
  • CEPH_AUDIT_APPROVED_BY / CEPH_AUDIT_RATIONALE are optional annotations recorded on the audit row (who/why); they are never required and never block.
  • Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with CEPH_RUNAWAY_MAX=0.
  • Destructive writes support --dry-run / dry_run=True and double confirmation at the CLI.
  • Reversible writes fetch the real before-state and record an inverse descriptor (osd_reweight→restore prior weight, cluster_flag_set→toggle back); irreversible ops (osd_purge, pool_delete, RBD deletes) record only the before-state.

References

  • references/capabilities.md — full tool → API-path → returns reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity

Questions people ask

How is the Dashboard password stored?
Encrypted locally at ~/.ceph-aiops/secrets.enc using Fernet with scrypt-derived keys. No plaintext credentials ever touch disk.
Can I preview a write operation before running it?
Yes. Every write tool supports dry-run mode to show what would happen. Irreversible operations like osd_purge and pool_delete additionally require double confirmation at the CLI.
What does 'governed writes' mean in practice?
Every operation — read and write — is logged to ~/.ceph-aiops/audit.db with params, result, status, duration, and risk tier. Reversible writes capture a before-state descriptor so you can undo them; irreversible writes record only the before-state for reference. A runaway guard acts as a safety backstop against rapid repeated calls.
Does this work with Kubernetes or geht?
No. This targets vanilla ceph-mgr Dashboard — cephadm, hypervisor-bundled mgr, or MicroCeph. It is not designed for Kubernetes operators or croit.
Who maintains this and is it affiliated with the Ceph project?
Community-maintained open-source under MIT license at github.com/AIops-tools/Ceph-AIops. It is not affiliated with, endorsed by, or sponsored by the Ceph project or any storage vendor.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Measure whether a transformation changed your growth engine or just added a one-time bump.

by deciqai1 installs2 stars

Turn China 3C launch inputs into executable routes, messaging, channel actions, risk checks, and review decisions.

by killsnake0126 installs112 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

Replace vague hunches with calibrated probability estimates you can track and improve over time.

by deciqai2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs