Memory

minio-aiops

Try it

Operate and diagnose MinIO clusters — 48 governed tools with audit logs, undo support, and plain-language root-cause analysis.

What it does

A command-line and MCP-driven toolkit for MinIO object storage operations. It wraps 48 tools across health checks, capacity analysis, bucket auditing, lifecycle management, healing diagnostics, object lock (WORM), and IAM — all recorded in a local audit log with reversible writes. Six flagship analyses turn raw MinIO state into ranked findings with suggested actions: why storage is filling up, which buckets are publicly exposed, where lifecycle policies are not reclaiming space, how close the cluster is to write-quorum failure, whether retention gaps exist, and what IAM users are over-privileged. Secret keys are stored encrypted (Fernet + scrypt); the skill records every call but does not g…

When to use it

  • Cluster reports storage full or writes refusing — run capacity RCA then target the biggest buckets
  • Auditor asks whether any buckets allow anonymous access or have hygiene gaps
  • Lifecycle rules exist but space is not being reclaimed — check noncurrent versions and incomplete uploads
  • A drive failed — assess erasure-set write quorum and heal backlog before scheduling maintenance

The skill document

MinIO AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by MinIO, Inc. or any storage vendor. Product and trademark names belong to their owners. Source at github.com/AIops-tools/MinIO-AIops under the MIT license.

Governed MinIO object-storage operations — 48 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.minio-aiops/, a token/runaway budget guard, undo-token recording, and a descriptive risk tier on every audit row. The secret key is stored encrypted (~/.minio-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk. Six flagship analyses turn raw state into plain-language cause + suggested action: capacity_rca, bucket_exposure_audit, lifecycle_gap_analysis, healing_health, diagnose_retention_gaps, diagnose_iam_exposure.

Standalone: the governance harness is bundled in the package (minio_aiops.governance) — minio-aiops has no external skill-family dependency. Verification status and the live-run checklist are in docs/VERIFICATION.md.

Authorization is not this tool's job: whether a write is permitted is the agent's judgement or the permission of the access key you connect with (a read-only IAM policy makes writes fail at the server). There is no read-only switch, policy file, or approval gate — the guarantee is that every call is audited. Driving this with a smaller / local model? See references/agent-guardrails.md.

What This Skill Does

GroupToolsCountRead or Write
Healthhealth_live, health_ready, health_cluster, cluster_status, fleet_overview55 read
Capacitycapacity_rca (flagship), usage_by_bucket22 read
Healinghealing_health (flagship), drive_status, node_status33 read
Exposure / ILMbucket_exposure_audit (flagship), lifecycle_gap_analysis (flagship)22 read
Bucketsbucket_ls, bucket_info, bucket_policy_get, bucket_lifecycle_get, bucket_versioning_get, bucket_quota_get, object_ls, incomplete_uploads_ls, server_info99 read
Writesset_bucket_policy, delete_bucket_policy, set_versioning, set_lifecycle, delete_lifecycle, set_bucket_quota, bucket_delete, remove_incomplete_uploads88 write
Object lock (WORM)bucket_lock_config, object_lock_status, diagnose_retention_gaps (flagship)33 read
bucket_create, set_default_retention, clear_default_retention, set_legal_hold, set_object_retention55 write
IAMiam_users, iam_groups, iam_policies, diagnose_iam_exposure (flagship)44 read
create_user, set_user_status, remove_user, attach_user_policy, detach_user_policy55 write
Undoundo_list, undo_apply2read + replay

Totals: 48 tools — 28 read, 18 write, 2 undo. The MCP server exposes all 48; the CLI is a convenience subset.

Quick Install

uv tool install minio-aiops
minio-aiops init       # interactive wizard: endpoint + access key + encrypted secret key
minio-aiops doctor

When to Use This Skill

  • "Storage is filling up / writes are failing" → capacity_rca (capacity vs used, offline drives/nodes, hotspots — cause + action per finding), then usage_by_bucket for the biggest consumers
  • "Is anything exposed?" → bucket_exposure_audit (ranked: public read/write policies, missing encryption, versioning off, no lifecycle)
  • "Where did my space go?" → lifecycle_gap_analysis (unbounded noncurrent versions, incomplete multipart uploads, reclaimable estimate) then remove_incomplete_uploads + set_lifecycle to fix it for good
  • "How many more drives can fail?" → healing_health (per-erasure-set online drives vs write quorum, heal backlog/errors)
  • One-shot triage (fleet_overview / minio-aiops overview): health + capacity headline + exposure headline
  • Per-bucket questions: bucket_info (policy/versioning/lifecycle/encryption/quota/tags in one answer)
  • Safely change policy/versioning/lifecycle/quota (reversible, undo recorded) or delete an empty bucket (governed, dry-run + double confirm)

Do NOT use when the target is not MinIO — for Ceph/RGW use ceph-aiops; for TrueNAS use truenas-aiops; for a hypervisor, backup product, container cluster, or network device route to the appropriate other AIops-tools skill.

If the user wants…Use
MinIO: capacity RCA, bucket exposure, ILM gaps, healing, bucket writesminio-aiops (this skill)
Ceph/RGW storageceph-aiops
TrueNAS storage appliancestruenas-aiops
Any other target (hypervisor, backup, cluster, network)the appropriate other AIops-tools skill

Common Workflows

Each recipe starts from an RCA read and ends in a governed, reversible write. Every write step accepts --dry-run; irreversible ones also double-confirm.

1. "Backups are failing — the cluster says it's out of space"

  1. minio-aiops overview → one-shot triage: health + capacity headline + exposure headline.
  2. minio-aiops capacity rca → ranked findings with cause + action (CLUSTER_NEARFULL, DRIVES_OFFLINE, DRIVE_HOTSPOT, …). Note which finding is actually driving the fill.
  3. minio-aiops capacity usage → the biggest buckets, largest first, so you know where the bytes live.
  4. minio-aiops bucket ilm-gap → how much of that is reclaimable: unbounded noncurrent versions and abandoned multipart uploads, with an estimate.
  5. Reclaim the abandoned uploads: minio-aiops bucket purge-uploads --older-than-days 7 --dry-run, then re-run without --dry-run (double confirm — this one is irreversible, only uploads older than the window are aborted).
  6. Cap version growth: minio-aiops bucket lifecycle-set --noncurrent-days 30 (reversible — the prior lifecycle config is captured). Abandoned uploads have no server-side rule — MinIO does not honour a lifecycle abort-incomplete rule — so re-run purge-uploads periodically instead.
  7. minio-aiops capacity rca again to confirm the finding cleared.

Failure branch: if capacity rca reports DRIVES_OFFLINE rather than genuine data growth, stop — do not delete anything. The space is not gone, it is unavailable. Go to recipe 3 and restore drive/erasure-set health first; purging data under a degraded erasure set removes redundancy you may need.

2. "Someone says one of our buckets is readable from the internet"

  1. minio-aiops bucket audit → ranked exposure findings, riskiest first (PUBLIC_WRITE_POLICY, PUBLIC_READ_POLICY, missing encryption, versioning off, no lifecycle).
  2. minio-aiops bucket info → the full per-bucket picture (policy, versioning, lifecycle, encryption, quota, tags) so you fix the right thing.
  3. Decide the fix. To remove anonymous access entirely: minio-aiops bucket policy-set --file restricted-policy.json --dry-run, then re-run for real. To drop the policy altogether, use the delete_bucket_policy MCP tool. Both are reversible — the prior policy JSON is captured.
  4. minio-aiops undo list → confirm an undo token was recorded for the change you just made.
  5. minio-aiops bucket audit → confirm the finding is gone.

Failure branch: if the write is refused by the server with an access-denied error, the access key you connected with lacks permission for that operation — the tool does not gate the write, the key does. Connect with a key whose IAM policy allows it (or ask whoever owns the key). If the new policy breaks a legitimate consumer, minio-aiops undo apply restores the exact prior policy document.

3. "A drive died — how many more failures can we take?"

  1. minio-aiops health check and minio-aiops health status → is the cluster serving reads and writes at all right now?
  2. minio-aiops heal status → per erasure set: online drives vs write quorum, failureToleranceRemaining, healing drives, heal backlog and errors.
  3. minio-aiops heal drives → which specific drives are offline or healing.
  4. minio-aiops heal nodes → whether the failures cluster on one node (a node problem, not a drive problem).
  5. If WRITE_QUORUM_AT_EDGE appears, the next failure stops writes: replace drives before any maintenance, and re-run heal status until the backlog drains.

Failure branch: if the heal backlog is not shrinking between runs, do not start more maintenance. Check heal nodes for an offline node first — a down node makes its drives look like many simultaneous drive failures, and replacing hardware will not fix it.

4. "The auditor asked whether our retained data is actually protected"

  1. minio-aiops lock gaps → ranked WORM findings across every bucket. The two that matter most: LOCK_ENABLED_NO_DEFAULT_RETENTION (the bucket advertises object lock but retains nothing unless each upload asks) and LIFECYCLE_CANNOT_EXPIRE_UNDER_RETENTION (an expiry rule that retention outlives, so the capacity never returns — both day counts are in the finding).
  2. minio-aiops lock config → whether object lock is enabled at all. objectLockEnabled: false is terminal: S3 accepts the flag only at bucket creation, so the answer is a new bucket plus a migration, not a setting.
  3. minio-aiops lock status → for a specific object: retention mode, days remaining, legal hold, and protection.versionDestroyable with what is blocking it. Note deleteMarkerStillPossible: object lock protects the bytes, not the key's visibility.
  4. Close the gap for future uploads: minio-aiops lock default-set GOVERNANCE --days 365 --dry-run, then for real (reversible — the prior rule is captured).
  5. minio-aiops lock gaps again → confirm the finding cleared, and check whether step 4 introduced the lifecycle contradiction from step 1.

Failure branch: lock default-set refused with "does not have object lock enabled" means the bucket can never be made WORM in place — create one with minio-aiops lock bucket-create --object-lock and migrate. If the auditor requires retention nobody can lift, that is COMPLIANCE, not GOVERNANCE — and it is genuinely permanent: verified on a live server, root with --bypass could not clear it, downgrade it, or delete the version. set_object_retention therefore records no undo token, refuses any call that would shorten retention already in force, and requires acknowledge_irreversible=True. Use GOVERNANCE unless permanence is the actual requirement.

5. "Decommission a retired bucket"

  1. minio-aiops bucket ls → confirm the exact bucket name.
  2. minio-aiops bucket info → verify it is genuinely retired (check versioning, lifecycle and quota, not just the object count).
  3. minio-aiops bucket uploads → surface incomplete multipart uploads, which keep a bucket non-empty even when it looks empty.
  4. minio-aiops bucket purge-uploads --older-than-days 7 if any remain (dry-run first, double confirm).
  5. minio-aiops bucket delete --dry-run → shows the API call, changes nothing.
  6. Re-run without --dry-run: double confirm, high risk. Execution re-checks emptiness (versions and delete markers included) and refuses otherwise — this tool never mass-deletes data.

Failure branch: if the delete is refused as non-empty, that is the guard doing its job — the bucket still holds objects, noncurrent versions, delete markers, or incomplete uploads. Go back to step 3, and never work around the guard by deleting data out-of-band; this tool deliberately has no mass-delete path.

Governance & Safety

  • The skill delivers reads and writes and records them; it does not decide whether a write is permitted — that is the agent's judgement or the permission of the access key you connect with (a read-only IAM policy makes writes fail at the server). There is no read-only switch, policy file, or approval gate.
  • Audit is the guarantee: every tool — MCP and CLI alike — is audited to ~/.minio-aiops/audit.db (relocatable via MINIO_AIOPS_HOME). MINIO_AUDIT_APPROVED_BY / MINIO_AUDIT_RATIONALE are optional annotations recorded when set, never required.
  • The declared risk_level (bucket_delete is high) is carried into the audit row as a descriptive tier — a label for the reviewer, not a gate.
  • Destructive writes support --dry-run and double confirmation at the CLI.
  • Reversible writes record an inverse descriptor capturing the real prior state (policy JSON, lifecycle XML, versioning state, quota).

References

  • references/capabilities.md — full tool → API-surface → returns reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity

Questions people ask

If I accidentally break a bucket policy, can I undo it?
Yes. Every governed write records an undo token in ~/.minio-aiops/. The prior state (policy JSON, lifecycle config, versioning, quota) is captured before any change. Use undo_apply to restore it. The exception is object lock in COMPLIANCE mode — it is genuinely permanent and records no undo token.
How are credentials stored?
Secret keys are encrypted at rest in ~/.minio-aiops/secrets.enc using Fernet symmetric encryption with a scrypt-derived key. Plaintext credentials never touch the disk.
Can I use this with a read-only access key?
The tool itself does not have a read-only switch — it records every call. A read-only IAM policy on your access key will cause writes to fail at the MinIO server side with an access-denied error. This is the intended guard when operating from a less-privileged key.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Make irreversible life decisions by projecting to 80 and naming which regret you'd rather live with.

by deciqai1 installs2 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Turn China 3C launch inputs into executable routes, messaging, channel actions, risk checks, and review decisions.

by killsnake0126 installs112 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs