Memory

container-host-aiops

Try it

Diagnose and operate a single Docker, Portainer, or Podman host — with audit logs, undo tokens, and guided write workflows.

What it does

38 MCP tools for governed single-host container operations. Covers container/image/volume/network inspection and system metrics across Docker Engine API, Portainer, and Podman. Three flagship analyses handle restart-loop RCA, resource-pressure hotspots, and image/volume bloat with reclaimable bytes. Write operations (restart/stop/start/remove containers, prune, update limits, recreate stacks) require --dry-run and capture undo tokens; all actions log to a local audit database. Supports offline analysis by passing exported data directly to the analysis tools. Portainer API tokens are stored encrypted.

When to use it

  • Container is crash-looping — find the root cause and act on it
  • Host is running out of disk — identify and preview prune candidates
  • Everything on the host is slow — rank containers by resource pressure
  • Recreate a drifting Portainer stack after a bad deploy

The skill document

Container Host AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by Docker, Inc., Portainer.io, or any container-platform vendor. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Container-Host-AIops under the MIT license.

Governed Docker + Portainer + Podman container-host operations — 38 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.container-host-aiops/ (MCP + CLI alike), a runaway/budget safety guard, and undo-token recording. It records every operation; whether a write is permitted is the agent's or the account's call, not the skill's. A Docker target speaks the Docker Engine API over a unix socket or TCP; a Portainer target speaks the Portainer API (and proxies Docker); a Podman target speaks over its rootful/rootless socket (Docker-compat + libpod). The Portainer API token is stored encrypted (~/.container-host-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk; a local Docker/Podman socket needs no secret.

Standalone: the governance harness is bundled in the package (container_host_aiops.governance) — container-host-aiops has no external skill-family dependency. Verification: Exercised against a live Docker Engine 27.5.1 daemon (doctor, overview, the three flagship analyses, and a governed stop_container with audit + undo recorded); the Portainer and Podman API paths are covered by the mock suite only. See docs/VERIFICATION.md.

What This Skill Does

DomainToolsCountRead or Write
Overviewone-shot host health11 read
Containerslist/inspect, logs, stats, top, restart summary66 read
Imageslist, inspect (+history), dangling, disk usage44 read
Volumeslist, inspect, dangling33 read
Networkslist, inspect22 read
Systeminfo, version, df, events44 read
Stacksendpoints, stacks, stack detail (Portainer), compose-stack rollup (docker+podman)44 read
Pods (Podman)list pods (libpod)11 read
Analyses (flagship)restart-loop RCA, resource pressure, image/volume bloat33 read
Writesremove container, prune images, prune volumes, recreate stack44 write (high)
restart, stop, start, update container44 write (medium)

The three analyses accept injected data for offline analysis, or pull live from a configured target. Portainer endpoints/stacks require a portainer target; list_compose_stacks works on docker or podman; list_pods requires a podman target.

Quick Install

uv tool install container-host-aiops
container-host-aiops init       # interactive wizard: Docker/Podman socket or Portainer target
container-host-aiops doctor

When to Use This Skill

  • Triage a host (overview): version + container state rollup + disk headline
  • Find crash-looping containers (analyze restart-loop / restart_loop_rca): ranked by restart count with a likely cause and action from the exit code, plus a log tail
  • Spot resource pressure (analyze resource-pressure / resource_pressure_analysis): CPU%/mem% vs each container's limits, worst first, with a recommendation
  • Reclaim disk (analyze bloat / image_and_volume_bloat): dangling images + volumes + build cache as prune candidates with reclaimable bytes
  • List/inspect containers, images, volumes, networks; tail logs; read stats/top
  • Restart/stop/start a container, update its resource limits (reversible), remove a container, prune images/volumes, or recreate a Portainer stack — all with dry-run + double-confirm

Do NOT use when the target is a cluster orchestrator, a hypervisor, a storage appliance, a backup product, network device config, or OT/industrial equipment.

If the user wants…Use
Docker / Portainer single-host container opscontainer-host-aiops (this skill)
A cluster orchestrator's workloads/rolloutsa cluster ops skill
Hypervisor VM lifecycle (power, snapshot, migrate)a hypervisor ops skill
OT / industrial edge (Modbus, OPC-UA, PLC)the industrial-aiops line

Common Workflows

1. A container is crash-looping

  1. container-host-aiops doctor → confirm the socket/endpoint is reachable before you trust any read.
  2. container-host-aiops analyze restart-loop → containers ranked by restart count, each with a likely cause read off the real exit code (137 OOM/SIGKILL, 143 SIGTERM, 139 segfault, 127 bad entrypoint, …), a recommended action, and a log tail.
  3. container-host-aiops container logs --tail 200 → read the actual crash output; container-host-aiops container inspect → confirm the exit code, restart policy, and configured limits the RCA cited.
  4. If the cause is memory: container-host-aiops analyze resource-pressure --mem 75 → see how close the container runs to its ceiling, then container-host-aiops manage update '{"Memory": 1073741824}' --dry-run and re-run without --dry-run (double-confirm; the write captures the prior limits as its undo descriptor).
  5. container-host-aiops manage restart → bring it up on the new limit, then re-run analyze restart-loop to confirm the loop stopped.
  6. Failure branch: if it still loops, the limit was not the cause — reverse the change with container-host-aiops undo list → undo apply (restores the prior limits, not a guess) and go back to step 3 with the fresh log tail. If the container will not stop at all, manage remove --force --dry-run first: force-remove is high-risk and irreversible, so read the dry-run before committing.

2. The host is out of disk

  1. container-host-aiops system df → where the space actually went (images vs containers vs volumes vs build cache).
  2. container-host-aiops analyze bloat → dangling images, dangling volumes, and build cache as ranked prune candidates with reclaimable bytes per item.
  3. container-host-aiops image dangling and container-host-aiops volume dangling → eyeball the concrete list before deleting anything. A "dangling" volume holding data you still want is the classic way this goes wrong.
  4. container-host-aiops manage prune-images --dry-run → exactly what would be removed; re-run without --dry-run (double-confirm, high risk).
  5. container-host-aiops manage prune-volumes --dry-run → read this one carefully; volume pruning destroys data and records no undo. Note that Docker's default prune removes only ANONYMOUS unused volumes — the preview reports the named unused ones it will not touch as alsoUnusedNamed*; add --all to include them. Only then re-run for real.
  6. container-host-aiops system df again → confirm the space came back.
  7. Failure branch: pruning is not reversible. If you removed a volume you needed, the undo store cannot help — restore from your backup. The dry-run in steps 4–5 is the only safety net, which is why both are separate confirm-gated steps.

3. "Everything on this box is slow"

  1. container-host-aiops overview → one-shot: platform/version, container counts by state, and the headline resource picture.
  2. container-host-aiops analyze resource-pressure --cpu 80 --mem 80 → running containers ranked against their own limits, each row citing the measured percentage rather than a verdict.
  3. container-host-aiops container stats and container-host-aiops container top → confirm the top offender at the process level before you act on it.
  4. container-host-aiops system events → correlate the pressure with what changed (a recent deploy, restart storm, or image pull).
  5. Act on the worst offender: manage update '{"NanoCpus": 2000000000}' to cap it (dry-run first, undo-recorded), or manage stop to shed it entirely.
  6. Failure branch: if capping the top container just moves the pressure elsewhere, the host is genuinely undersized rather than misconfigured — reverse your change with undo apply so you are not left with a half-applied limit, and take the sizing result to whoever owns capacity.

4. Stack drift after a bad deploy (Portainer)

  1. container-host-aiops stack endpoints → the endpoints this Portainer manages; container-host-aiops stack list → the stacks on the one you care about.
  2. container-host-aiops stack detail → the stack's current definition; container-host-aiops stack compose → the compose file it is running from.
  3. container-host-aiops container list --running and container-host-aiops container restarts → which of the stack's containers are actually unhealthy versus merely restarted.
  4. container-host-aiops manage recreate-stack --dry-run → preview the redeploy; re-run without --dry-run (double-confirm, high risk).
  5. Validate with container-host-aiops overview and analyze restart-loop.
  6. Failure branch: recreate-stack redeploys from the stack's stored definition — if that definition is itself the broken thing, recreating will faithfully reproduce the breakage. Fix the compose source in Portainer first, and use container-host-aiops undo list to check what the session already changed before layering another write on top.

Offline analysis (no live host)

Pass data straight to the analysis tools — restart_loop_rca(containers=[...]), resource_pressure_analysis(samples=[...]), or image_and_volume_bloat(dangling_images=..., dangling_volumes=..., df=...) — to analyse an exported dataset without connecting to a host.

Governance & Safety

The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (a read-only Docker socket, a Portainer account without write scope — writes then fail at the server). There is no read-only switch, policy file, or approval gate.

  • Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to ~/.container-host-aiops/audit.db (relocatable via CONTAINER_HOST_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.
  • CONTAINER_HOST_AUDIT_APPROVED_BY / CONTAINER_HOST_AUDIT_RATIONALE are optional annotations recorded on the audit row (who/why); they are never required and never block.
  • Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with CONTAINER_HOST_RUNAWAY_MAX=0.
  • Writes support --dry-run / dry_run=True and double confirmation at the CLI; prune previews list what would be removed + reclaimable bytes.
  • Mutating/reversible writes fetch the real before-state and record an inverse descriptor (stop→start, update_container→restore prior limits); irreversible ops record only the before-state.

References

  • references/capabilities.md — full tool + field reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity

Questions people ask

Does this work with Podman?
Yes. It speaks over the Podman rootful or rootless socket (Docker-compatible API + libpod). The list_pods tool requires a podman target. Podman API paths are covered by the mock test suite.
How does the safety mechanism work?
Every operation — read and write — is logged to ~/.container-host-aiops/audit.db. Write tools support --dry-run and double-confirm at the CLI. Reversible writes (restart/stop/start, update limits) record an undo descriptor; irreversible ops (remove, prune) record only the before-state. A runaway guard acts as a circuit breaker for repeated calls.
Does it support read-only connections?
There is no built-in read-only switch. Write operations will fail at the server if you connect through a read-only Docker socket or a Portainer account without write scope. The skill itself does not gate writes — it logs them.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Turn China 3C launch inputs into executable routes, messaging, channel actions, risk checks, and review decisions.

by killsnake0126 installs112 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs