Memory

queue-aiops

Try it

Redis and RabbitMQ diagnostics with memory/latency RCA, queue backlog triage, and governed write operations backed by audit and undo.

What it does

28 MCP tools covering Redis (RESP) and RabbitMQ (management HTTP API), every one wrapped in a @governed_tool harness that produces a local audit log, reversible writes with undo, and descriptive risk tiers. Read operations include: broker overview, memory posture (used vs maxmemory, eviction policy, fragmentation), SLOWLOG digest, SCAN-budgeted big-key sampling (never KEYS *), connected clients, and queue/backlog detail. Four flagship RCAs report measured numbers with a cause and action: memory pressure, latency (command patterns, fork/AOF stalls), queue backlog classification (no consumers / unacked pileup / rate deficit), and connection churn on both platforms. Write operations include co…

When to use it

  • Investigate cache memory pressure when approaching maxmemory — identifies eviction policy, fragmentation, or oversized keys
  • Chase Redis latency — digests SLOWLOG by command pattern, flags O(N) commands, fork stalls, blocked clients
  • Triage a growing RabbitMQ queue backlog — classifies per-queue cause and checks memory/disk watermark alarms
  • Safely change broker state — config changes and policies are reversible with undo; high-risk operations require dry-run and double confirmation

The skill document

Queue AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by the Redis or RabbitMQ projects or their respective owners. Redis and RabbitMQ are trademarks of their respective owners. Source at github.com/AIops-tools/Queue-AIops under the MIT license.

Governed broker operations — 28 MCP tools across redis (RESP client) and rabbitmq (management HTTP API), every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.queue-aiops/, policy engine, token/runaway budget guard, undo-token recording, and descriptive risk tiers. A per-target platform field selects the protocol shape, so one config can span a mixed estate. The redis password / rabbitmq management password is stored encrypted (~/.queue-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.

Standalone: the governance harness is bundled in the package (queue_aiops.governance) — no external skill-family dependency. Behaviour is covered by a mock-based test suite; docs/VERIFICATION.md is the checklist for a live run (both platforms are free/self-hostable, so a lab container is enough).

What This Skill Does

GroupToolsCountR/W
Overviewqueue_overview1read
redisredis_server_info, redis_memory_stats, redis_clients, redis_slowlog, redis_config_get, redis_keyspace, redis_big_keys7read
rabbitmqrabbitmq_overview, list_queues, queue_detail, list_connections, list_channels, list_policies, node_health7read
Flagship analysesredis_memory_pressure_rca, redis_latency_rca, rabbitmq_queue_backlog_rca, connection_churn_analysis4read
Writesredis_config_set, redis_kill_client, declare_queue, set_policy, delete_policy5write (med)
Writespurge_queue, delete_queue2write (high)
Undoundo_list, undo_apply2read / write

The four flagship analyses are transparent heuristics that report their numbers, never a black-box verdict: redis_memory_pressure_rca reads used-vs-maxmemory + eviction policy + fragmentation + the big-key sample into a cause + action; redis_latency_rca digests the SLOWLOG by command pattern and adds fork/AOF stall signals; rabbitmq_queue_backlog_rca classifies each deep queue (no consumers / unacked pileup / rate deficit) and reports watermark alarms that block all publishers; connection_churn_analysis works on both platforms and pins churn to client sources.

Quick Install

uv tool install queue-aiops
queue-aiops init       # wizard: pick platform (redis/rabbitmq) + encrypted secret
queue-aiops doctor

When to Use This Skill

  • Get a one-shot snapshot (overview / redis_server_info / rabbitmq_overview)
  • Investigate cache memory pressure (analyze memory) → cause + action (raise maxmemory vs fix eviction policy vs split big keys)
  • Chase latency (analyze latency, redis slowlog) → O(N) command patterns, blocked clients, fork/AOF stalls
  • Triage a growing queue (analyze backlog, rabbitmq queues) → per-queue cause (no consumers / unacked pileup / slow consumers) + watermark alarms
  • Spot connection churn or leaks (analyze churn, redis clients, rabbitmq connections) — clients grouped by source
  • Safely change state: redis config-set (undo = prior value), rabbitmq set-policy/delete-policy (undo = prior policy), declare-queue, and the high-risk purge/delete-queue (dry-run + double confirmation; messages are not restorable)

Do NOT use when the target is not a redis/rabbitmq broker — route hypervisor, storage, backup, cluster/orchestration, database, network, monitoring-stack, or OT/industrial work to the appropriate other AIops-tools skill.

If the user wants…Use
redis / rabbitmq cache & broker opsqueue-aiops (this skill)
A non-broker platform (hypervisor, storage, backup, cluster, database, network, monitoring stack, OT edge)the appropriate other AIops-tools skill
Managed cloud queue services / other broker productsout of scope for this tool

Common Workflows

1. Redis is near maxmemory and starting to evict

  1. queue-aiops doctor → confirm the broker is reachable and the credential (if any) works before you read numbers off it.
  2. queue-aiops redis memory → used vs maxmemory, the eviction policy in force, and the fragmentation ratio, straight from INFO memory.
  3. queue-aiops analyze memory --used-pct 85 → ranked findings, each citing its measured number: noeviction near the limit (the dangerous one — writes will start failing with OOM rather than evicting), active eviction in progress, fragmentation versus real swapping, and oversized keys.
  4. queue-aiops redis bigkeys --count 500 → a SCAN-budgeted sample of the largest keys, reported with its coverage % so you know how much of the keyspace was actually examined. A low coverage number means "no big key found" is not yet an answer.
  5. queue-aiops redis keyspace → which database the growth is in, to point the fix at the right workload.
  6. If the policy is the problem: queue-aiops redis config-get maxmemory* to read the current values, then queue-aiops redis config-set maxmemory-policy allkeys-lru --dry-run and re-run for real (double-confirm; the prior value is captured from CONFIG GET as the undo descriptor).
  7. Failure branch: switching a cache from noeviction to an eviction policy means Redis will start deleting data — if that instance is being used as a datastore rather than a cache, this is the wrong fix and you need memory instead. Reverse it immediately with queue-aiops undo list → undo apply , which restores the prior policy rather than a default. Note config-set changes the running config only; if it must survive a restart, persist it in the config file too — the undo store cannot help you with a value the broker forgot on its own.

2. Redis got slow

  1. queue-aiops redis slowlog --limit 50 → the slowest entries the broker itself recorded.
  2. queue-aiops analyze latency --slow-us 10000 → the slowlog digested by command pattern, with O(N) commands flagged and an incremental variant suggested (KEYS → SCAN, SMEMBERS → SSCAN), plus blocked-client counts and fork/AOF stall signals read out of INFO persistence.
  3. queue-aiops redis info → confirm whether the stalls line up with background saves (a fork stall is a persistence problem, not a query problem).
  4. queue-aiops redis clients → who is connected, and how many are blocked; queue-aiops analyze churn → whether clients are reconnecting constantly, which shows up as latency but is really a client-configuration bug.
  5. If one client is pathological: queue-aiops redis kill-client --addr --dry-run then for real (double-confirm).
  6. Failure branch: kill-client is irreversible and records no undo — a killed connection cannot be un-killed, and a well-behaved client will simply reconnect, which fixes nothing while your kill is audited as a write. If latency does not improve, the cause is more likely the O(N) command pattern from step 2: fix the caller, not the connection. Killing clients in a loop will trip the runaway budget guard.

3. A RabbitMQ queue keeps growing

  1. queue-aiops rabbitmq overview and queue-aiops rabbitmq nodes → check first for memory or disk watermark alarms, which block every publisher on the node and make every queue look broken at once.
  2. queue-aiops rabbitmq queues --vhost / → queues sorted deepest-backlog-first.
  3. queue-aiops analyze backlog --vhost / --top 20 → the per-queue cause: no consumers attached, consumers connected but not acking (an unacked pile-up), or a publish rate simply outpacing delivery — each citing the measured counts.
  4. queue-aiops rabbitmq queue → that queue's detail: consumer count, ready vs unacked split, and its arguments.
  5. queue-aiops rabbitmq connections and queue-aiops rabbitmq channels → confirm whether the consumers exist at all, and whether their prefetch is starving throughput.
  6. To cap unbounded growth while the consumer is fixed: queue-aiops rabbitmq set-policy backlog-cap '^orders\.' '{"max-length": 100000}' --apply-to queues --dry-run, then re-run for real (reversible — the prior policy is captured for undo).
  7. Failure branch: do not reach for rabbitmq purge as a first response — it is irreversible, risk=high, and destroys real messages; if the cause from step 3 was "no consumers", those messages are the backlog your consumers still need. Purge only with explicit sign-off, a --dry-run read first, and the CLI double confirmation. A max-length policy also drops messages once the cap is hit — if that is not acceptable, queue-aiops undo apply restores the prior policy and the real fix is consumer capacity.

4. Retire a queue, reversibly

  1. queue-aiops overview → the broker-wide picture across configured targets.
  2. queue-aiops rabbitmq queue --vhost / → confirm it is genuinely idle: zero consumers, zero ready, zero unacked. A queue with messages is not a queue you retire.
  3. queue-aiops rabbitmq policies → check no policy still targets its name pattern, so you are not leaving a dangling rule behind.
  4. queue-aiops rabbitmq delete-queue --vhost / --dry-run → preview.
  5. Re-run without --dry-run (double-confirm, risk=high) — the write captures the queue's definition first, so the undo descriptor re-declares exactly that queue (durability and auto-delete flags included).
  6. queue-aiops rabbitmq queues → confirm it is gone and nothing else changed.
  7. Failure branch: queue-aiops undo apply re-declares the queue from the captured definition — but it restores the queue, not its messages, which are gone with it. Bindings created outside this tool are not captured either. If the queue turns out to have been in use, expect to re-create bindings by hand; that asymmetry is why step 2 (proving it is idle) matters more than the undo does.

Governance & Safety

The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (a Redis ACL user restricted to read commands, a RabbitMQ management user with only the monitoring tag — writes then fail at the broker). There is no read-only switch, policy file, or approval gate.

  • Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to ~/.queue-aiops/audit.db (relocatable via QUEUE_AIOPS_HOME): params (secrets redacted), result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.
  • QUEUE_AUDIT_APPROVED_BY / QUEUE_AUDIT_RATIONALE are optional annotations recorded on the audit row (who/why); they are never required and never block.
  • Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with QUEUE_RUNAWAY_MAX=0.
  • Writes support --dry-run / dry_run=True and double confirmation at the CLI; CLI writes execute through the governed twins, so they are audited too.
  • Reversible writes capture the real fetched before-state and record an inverse descriptor; purge_queue and redis_kill_client are irreversible and record priorState only. delete_queue's undo restores the queue definition, never its messages — the descriptor says so.
  • Big-key sampling is SCAN-budgeted (never KEYS *); the redis surface is a typed command allow-list; rabbitmq paths are centrally percent-encoded (default vhost / included).

References

  • references/capabilities.md — full tool + platform + API/command reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity

Questions people ask

What guardrails apply to write operations?
All writes are audited to a local SQLite log. Reversible writes (config-set, policy changes, queue declaration) capture the prior state as an undo descriptor. Irreversible writes (purge_queue, kill_client) record priorState only and are flagged risk=high. The CLI requires double confirmation and dry-run support exists for every write.
How does the memory pressure RCA work?
redis_memory_pressure_rca reads used-vs-maxmemory, eviction policy, fragmentation, and the big-key sample. It reports each finding with its measured number, a cause statement, and a recommended action — not a black-box verdict.
Does this replace Prometheus/Datadog monitoring?
No. This skill provides on-demand diagnostics and governed write operations. It does not run continuous collection, alerting, or historical trending. Audit logs are local and operational, not metrics dashboards.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

End-of-day options analytics ranked against each ticker's own history: IV rank, put/call percentile, skew, max pain, and unusually active contracts.

by thesentitrader2 installs2 stars

Diagnose which mental domain is holding you back before choosing a cognitive intervention.

by deciqai1 installs3 stars

Prioritize growth directions with a 2×2 risk framework — pick one bet and commit.

by deciqai2 installs2 stars

Make irreversible life decisions by projecting to 80 and naming which regret you'd rather live with.

by deciqai1 installs2 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs