Coding

endpoint-aiops

Try it

Diagnose managed endpoint fleets: login storms, health scores, drift detection, and guarded writes with full audit trail.

What it does

A fleet operations toolkit for centrally-managed endpoints (thin clients, VDI). Provides read-only fleet triage, composite health scoring, and two guarded write operations. Reads include fleet health overview, endpoint inventory, per-endpoint health score (0–100, worst first with cited deductions), login/boot session analysis, login-storm detection, and patch/config drift reports. Writes include profile assignment (reversible, captures prior state) and endpoint reboot (no undo). Every operation — reads and writes — is logged to a local SQLite audit file with parameters, result, duration, and risk tier. Analysis tools accept injected records for offline post-incident work without live server…

When to use it

  • Morning login slow — find the slowest boot/login contributors across the fleet
  • Rank endpoints by health score to prioritize remediation
  • Find which endpoints have drifted from the fleet baseline or are behind on patches
  • Assign a config profile to a drifted endpoint and revert if needed

The skill document

Endpoint AIops

Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by any endpoint-management vendor. Product and trademark names belong to their owners. Source at github.com/AIops-tools/Endpoint-AIops under the MIT license.

Governed managed-endpoint operations — 13 MCP tools, every one wrapped with the bundled @governed_tool harness: a local unified audit log under ~/.endpoint-aiops/, token/runaway budget guard, undo-token recording, and descriptive risk tiers. The management-server API key is stored encrypted (~/.endpoint-aiops/secrets.enc, Fernet + scrypt) — never plaintext on disk.

Standalone: the governance harness is bundled in the package (endpoint_aiops.governance) — endpoint-aiops has no external skill-family dependency. The test suite is mock-based; a live management server has not yet been exercised (see docs/VERIFICATION.md).

What This Skill Does

CategoryToolsCountRead or Write
Overviewfleet health overview11 read
Inventoryendpoint list, get, health score33 read
Sessionssession list, login-storm analysis22 read
Driftdrift report, patch status, patch compliance33 read
Remediationassign profile (high)11 write
reboot (medium)11 write

The analysis tools (login_storm_analysis, drift_report, patch_status, patch_compliance, endpoint_health_score) accept injected records for pure/offline analysis; endpoint_health_score and patch_compliance are injected-only, the others also pull live from a configured target.

Quick Install

uv tool install endpoint-aiops
endpoint-aiops init       # interactive wizard: connection + encrypted API key
endpoint-aiops doctor

When to Use This Skill

  • Triage a fleet (overview): online/offline counts, stale endpoints, agent/patch spread
  • Rank the fleet by risk (endpoint_health_score): a composite 0-100 per-endpoint score, worst first, with every deduction cited
  • Diagnose a morning login storm (session storm / login_storm_analysis) and find the slowest login/boot contributors
  • Find endpoints drifted from the fleet baseline (drift report) or behind on patches (drift patch)
  • Assign a config profile to an endpoint (reversible) or reboot one (dry-run + double-confirm)

Do NOT use when the target is OT/industrial equipment (use industrial-aiops), a hypervisor, a storage appliance, a backup product, a container cluster, or a network device.

If the user wants…Use
Managed-endpoint fleet: login storms, drift, profilesendpoint-aiops (this skill)
OT / industrial edge (Modbus, OPC-UA, PLC, PROFINET)the industrial-aiops line
Hypervisor VM lifecycle (power, snapshot, migrate)a hypervisor ops skill
Container/cluster lifecyclea cluster ops skill

Common Workflows

"Nobody can log in this morning" — diagnose the 9am login storm

  1. endpoint-aiops overview → is this fleet-wide (offline/stale counts spiking) or confined to logins?
  2. endpoint-aiops session storm --since-hours 12 --window-s 300 --min-concurrent 10 → storm episodes with peak concurrency and distinct users/endpoints, plus slowestByLogin / slowestByBoot
  3. endpoint-aiops session list --since-hours 12 → inspect the raw sessions behind a suspicious episode (confirm the timestamps, don't trust the summary alone)
  4. endpoint-aiops drift report → cross-check the laggards; a stray agent version or divergent profile is a common cause of slow logins
  5. Failure branch: if session storm reports no episodes but users still complain, widen the window (--window-s 900) and lower --min-concurrent before concluding there is no storm; if the CLI errors on connectivity, run endpoint-aiops doctor first — the analysis is only as good as the session feed.

Bring a drifted endpoint back to the fleet baseline (reversible)

  1. endpoint-aiops drift report → the drifted endpoints and exactly which fields deviate from the fleet-majority baseline
  2. endpoint-aiops endpoint get → confirm you are about to change the right device and note its current profile
  3. endpoint-aiops endpoint assign-profile --dry-run → preview the exact POST /endpoints//profile call, changes nothing
  4. endpoint-aiops endpoint assign-profile → double confirmation; high risk. The prior profile is captured and an inverse reassign undo descriptor is recorded
  5. Failure branch: if the endpoint misbehaves on the new profile, endpoint-aiops undo list then endpoint-aiops undo apply restores the captured prior profile (not a guess); re-run drift report to confirm the fleet picture.

Patch-compliance sweep before a maintenance window

  1. endpoint-aiops drift patch --target-patch 2024-06 → distribution of patch levels plus the endpoints behind the target
  2. endpoint-aiops endpoint list → resolve the behind-target ids to hostnames/owners for the change ticket
  3. endpoint-aiops overview → check how many of those are currently offline (an offline endpoint will not take the patch)
  4. Reboot a stuck endpoint that has staged its patch: endpoint-aiops endpoint reboot --dry-run, then without --dry-run (double confirmation)
  5. Failure branch: endpoint_reboot is medium risk and declares no undo — a reboot has no safe inverse. If the endpoint does not come back, the audit record in ~/.endpoint-aiops/audit.db holds its prior online state for the incident write-up; recovery is out-of-band (console/PXE), not via this tool.

Offline post-incident analysis (no live server)

  1. Export the incident's session and endpoint records from the management server into JSON
  2. Call the analysis tools with injected records — login_storm_analysis(sessions=[...]), drift_report(endpoints=[...]), patch_compliance(endpoints=[...]), endpoint_health_score(endpoints=[...]) — no connection or credentials required
  3. endpoint_health_score returns a composite 0-100 per endpoint, worst first, with every deduction cited — use it to rank the remediation queue
  4. Failure branch: if a tool rejects the injected records, the export is missing fields the analysis needs (e.g. session start/login-duration, or endpoint patch level) — re-export rather than hand-patching the data, so the numbers stay traceable to the source.

Governance & Safety

The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the account you connect it with (a management-console account or API token scoped to a read-only role — writes then fail at the server). There is no read-only switch, policy file, or approval gate.

  • Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to ~/.endpoint-aiops/audit.db (relocatable via ENDPOINT_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does.
  • ENDPOINT_AUDIT_APPROVED_BY / ENDPOINT_AUDIT_RATIONALE are optional annotations recorded on the audit row (who/why); they are never required and never block.
  • Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with ENDPOINT_RUNAWAY_MAX=0.
  • Writes support --dry-run / dry_run=True and double confirmation at the CLI.
  • Reversible writes fetch the real before-state and record an inverse descriptor (endpoint_assign_profile→restore prior profile); the reboot (no safe inverse) records only the before-state.

References

  • references/capabilities.md — full tool + field reference
  • references/cli-reference.md — CLI command reference
  • references/setup-guide.md — onboarding, credentials, and connectivity

Questions people ask

What does the health score mean?
The composite 0–100 score reflects patch currency, agent version, session lag, and drift status per endpoint. Every deduction is cited in the output so you can verify the reasoning. Endpoints rank worst-first.
Can I analyze login storms without a live connection to the management server?
Yes — login_storm_analysis, drift_report, patch_compliance, and endpoint_health_score all accept injected JSON records. Export the data from the management server, pass it in, and get the analysis offline. endpoint_health_score and patch_compliance are injected-only.
What happens if a write goes wrong?
Profile assignments capture the prior state and record an inverse descriptor — use 'undo list' and 'undo apply' to restore it. Reboots record only the before-state; there is no undo. Both writes support dry-run mode and double confirmation at the CLI. All operations, successful or failed, appear in the audit log.

Related skills

Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.

by zw0081 installs1 stars

Operate TaskTime Pro through MCP: manage tasks, track time, handle expenses, and prepare invoices from a paired browser session.

by tasktimepro1 installs1 stars

Escape the scarcity trap — diagnose bandwidth consumption and design protected slack to restore strategic capacity.

by deciqai1 installs2 stars

Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.

by johnpatternai22 installs8 stars

End-of-day options analytics ranked against each ticker's own history: IV rank, put/call percentile, skew, max pain, and unusually active contracts.

by thesentitrader2 installs2 stars

Turn China 3C launch inputs into executable routes, messaging, channel actions, risk checks, and review decisions.

by killsnake0126 installs112 stars

More from zw008

Browse all skills

Operate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.

by zw00878 installs1 stars

Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.

by zw00876 installs

Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.

by zw00853 installs

Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.

by zw00851 installs

Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.

by zw00851 installs

Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.

by zw00850 installs