Operate Kubernetes clusters with 55 audited tools — list resources, diagnose pod health, scale workloads, and manage rollouts safely.
Memory
cicd-aiops
Try itInspect and operate self-managed GitLab or Gitea CI/CD — pipelines, runners, artifacts, and root-cause analysis with governance and audit.
What it does
A standalone skill with 28 MCP tools covering server overview, project listing, pipelines with job trace tails, runner fleet status, merge/pull requests, branches, protection rules, releases, artifact inventory, and four RCA analyses: pipeline failure classification, runner health/saturation, artifact storage bloat, and stale work audit. Write operations (retry/cancel pipeline, pause/resume runner, update branch protection, delete artifacts) are wrapped with a governance harness that enforces audit logging, encrypted token storage, runaway circuit breaker, risk tiers, undo recording, and dry-run confirmation. Targets self-managed GitLab (/api/v4) and self-hosted Gitea (/api/v1) with a unifi…
When to use it
- Self-managed GitLab or Gitea instance is red — investigate which job failed and why
- Jobs stuck in queue — check runner status and per-tag saturation
- CI server out of disk — rank projects by artifact storage and identify reclaimable bytes
- Release freeze approaching — audit stale MRs, idle branches, and protection gaps
The skill document
CICD AIops
Disclaimer: Community-maintained open-source project, not affiliated with, endorsed by, or sponsored by GitLab Inc. or the Gitea project. GitLab and Gitea are trademarks of their respective owners. Source at github.com/AIops-tools/CICD-AIops under the MIT license.
Governed CI/CD operations — 28 MCP tools across self-managed GitLab (REST
/api/v4/...) and self-hosted Gitea (API /api/v1/...), every one wrapped with
the bundled @governed_tool harness: a local unified audit log under
~/.cicd-aiops/, token/runaway budget guard, undo-token recording, and
descriptive risk tiers. A per-target platform field selects the API shape, so
the same tools work on both servers and one config can span a mixed estate. The
access token is stored encrypted (~/.cicd-aiops/secrets.enc, Fernet +
scrypt) — never plaintext on disk.
Standalone: the governance harness is bundled in the package (
cicd_aiops.governance) — no external skill-family dependency. Both platforms are free/self-hostable, so a home lab is the cheapest live check; verification status and the checklist are indocs/VERIFICATION.md.
What This Skill Does
| Group | Tools | Count | R/W |
|---|---|---|---|
| Server | server_version, current_user, cicd_overview | 3 | read |
| Projects | list_projects, project_detail | 2 | read |
| Pipelines | list_pipelines, pipeline_detail, pipeline_jobs, job_trace_tail | 4 | read |
| Runners | list_runners, runner_detail | 2 | read |
| Repo surface | list_merge_requests, list_branches, list_protected_branches, list_releases | 4 | read |
| Artifacts | list_artifacts | 1 | read |
| Flagship analyses | pipeline_failure_rca, runner_health_rca, artifact_storage_bloat_analysis, stale_work_audit | 4 | read |
| Writes | retry_pipeline, cancel_pipeline, pause_runner, resume_runner, update_branch_protection | 5 | write (med) |
| Writes | delete_artifacts | 1 | write (high) |
| Undo | undo_list, undo_apply | 2 | read + replay |
The four flagship analyses are transparent heuristics that report their numbers,
never a black-box verdict: pipeline_failure_rca classifies each failed job from
its failure_reason + trace-tail markers (test-failure / dependency-network /
runner-timeout / oom / script-error) with matched evidence; runner_health_rca
flags offline/stale/paused runners, long-queued jobs, and per-tag saturation;
artifact_storage_bloat_analysis ranks projects by repo + artifact bytes and
estimates reclaimable bytes; stale_work_audit flags idle MRs/branches and
protection gaps.
Quick Install
uv tool install cicd-aiops
cicd-aiops init # wizard: pick platform (gitlab/gitea) + base URL + encrypted token
cicd-aiops doctor # version endpoint + token-scope probe per target
When to Use This Skill
- Get a one-shot snapshot (
overview/server_version/current_user) - Answer "why is CI red?" (
rca pipelines/pipeline_failure_rca) → cause + action per failed pipeline, with the trace evidence that drove the call - Find wedged capacity (
rca runners/runner_health_rca) → offline/stale runners, long-queued jobs, saturated tags - Reclaim disk (
rca storage/artifact_storage_bloat_analysis) → ranked projects + reclaimable bytes, feedingdelete_artifacts --dry-run - Repo hygiene (
rca stale/stale_work_audit) → idle MRs/branches, unprotected default branch, force-push gaps - Safely act:
retry_pipeline/cancel_pipeline,pause_runner/resume_runner(undo pair),update_branch_protection(undo replays prior settings),delete_artifacts(risk=high, dry-run + double confirm)
Do NOT use when the target is not a GitLab/Gitea CI/CD server — route hypervisor, storage, backup, database, network, or OT/industrial work to the appropriate other AIops-tools skill.
Related Skills — Skill Routing
| If the user wants… | Use |
|---|---|
| Self-managed GitLab / Gitea CI/CD ops | cicd-aiops (this skill) |
| Kubernetes deploy state (what the cluster is actually running) | k8s-aiops |
| A non-CI/CD platform (hypervisor, storage, backup, database, network, OT edge) | the appropriate other AIops-tools skill |
| GitLab.com / Gitea Cloud SaaS accounts | out of scope for this tool |
Common Workflows
Each recipe starts from one of the four RCAs and ends in a governed write.
Every CLI write accepts --dry-run and otherwise double-confirms. Note that
runner administration and pipeline retry/cancel are GitLab-only — on a
Gitea target those tools raise a teaching error listing what is available.
1. "The nightly pipeline has been red for three days"
cicd-aiops pipelines list dev/api --status failed -n 20→ the recent failed pipelines, newest first.cicd-aiops rca pipelines dev/api→ each failed job classified from itsfailure_reasonplus trace-tail markers: test-failure, dependency-network, runner-timeout, OOM, or script-error, with the evidence and a suggested action.cicd-aiops pipelines jobs dev/api→ which stage and job the classification came from.cicd-aiops pipelines trace dev/api -n 120→ the actual log tail, so you confirm the classification instead of trusting it.- Fix the cause. If the RCA said the failure was transient (dependency-network
or runner-timeout):
cicd-aiops pipelines retry dev/api --dry-run, then re-run for real (double confirm). cicd-aiops rca pipelines dev/apiagain to confirm the class of failure is gone rather than merely quieter.
Failure branch: if the RCA classifies the failures as test-failure or script-error, do not retry — the code is broken and a retry burns runner minutes to reach the same red. Retry is only honest for transient classes. If the retry itself fails to submit on a Gitea target, that is the teaching error: retry/cancel have no Gitea API v1 equivalent, so re-run the job from the Gitea UI instead.
2. "Jobs are sitting in the queue and nothing is picking them up"
cicd-aiops runners list --status offlineand--status paused→ the obvious suspects first.cicd-aiops rca runners→ stale contact ages, long-queued jobs, and which tag is saturated (queued jobs versus online runners carrying that tag).cicd-aiops runners show→ the specific runner's tags, last contact and status.- If a needed runner was paused:
cicd-aiops runners resume --dry-run, then for real (reversible — an inversepause_runneris recorded). - If a wedged runner is grabbing jobs and failing them:
cicd-aiops runners pauseto take it out of rotation (reversible — inverseresume_runnerrecorded). cicd-aiops rca runnersagain to confirm the queue is draining.
Failure branch: if the RCA shows a saturated tag rather than a
down runner, resuming runners will not help — no online runner carries the tag
those jobs require, so you need to add or retag capacity. And if you paused a
runner and the queue got worse, cicd-aiops undo apply resumes exactly
the runner you paused. Runner administration is GitLab-only; a Gitea target
raises a teaching error here.
3. "The CI server is out of disk"
cicd-aiops rca storage --old-days 30→ projects ranked by repo + artifact bytes with a reclaimable estimate at that age threshold.cicd-aiops artifacts list dev/api→ the actual artifact files, their sizes and their expiry, so you see what "reclaimable" really refers to.cicd-aiops projects --limit 50→ cross-check that the top consumer is the project you expect.cicd-aiops artifacts delete dev/api --older-than-days 30 --dry-run→ shows the scope, deletes nothing.- Re-run without
--dry-run: double confirm, high risk. Optionally setCICD_AUDIT_APPROVED_BY+CICD_AUDIT_RATIONALEto annotate who/why on the audit row. This is irreversible — priorState records the destroyed count and bytes for the audit trail, but there is no undo. cicd-aiops rca storageagain to confirm the reclaimed bytes landed.
Failure branch: never run artifacts delete with --older-than-days 0 as
a first move — 0 means ALL artifacts, including the ones a release or a
running deploy depends on. If the dry-run scope reads "ALL" and you did not
intend that, stop and set an age. Since there is no undo for this operation,
the dry-run is the only safety net you get; if the bloat is mostly repo
bytes rather than artifact bytes, deleting artifacts will not help at all.
4. "Tighten repo hygiene before the release freeze"
cicd-aiops rca stale dev/api --mr-days 14 --branch-days 60→ stale open merge requests, idle branches, and protection gaps such as an unprotected default branch or force-push left allowed.- Confirm the current state via MCP
list_protected_branchesandlist_branchesfor the project. - MCP
list_merge_requests→ the stale MRs the audit named, so you can close or revive them with their owners rather than in bulk. - Close the protection gap with MCP
update_branch_protection(e.g.allow_force_push=Falseon the default branch) — it fetches and captures the prior settings and records an undo that replays them exactly. cicd-aiops undo list→ confirm the protection change is reversible.cicd-aiops rca stale dev/apiagain to confirm the gap closed.
Failure branch: if tightening protection blocks a legitimate workflow — a
release automation that force-pushes tags, say — cicd-aiops undo apply
restores the exact prior protection settings rather than a guessed default.
Fix the automation before re-applying, and do not disable protection
fleet-wide to unblock one job.
Governance & Safety
The skill delivers reads and writes and records them; it does not decide whether a write is permitted. That is your agent's judgement, or the permission of the token you connect it with (a GitLab/Gitea access token without write scope — writes then fail at the server). There is no read-only switch, policy file, or approval gate.
- Audit is the guarantee, and it is not bypassable. Every operation — MCP and CLI alike — is logged to
~/.cicd-aiops/audit.db(relocatable viaCICD_AIOPS_HOME): params, result, status, duration, and the risk tier. The CLI writes the same row the MCP path does. CICD_AUDIT_APPROVED_BY/CICD_AUDIT_RATIONALEare optional annotations recorded on the audit row (who/why); they are never required and never block.- Runaway guard — a safety backstop, not authorization: the same call looped in a tight window trips a circuit breaker. Disable with
CICD_RUNAWAY_MAX=0. - Destructive writes support
--dry-run/dry_run=Trueand double confirmation at the CLI.delete_artifactsis irreversible (priorState records the destroyed count/bytes; audit only). - Reversible writes fetch the real before-state and record an inverse descriptor (pause_runner↔resume_runner, update_branch_protection→prior settings); irreversible ops (retry_pipeline, cancel_pipeline, delete_artifacts) record only the before-state.
References
references/capabilities.md— full tool + platform + API-path referencereferences/cli-reference.md— CLI command referencereferences/setup-guide.md— onboarding, credentials, and connectivityreferences/agent-guardrails.md— which guardrails the harness enforces, the GitLab-only vs Gitea platform asymmetry, and a ready-made system prompt for smaller / local models
Questions people ask
- What platforms does this skill support?
- Self-managed GitLab (any version with REST API v4) and self-hosted Gitea (any version with API v1). GitLab.com SaaS and Gitea Cloud are out of scope.
- How does the governance harness work?
- Every operation logs to ~/.cicd-aiops/audit.db: params, result, status, duration, and risk tier. Tokens are encrypted with Fernet + scrypt. Write operations include undo recording for reversible actions (pause/resume runner, update branch protection) and priorState capture for irreversible ones (delete artifacts). All destructive writes require --dry-run and confirmation.
- What does pipeline_failure_rca actually classify?
- It reads each failed job's failure_reason field and trace-tail markers, then classifies into five categories: test-failure, dependency-network, runner-timeout, OOM, and script-error — each with matched evidence and a suggested action. It never produces a black-box verdict; you can always verify with job_trace_tail.
Related skills
Join video meetings as a voice bot, visual avatar, or avatar with live screen sharing.
Turn China 3C launch inputs into executable routes, messaging, channel actions, risk checks, and review decisions.
Measure whether a transformation changed your growth engine or just added a one-time bump.
Diagnose which founder behavior is capping your growth and get a specific 30-day upgrade move.
Diagnose which mental domain is holding you back before choosing a cognitive intervention.
More from zw008
Browse all skillsOperate VMware VMs, deployments, clusters, guest tasks, and alarms with plan and rollback support.
Inspect VMware health, inventory, alarms, events, and performance without changing infrastructure.
Query Aria Operations metrics, alerts, capacity forecasts, anomalies, and reports from CLI or MCP.
Manage AVI services and pools, and diagnose AKO ingress, sync, certificates, analytics, and health.
Manage Supervisor Namespaces and TKC cluster lifecycles in vSphere Kubernetes Service.
Manage NSX segments, gateways, routing, IP pools, health checks, and connectivity diagnostics.