.fyi
SkillsMCPPluginsSubagents

Browse by category

DevOps & CI/CD SkillsProductivity & Workflow SkillsOther SkillsProduct & Project Management SkillsDocumentation & Knowledge SkillsCode Review & Refactor SkillsBackend & APIs SkillsAgent Meta & Communication SkillsResearch SkillsSecurity SkillsUX UI & Design SkillsTesting & QA SkillsSee all →

Every Claude Code skill, MCP server, plugin and subagent in one directory. Searchable, comparable, and one command from installed. Live stats from GitHub, npm and PyPI.

We're on Product HuntYour agent's app storeCheck it out →
Agent SkillsMCP ServersPluginsSubagentsCoding Agents
CollectionsOfficial publishersGlossaryFAQBlogSearchSavedFeedback
PrivacyTermsllms.txtSitemap

made with ♥ · © 2026 aaaa.fyi

Independent project · real data from public registries

…/bregman-arie/devops-sre-skills
home/skills/bregman-arie/devops-sre-skills
bregman-arie avatar

bregman-arie/devops-sre-skills

17 skills

View on GitHub
$npx skills add bregman-arie/devops-sre-skills
SkillInstalls
_templateOne-line description of when this skill is used.—diagnose-crashloopTriage pods restarting repeatedly and identify the most likely root cause.—diagnose-imagepullbackoffDetermine why a pod cannot pull its container image and resolve safely.—investigate-driftDetermine why actual infrastructure differs from Terraform state and choose a safe reconciliation path.—recover-state-lockSafely assess and recover from a stuck Terraform state lock.—sev1-first-15-minutesExecute the initial incident workflow to stabilize, communicate, and delegate.—triage-access-deniedIdentify why an AWS API call is denied and what policy element blocks it.—triage-app-outofsyncIdentify why an Argo CD application is OutOfSync and resolve safely.—triage-cost-spikeIdentify the primary drivers of a sudden cloud cost increase and implement safe mitigations.—triage-eks-node-notreadyDiagnose EKS worker nodes in NotReady and determine safe remediation.—triage-error-budget-burnInvestigate rapid SLO burn and identify whether the driver is errors, latency, or availability.—triage-latency-regressionIdentify what changed and where latency increased, using logs/metrics/traces.—triage-node-pressureDiagnose node-level memory/disk/pid pressure and determine safe mitigations.—triage-pending-podsDiagnose pods stuck in Pending and identify scheduling constraints.—triage-quota-exceededDiagnose GCP quota errors and identify the quota, scope, and fastest safe mitigation.—triage-service-dnsDiagnose in-cluster DNS resolution failures and isolate root causes.—triage-suspected-secret-exposureContain and respond to a suspected credential/secret exposure without increasing blast radius.—