.fyi
SkillsMCPPluginsSubagents

Browse by category

DevOps & CI/CD SkillsProductivity & Workflow SkillsOther SkillsProduct & Project Management SkillsDocumentation & Knowledge SkillsCode Review & Refactor SkillsBackend & APIs SkillsAgent Meta & Communication SkillsResearch SkillsSecurity SkillsUX UI & Design SkillsTesting & QA SkillsSee all →

Every Claude Code skill, MCP server, plugin and subagent in one directory. Searchable, comparable, and one command from installed. Live stats from GitHub, npm and PyPI.

We're on Product HuntYour agent's app storeCheck it out →
Agent SkillsMCP ServersPluginsSubagentsCoding Agents
CollectionsOfficial publishersGlossaryFAQBlogSearchSavedFeedback
PrivacyTermsllms.txtSitemap

made with ♥ · © 2026 aaaa.fyi

Independent project · real data from public registries

…/vaquarkhan/data-engineering-agent-skills
home/skills/vaquarkhan/data-engineering-agent-skills
vaquarkhan avatar

vaquarkhan/data-engineering-agent-skills

53 skills

View on GitHub
$npx skills add vaquarkhan/data-engineering-agent-skills
SkillInstalls
airflow-and-workflow-orchestrationGuides agents through workflow orchestration design and operation across Airflow-style DAGs, cloud-native schedulers, and event-driven pipeline control planes.—apache-beam-unified-batch-and-streamGuides agents through Apache Beam pipelines that unify batch and streaming logic.—apache-hudi-lakehouseGuides agents through Apache Hudi lakehouse design. Use when managing incremental upserts, record-level mutations, timeline behavior, compaction, and…—api-and-saas-ingestion-patternsGuides agents through API and SaaS ingestion workflows. Use when extracting data from REST, GraphQL, or SaaS platforms with pagination, rate limits, auth…—avro-protobuf-json-schema-registryGuides agents through schema-registry-backed event contracts. Use when managing Avro, Protobuf, or JSON Schema for event streams, compatibility policies,…—bigquery-and-dataform-platform-engineeringGuides agents through BigQuery- and Dataform-centered data engineering workflows.—cdc-and-incremental-loadingGuides agents through change data capture and incremental load design.—clickhouse-real-time-analyticsGuides agents through ClickHouse-based real-time analytics design.—data-catalog-and-discoveryGuides agents through data catalog, discovery, and metadata quality workflows.—data-contract-testing-with-schema-registryGuides agents through data-contract testing using schema registries and compatibility checks.—data-lake-and-zone-architectureGuides agents through data lake and zone architecture design. Use when defining raw, refined, curated, or publish layers; storage organization; retention; and…—data-mesh-and-domain-oriented-designGuides agents through domain-oriented data product and data mesh design.—data-migration-and-platform-cutoverGuides agents through data migration and platform cutover workflows.—data-observability-and-sla-managementGuides agents through data observability and service-level management.—data-platform-ci-cd-and-release-managementGuides agents through CI/CD and release management for data platforms.—data-platform-disaster-recovery-and-business-continuityGuides agents through disaster recovery and business continuity planning for data platforms.—data-platform-operating-model-and-service-ownershipGuides agents through data platform operating model and ownership design.—data-quality-and-contract-testingDrives data implementation with contracts, assertions, and validation evidence.—data-quality-platforms-and-rule-managementGuides agents through data-quality operating models and tool selection.—data-reconciliation-and-financial-controlsGuides agents through reconciliation and control design for business-critical data.—data-resiliency-testing-and-failure-injectionGuides agents through resiliency testing for data platforms. Use when designing or running failure drills, recovery validation, failover tests, replay-safety…—data-security-compliance-and-regulated-dataGuides agents through regulated-data security and compliance workflows for PII, PCI, HIPAA, PHI, and similar obligations.—data-sharing-and-publishing-contractsGuides agents through publishing data products for internal or external consumers.—data-specificationCreates structured specifications for data products and pipeline changes.—dataplex-and-bigquery-governanceGuides agents through GCP-native data governance workflows with Dataplex and BigQuery.—dbt-and-analytics-engineeringGuides agents through analytics engineering workflows with dbt.—debezium-and-kafka-connect-cdcGuides agents through Debezium and Kafka Connect CDC workflows.—delta-lake-and-medallion-architectureGuides agents through Delta Lake and medallion-style lakehouse design.—duckdb-local-analytics-and-devGuides agents through DuckDB-based local analytics and development workflows.—enterprise-etl-and-data-integration-modernizationGuides agents through operating, hardening, and modernizing enterprise ETL and integration stacks such as Informatica, Talend, DataStage, SSIS, and Matillion.—esg-and-sustainability-regulatory-reportingGuides agents through ESG, sustainability, and regulatory reporting data products.—etl-elt-and-modernization-strategyGuides agents through ETL, ELT, and transformation-modernization decisions.—feature-store-and-ml-data-pipelinesGuides agents through machine-learning data pipelines and feature serving workflows.—file-and-partner-feed-ingestionGuides agents through file-based and partner-feed ingestion workflows.—glue-data-catalog-and-lake-formation-governanceGuides agents through AWS-native data catalog and lake governance workflows.—great-expectations-deequ-and-cualleeGuides agents through data-quality frameworks such as Great Expectations, Deequ, and Cuallee.—incident-triage-and-pipeline-recoveryGuides agents through production data incidents. Use when a pipeline fails, publishes bad data, misses an SLA, partially loads, corrupts state, or requires…—java-data-engineering-and-integration-servicesGuides agents through Java-based data engineering services and processors.—kafka-resilience-and-schema-evolutionEnforces production Kafka guardrails including non-breaking schema evolution, dead-letter queues for poison messages, and acks=all producer durability.—lakefs-and-data-versioningGuides agents through data versioning workflows using lakeFS or similar systems.—lakehouse-table-format-engineeringGuides agents through lakehouse table design and open table format decisions.—lineage-pii-and-governanceApplies governance, lineage, ownership, and sensitive-data controls to data changes.—lower-environment-data-masking-and-obfuscationGuides agents through masking, obfuscating, and safely promoting production-like data into lower environments.—mainframe-modernization-and-data-offloadGuides agents through mainframe data modernization and offload workflows.—master-data-and-entity-resolutionGuides agents through master data and entity resolution workflows.—mcp-data-observability-integrationGuides agents to wire Model Context Protocol servers for live data platform observability including Spark execution plans, OOM diagnosis, Kafka consumer lag,…—microsoft-purview-and-azure-data-governanceGuides agents through Microsoft Purview and Azure-native data governance workflows.—notebook-to-production-hardeningGuides agents through converting exploratory notebooks into production-ready data jobs.—openmetadata-datahub-and-openlineageGuides agents through metadata platform and lineage workflows using OpenMetadata, DataHub, or OpenLineage-compatible systems.—operational-datastore-selection-relational-and-nosqlGuides agents through choosing relational operational stores such as MySQL versus NoSQL options such as document, key-value, wide-column, or cache-backed…—orchestration-and-backfillsDesigns scheduling, reruns, and backfills safely for data systems.—pipeline-planning-and-task-breakdownBreaks approved data specifications into safe, verifiable implementation tasks.—privacy-retention-and-right-to-deleteGuides agents through privacy, retention, and deletion workflows in data systems.—