Skills
MCP
Plugins
Subagents
.fyi
.fyi
Search…
⌘K
…
/
vaquarkhan
/
data-engineering-agent-skills
home
/
skills
/
vaquarkhan
/
data-engineering-agent-skills
vaquarkhan/data-engineering-agent-skills
60 skills
View on GitHub
$
npx skills add vaquarkhan/data-engineering-agent-skills
Skill
Installs
airflow-and-workflow-orchestration
Guides agents through workflow orchestration design and operation across Airflow-style DAGs, cloud-native schedulers, and event-driven pipeline control planes. Use when building or modifying workflow
—
apache-beam-unified-batch-and-stream
Guides agents through Apache Beam pipelines that unify batch and streaming logic. Use when designing Beam transforms, windowing, runners, replay behavior, or portability across execution backends.
—
apache-hudi-lakehouse
Guides agents through Apache Hudi lakehouse design. Use when managing incremental upserts, record-level mutations, timeline behavior, compaction, and Hudi-based lakehouse tables.
—
api-and-saas-ingestion-patterns
Guides agents through API and SaaS ingestion workflows. Use when extracting data from REST, GraphQL, or SaaS platforms with pagination, rate limits, auth rotation, backfills, or unstable source contra
—
avro-protobuf-json-schema-registry
Guides agents through schema-registry-backed event contracts. Use when managing Avro, Protobuf, or JSON Schema for event streams, compatibility policies, producer and consumer evolution, or contract e
—
bigquery-and-dataform-platform-engineering
Guides agents through BigQuery- and Dataform-centered data engineering workflows. Use when designing BigQuery physical models, ingestion boundaries, Dataform transformation workflows, slot or cost con
—
cdc-and-incremental-loading
Guides agents through change data capture and incremental load design. Use when building or modifying watermark-based loads, upserts, deduplication, merge logic, late data handling, or replayable incr
—
clickhouse-real-time-analytics
Guides agents through ClickHouse-based real-time analytics design. Use when building fast analytical serving layers, event aggregations, materialized views, or low-latency metric access patterns.
—
data-catalog-and-discovery
Guides agents through data catalog, discovery, and metadata quality workflows. Use when publishing datasets, improving discoverability, curating lineage metadata, or making data products easier for ot
—
data-contract-testing-with-schema-registry
Guides agents through data-contract testing using schema registries and compatibility checks. Use when validating event contracts, stream schema evolution, consumer compatibility, or release gates for
—
data-lake-and-zone-architecture
Guides agents through data lake and zone architecture design. Use when defining raw, refined, curated, or publish layers; storage organization; retention; and operational boundaries for a data lake.
—
data-mesh-and-domain-oriented-design
Guides agents through domain-oriented data product and data mesh design. Use when organizing ownership, domain boundaries, federated governance, and shared platform responsibilities across multiple te
—
data-migration-and-platform-cutover
Guides agents through data migration and platform cutover workflows. Use when moving pipelines, tables, contracts, orchestration, or workloads between systems, clouds, warehouses, lakehouses, or servi
—
data-observability-and-sla-management
Guides agents through data observability and service-level management. Use when defining or improving freshness, completeness, anomaly detection, alerting, lag tracking, run metadata, and ownership fo
—
data-platform-ci-cd-and-release-management
Guides agents through CI/CD and release management for data platforms. Use when promoting pipeline code, SQL models, contracts, infra, or configuration across environments with validation gates, stage
—
data-platform-disaster-recovery-and-business-continuity
Guides agents through disaster recovery and business continuity planning for data platforms. Use when defining region or account failover, backup and restore, RTO or RPO targets, control-plane recover
—
data-platform-operating-model-and-service-ownership
Guides agents through data platform operating model and ownership design. Use when defining platform team responsibilities, service tiers, golden paths, escalation boundaries, onboarding flows, or han
—
data-quality-and-contract-testing
Drives data implementation with contracts, assertions, and validation evidence. Use when adding or changing ingestion logic, transformations, schemas, or published data products.
—
data-quality-platforms-and-rule-management
Guides agents through data-quality operating models and tool selection. Use when designing rule portfolios, severity levels, ownership, evidence, and enforcement across dbt tests, Great Expectations,
—
data-reconciliation-and-financial-controls
Guides agents through reconciliation and control design for business-critical data. Use when validating financial, operational, or audit-sensitive metrics with source-to-target totals, control balance
—
data-resiliency-testing-and-failure-injection
Guides agents through resiliency testing for data platforms. Use when designing or running failure drills, recovery validation, failover tests, replay-safety checks, dependency outage exercises, or fa
—
data-security-compliance-and-regulated-data
Guides agents through regulated-data security and compliance workflows for PII, PCI, HIPAA, PHI, and similar obligations. Use when data products handle sensitive fields, regulated records, control evi
—
data-sharing-and-publishing-contracts
Guides agents through publishing data products for internal or external consumers. Use when sharing tables, files, extracts, APIs, or reverse-ETL-ready outputs that require stable contracts, ownership
—
data-specification
Creates structured specifications for data products and pipeline changes. Use when starting a new pipeline, model, ingestion flow, or any significant change with unclear requirements.
—
dataplex-and-bigquery-governance
Guides agents through GCP-native data governance workflows with Dataplex and BigQuery. Use when designing lakes, zones, policy tags, metadata quality, lineage, discovery, and governed publishing acros
—
dbt-and-analytics-engineering
Guides agents through analytics engineering workflows with dbt. Use when building or modifying staging models, marts, tests, snapshots, documentation, exposures, or semantic-layer-facing models.
—
debezium-and-kafka-connect-cdc
Guides agents through Debezium and Kafka Connect CDC workflows. Use when streaming database changes into Kafka topics, managing connectors, snapshots, schema evolution, or downstream CDC consumers.
—
delta-lake-and-medallion-architecture
Guides agents through Delta Lake and medallion-style lakehouse design. Use when building or modifying bronze, silver, and gold layers, Delta Lake mutation patterns, streaming-to-batch lakehouse flows,
—
duckdb-local-analytics-and-dev
Guides agents through DuckDB-based local analytics and development workflows. Use when prototyping models locally, validating transformations, reproducing data issues quickly, or building lightweight
—
enterprise-etl-and-data-integration-modernization
Guides agents through operating, hardening, and modernizing enterprise ETL and integration stacks such as Informatica, Talend, DataStage, SSIS, and Matillion. Use when legacy mappings, job orchestrati
—
esg-and-sustainability-regulatory-reporting
Guides agents through ESG, sustainability, and regulatory reporting data products. Use when building governed metrics, traceable evidence, and audit-ready data pipelines for frameworks such as CSRD/ES
—
etl-elt-and-modernization-strategy
Guides agents through ETL, ELT, and transformation-modernization decisions. Use when choosing execution boundaries, redesigning transformation layers, or moving from legacy ETL estates to warehouse- o
—
feature-store-and-ml-data-pipelines
Guides agents through machine-learning data pipelines and feature serving workflows. Use when designing feature generation, offline and online consistency, training-serving parity, point-in-time corre
—
file-and-partner-feed-ingestion
Guides agents through file-based and partner-feed ingestion workflows. Use when landing data from SFTP, managed file transfer, shared buckets, recurring flat files, manifests, or externally supplied f
—
glue-data-catalog-and-lake-formation-governance
Guides agents through AWS-native data catalog and lake governance workflows. Use when designing or reviewing Glue Data Catalog, Lake Formation permissions, governed sharing, metadata quality, and acce
—
great-expectations-deequ-and-cuallee
Guides agents through data-quality frameworks such as Great Expectations, Deequ, and Cuallee. Use when implementing framework-based validation suites, reusable checks, or evidence-driven data-quality
—
incident-triage-and-pipeline-recovery
Guides agents through production data incidents. Use when a pipeline fails, publishes bad data, misses an SLA, partially loads, corrupts state, or requires rollback, replay, or stakeholder communicati
—
java-data-engineering-and-integration-services
Guides agents through Java-based data engineering services and processors. Use when building connectors, ingestion services, stream processors, metadata services, JVM batch tools, or operational integ
—
kafka-resilience-and-schema-evolution
Enforces production Kafka guardrails including non-breaking schema evolution, dead-letter queues for poison messages, and acks=all producer durability. Use when designing or changing Kafka topics, pro
—
lakefs-and-data-versioning
Guides agents through data versioning workflows using lakeFS or similar systems. Use when branching data, validating changes before publish, or controlling risky lakehouse operations with versioned da
—
lakehouse-table-format-engineering
Guides agents through lakehouse table design and open table format decisions. Use when designing or changing Iceberg, Delta, Hudi, partitioning, schema evolution, compaction, or batch and streaming in
—
lineage-pii-and-governance
Applies governance, lineage, ownership, and sensitive-data controls to data changes. Use when a pipeline touches published datasets, regulated information, or shared business metrics.
—
lower-environment-data-masking-and-obfuscation
Guides agents through masking, obfuscating, and safely promoting production-like data into lower environments. Use when QA, development, or staging needs realistic data without exposing production-sen
—
mainframe-modernization-and-data-offload
Guides agents through mainframe data modernization and offload workflows. Use when migrating or exposing data from COBOL, JCL, VSAM, IMS, DB2 for z/OS, or batch-oriented mainframe estates into modern
—
master-data-and-entity-resolution
Guides agents through master data and entity resolution workflows. Use when matching identities across systems, defining canonical entities, resolving duplicates, or building golden records for shared
—
mcp-data-observability-integration
Guides agents to wire Model Context Protocol servers for live data platform observability including Spark execution plans, OOM diagnosis, Kafka consumer lag, and orchestration run state. Use when agen
—
microsoft-purview-and-azure-data-governance
Guides agents through Microsoft Purview and Azure-native data governance workflows. Use when designing collections, scans, classifications, lineage, policy boundaries, and governed publishing across A
—
notebook-to-production-hardening
Guides agents through converting exploratory notebooks into production-ready data jobs. Use when operationalizing notebooks from Databricks, Jupyter, or similar environments into tested, packaged, rep
—
openmetadata-datahub-and-openlineage
Guides agents through metadata platform and lineage workflows using OpenMetadata, DataHub, or OpenLineage-compatible systems. Use when improving discovery, lineage quality, metadata governance, or pro
—
operational-datastore-selection-relational-and-nosql
Guides agents through choosing relational operational stores such as MySQL versus NoSQL options such as document, key-value, wide-column, or cache-backed systems. Use when deciding where application-a
—
orchestration-and-backfills
Designs scheduling, reruns, and backfills safely for data systems. Use when changing orchestration, retries, dependency timing, historical reprocessing, or publish sequencing.
—
pipeline-planning-and-task-breakdown
Breaks approved data specifications into safe, verifiable implementation tasks. Use when a data project spans multiple steps, systems, or files and needs dependency-aware sequencing.
—
privacy-retention-and-right-to-delete
Guides agents through privacy, retention, and deletion workflows in data systems. Use when handling personal data, retention limits, deletion requests, legal holds, or data minimization requirements a
—
python-data-engineering-and-pipeline-packaging
Guides agents through Python-based data engineering implementation. Use when building or modifying Python ingestion jobs, orchestration helpers, PySpark entry points, validation code, packaging, depen
—
regional-data-compliance-and-sovereignty
Guides agents through region-specific data compliance, residency, sovereignty, and transfer design. Use when data products must operate across jurisdictions such as Europe, the USA, India, Saudi Arabi
—
reverse-etl-and-operational-data-serving
Guides agents through reverse ETL and operational data serving workflows. Use when sending curated warehouse data to business systems, SaaS tools, APIs, activation layers, or operational applications
—
safe-backfill-and-replay-orchestration
Forces replay-safe rollout plans, reconciliation gates, and rollback paths before executing any data backfill or historical reprocessing. Use when running /backfill, rerunning pipelines, repairing pub
—
scala-data-engineering-on-jvm-runtimes
Guides agents through Scala-based data engineering on JVM runtimes. Use when building Spark, Flink, Kafka Streams, or other Scala data jobs that require explicit build, packaging, type, and runtime di
—
schema-evolution-and-contract-migrations
Guides agents through schema changes and contract migrations. Use when adding, renaming, removing, or changing columns, data types, nullability, keys, or downstream-facing data contracts.
—
semantic-layer-and-metric-governance
Guides agents through semantic layer and shared metric design. Use when defining business metrics, reusable dimensions, governed metric contracts, or shared semantic models consumed by dashboards, ana
—