About the role
Location - Irving, TX/Charlotte, NC
Contract
Responsibilities
Design, build, and harden production-grade data pipelines for ingest, curation, reconciliation, DQ monitoring, lineage, and notifications.
Contribute to and extend the generic pipeline framework so federated teams can onboard data consistently and at speed.
Engineer data transformations primarily using PySpark on GCP (Dataproc/Sparkflow) and BigQuery as the main query engine.
Implement best practices for schema management, performance, reliability, and cost efficiency across large-scale datasets (Parquet, Iceberg).
Integrate with centralized data cataloging and governance (e.g., Dataplex/Knowledge Catalog or equivalent), and support automated lineage harvesting.
Collaborate with platform, security, and central cloud teams; support hybrid access patterns using Starburst where needed.
Operate with an ownership mindset in an on-time, production SLAs environment, ensuring resilient, observable, and recoverable data flows.
Key Qualifications Required (for all levels):
3+ years (L3) / 6+ years (L4) hands-on data engineering in production environments.
Strong SQL skills and proficiency building data transformations at scale.
Practical experience with Spark (preferably PySpark) and big data file formats (Parquet; familiarity with Iceberg concepts).
Experience delivering reliable, scheduled pipelines with monitoring, alerting, and incident response.
Understanding of data cataloging, metadata, and lineage concepts, and why they matter for governance and reuse.
Exposure to GCP data services (e.g., Dataproc/Spark, BigQuery). Deep infra setup knowledge is not required; collaboration with central cloud teams is expected.
Solid software engineering practices: version control, CI/CD, testing, code reviews, and documentation.
Preferred:
Experience building reusable pipeline frameworks or platform components adopted by multiple teams.
Familiarity with Starburst/Trino for federated query across hybrid environments.
Knowledge of data quality frameworks, reconciliation, and auditability in regulated or enterprise settings.
Python proficiency; Java experience a plus (willingness to work primarily in Python).
Background in operational data environments with time-bound SLAs and prod support rotations.
Exposure to AI/ML model lifecycle platforms or MLOps concepts (platform enablement rather than model authoring).
Soft Skills
High ownership, bias for action, and clear communication with technical and non-technical partners.
Comfort operating in a fast-scaling organization with evolving scope. has context menu