Data Engineer
W3Global
Role: Sr Data Engineer - Disney
Location: Glendale, CA (Hybrid - 2- 4 Days Onsite)
Full Time
In Person is interview is must
Mandatory skills: Databricks experience (primary requirement), Apache Airflow, Advanced SQL skills, Python, Spark / PySpark, Scala, Experience building and maintaining data pipelines and workflows
Key Responsibilities
- Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python
- Meet with stakeholders to gather requirements and translate them into scalable data platform solutions
- Understanding of Databricks platform and developer tooling to diagnose errors, audit platform activity, and automate updates across pipelines, objects, and integrations
- Ability to explain Spark architecture and pipeline behavior to stakeholders to diagnose root causes and recommend solutions
- Provide solution architecture across AWS, Databricks, Kubernetes, and Airflow (MWAA), including cross-platform integrations
- Manage Databricks platform governance, including Unity Catalog, ACLs, lineage, and data discovery and privacy tooling
- Build and maintain Kubernetes containers and containerized utilities supporting deployed data platform services
- Apply networking knowledge to troubleshoot connectivity and integration errors across platform components
- Perform platform administration: provision and remove access, assess resource utilization, monitor platform health and cost, and evaluate stakeholder requests
- Collaborate with engineers, architects, and product managers to drive Core Data platform success; participate in agile/scrum ceremonies
- Maintain documentation of platform changes, standards, and pipeline configurations to support data quality and governance
Qualifications
- 5+ years of data engineering experience developing and operating large-scale data pipelines
- Deep hands-on experience with Databricks and Apache Spark (batch and streaming), including pipeline development in PySpark and/or Scala
- Strong understanding of Spark architecture-executors, stages, partitioning, shuffle, and performance tuning-with ability to explain tradeoffs to technical and non-technical stakeholders
- Proficiency with Databricks platform tooling (API, SDK, CLI) for automation, auditing, governance, and operational troubleshooting
- Proficient in SQL with advanced performance tuning capabilities
- Hands-on production experience with Airflow (MWAA) for orchestrating data pipelines
- Experience managing Databricks platform governance: ACLs, Unity Catalog, lineage, and access provisioning
- Proficiency in Python and at least one additional language (Scala, Kotlin, or SQL-driven pipeline tooling)