Apply on company website Philadelphia, Pennsylvania, United States Hybrid Full-time

Medical Guardian is hiring a Principal Data Engineer

Responsibilities

  • Design, build, optimize, and operate production-grade batch and streaming data pipelines on Azure and Databricks, with a primary focus on real-time IoT and telemetry use cases within a Medallion architecture.
  • Develop ETL/ELT workflows to ingest, transform, validate, and serve large volumes of structured, semi-structured, unstructured, and streaming data.
  • Build and maintain reliable data products, data services, APIs, and microservices that support operational applications, analytics, software engineering, and ML/AI teams.
  • Use Python, PySpark, Spark SQL, SQL, Delta Lake, Databricks Workflows, CI/CD, and related tools to build maintainable, testable, and observable data systems.
  • Troubleshoot complex production pipeline issues across Databricks, Azure, streaming systems, APIs, and source systems, including root cause analysis, corrective action, and prevention planning.
  • Move quickly from rough business need to prototype, pilot, and production-ready data capability while maintaining appropriate engineering discipline.
  • Lead the design and delivery of real-time streaming ingestion and processing patterns for connected medical device telemetry, event data, and operational data feeds.
  • Implement streaming solutions using Azure Event Hubs, Azure Stream Analytics, Databricks, Delta Lake, and related Azure integration patterns.
  • Design cost-effective throughput, partitioning, delivery, retention, and replay strategies for high-volume event and telemetry workloads.
  • Create consumption patterns that support APIs, microservices, operational applications, near-real-time decisioning, analytics, and ML/AI use cases.
  • Define reliability, latency, quality, observability, and supportability expectations for production streaming systems.
  • Set direction for Databricks-based data engineering patterns, including Medallion architecture, Delta Lake, Spark optimization, data modeling, data quality, and reusable pipeline design.
  • Optimize production Databricks pipelines using PySpark, Spark SQL, Delta Lake, partitioning strategies, caching, shuffle optimization, cluster/job configuration, and cost-aware design.
  • Establish practical standards for pipeline structure, code organization, testing, deployment, monitoring, documentation, and ownership.
  • Partner with data platform, security, infrastructure, and engineering teams to ensure the data platform is scalable, secure, reliable, and aligned with enterprise architecture.
  • Make pragmatic architecture tradeoffs between speed, durability, cost, governance, performance, and downstream business impact.
  • Work directly with business, product, analytics, ML/AI, operations, software engineering, and leadership stakeholders to clarify what data is needed, why it matters, how it will be used, and what success looks like.
  • Translate ambiguous business needs into concrete data requirements, data product definitions, architecture options, delivery priorities, and implementation plans.
  • Ask practical questions early: who will use the data, what decision or workflow does it support, what latency and quality are required, what happens if the data is wrong or late, and how will we know the capability is creating value?
  • Help the organization avoid becoming a data ticket factory by shaping solutions, not just executing requests.
  • Communicate architecture decisions, tradeoffs, risks, dependencies, and delivery options clearly to technical and non-technical stakeholders.
  • Manage, mentor, and develop data engineers, providing clear expectations, technical guidance, prioritization support, feedback, and accountability.
  • Provide technical leadership through hands-on example, strong engineering judgment, clear recommendations, and pragmatic decision-making.
  • Lead design reviews, code reviews, production readiness reviews, incident reviews, and architecture discussions across data engineering initiatives.
  • Establish and improve engineering standards for data quality, testing, CI/CD, observability, documentation, runbooks, cost management, privacy, and security.
  • Proactively identify platform risks, data gaps, unclear ownership, operational weaknesses, and opportunities to improve reliability, scalability, and delivery speed.
  • Influence without relying only on formal authority by building trust, framing tradeoffs, and helping cross-functional teams get to decisions.
  • Partner with ML engineers, data scientists, analytics engineers, and analysts to deliver reliable data pipelines, feature pipelines, training datasets, scoring inputs, and feedback loops.
  • Support the data foundation for predictive models, risk scores, operational decisioning, GenAI workflows, RAG, document intelligence, summarization, and AI-enabled automation.
  • Help define data contracts, model-ready datasets, feature definitions, lineage, and monitoring expectations for ML/AI and analytics use cases.
  • Ensure downstream consumers understand the meaning, freshness, quality, limitations, and appropriate use of the data products they depend on.
  • Apply privacy-first, security-aware, and governance-aligned practices for regulated, sensitive, and operationally critical data.
  • Design and implement data quality checks, validation rules, anomaly detection, schema expectations, alerting, and operational monitoring.
  • Ensure production pipelines and services are supportable, observable, documented, recoverable, and aligned with business continuity needs.
  • Drive incident response and continuous improvement for data platform and pipeline issues, including root cause analysis and preventative remediation.
  • Balance innovation with reliability, compliance, privacy, cost discipline, and operational usefulness.

Requirements

  • Significant hands-on experience with Databricks, Spark, SQL, Python/PySpark, Azure services, streaming architectures, data quality frameworks, pipeline automation, CI/CD, and production troubleshooting.
  • Proven ability to operate as a principal-level leader: shaping unclear requirements, making pragmatic technical decisions, managing and mentoring engineers, and driving work forward without waiting for perfect specifications.
  • Experience in designing, building, operating, and improving data pipelines and data products with a focus on real-time streaming, IoT telemetry, Databricks, Azure, data services for APIs and microservices, and reliable data products for downstream consumption.
  • Strong understanding that data pipelines and data services are product capabilities supporting real users, real workflows, operational decisions, ML/AI systems, APIs, analytics, and measurable business outcomes.
  • Ability to troubleshoot complex production pipeline issues across Databricks, Azure, streaming systems, APIs, and source systems, including root cause analysis, corrective action, and prevention planning.
  • Experience working directly with business, product, software engineering, analytics, ML/AI, operations, and leadership stakeholders to define data needs, consumption patterns, production guarantees, and success metrics.
  • Demonstrated experience in commercial software, SaaS, digital products, healthtech, fintech, IoT, data platforms, or other product-driven environments.

Nice to Have

  • Background in healthtech, fintech, IoT, or regulated product environments is strongly preferred.
  • Experience with Medallion architecture, Delta Lake, and Spark optimization in Databricks environments.
  • Familiarity with Azure Event Hubs, Azure Stream Analytics, and related Azure integration patterns.
  • Experience supporting ML/AI systems, including feature pipelines, training datasets, scoring inputs, and feedback loops.
  • Experience with GenAI workflows, RAG, document intelligence, summarization, and AI-enabled automation.

Work Arrangement

Hybrid

Required Skills
DatabricksApache SparkSQLCI/CDSaaSFintechIoTRAG
Test your skills for this role

Take a short quiz and show this employer what you can do.

Job Details
Location Philadelphia, Pennsylvania, United States
Work mode Hybrid
Employment Full-time
Department Engineering
Category Data & ML
Posted 2 months ago
Application On company website
or drop your CV first
About company
Medical Guardian logo
Founded in 2005, Medical Guardian is a leading provider of innovative senior health solutions, with 625,000+ active members across the country. The company offers a full suite of connected-care medical alert systems and engagement services that empower older adults to live a life without limits and age safely at home.
All jobs at Medical Guardian Visit website