Philadelphia, PA, USA Hybrid Full-time

IntegriChain is hiring a Senior Database Engineer - Platform Engineering

Responsibilities

  • Design and implement hybrid data architectures using relational, columnar, and NoSQL databases, selecting optimal technologies based on workload requirements.
  • Build cloud-native data lakehouse systems on AWS with S3, Glue, Lake Formation, and open data formats, targeting Azure Data Lake as a secondary environment.
  • Apply Medallion Architecture principles to structure data pipelines into raw, cleaned, and business-ready layers for analytics and reporting.
  • Develop integrated data platforms that bridge transactional databases with analytical warehouses for consistent data access.
  • Create semantic layers and reusable analytics models to standardize metrics across business intelligence, analytics, and machine learning applications.
  • Design efficient data models, ETL/ELT workflows, and query optimizations for both operational and analytical performance.
  • Engineer automated replication, partitioning, and lifecycle management strategies using infrastructure-as-code to eliminate manual database administration.
  • Implement version-controlled schema migration pipelines using tools like Flyway or Liquibase within CI/CD environments.
  • Design API-first data access interfaces to decouple applications from direct database connections and improve maintainability.
  • Develop batch, micro-batch, and near-real-time data pipelines using AWS services such as Glue, Kinesis, MSK, and orchestration tools like dbt and Airflow.
  • Construct streaming data architectures with Kinesis and MSK to support low-latency ingestion into multiple database systems.
  • Enforce data quality, schema validation, lineage tracking, and monitoring across all data workflows.
  • Optimize data platform performance, cost, and scalability from ingestion through transformation to consumption.
  • Deploy change data capture solutions using AWS DMS, Debezium, or native database features to replicate data across systems.
  • Design high-performance DynamoDB schemas using single-table patterns, indexes, and stream processing for event-driven applications.
  • Architect DocumentDB clusters for workloads requiring flexible document structures and hierarchical data modeling.
  • Deploy and manage OpenSearch or ElasticSearch clusters for search, log analysis, and system observability.
  • Assess and recommend appropriate NoSQL solutions based on query patterns, latency needs, and cost efficiency.
  • Implement caching strategies using TTL, DAX, and ElastiCache to support high-throughput data access.
  • Leverage AI and ML techniques for data quality monitoring, anomaly detection, schema drift identification, and workload analysis using SageMaker and Bedrock.
  • Construct data pipelines and storage structures tailored for machine learning, including feature stores and training datasets.
  • Integrate generative AI capabilities via Bedrock to automate data documentation, query creation, and cataloging.
  • Deploy vector databases using pgvector or OpenSearch k-NN for AI-powered similarity search and retrieval-augmented generation.
  • Enhance observability with machine learning to detect anomalies in pipeline behavior, query patterns, and data quality metrics.
  • Manage all data infrastructure through code using Terraform and AWS CDK, covering databases, networking, and IAM configurations.
About company
IntegriChain

We empower biopharma to ensure patient access to therapy and sustain research intensity by optimizing net revenue performance and integrity.

IntegriChain is trusted by biopharma companies, from emerging biotech to Top 50 global manufacturers, to drive net revenue performance and compliance. We empower organizations to ensure patient access today while reinvesting in R&D for tomorrow.

IntegriChain is powered by ICyte, the industry's only integrated end-to-end platform that manages, adjudicates, and connects pricing, rebates, claims, channel, and patient data for visibility, control, and decisioning support. By aggregating and operationalizing all data channels that drive gross-to-net, we turn complexity into clarity, minimizing revenue leakage while maximizing demand conversion by ensuring patient speed to therapy.

All jobs at IntegriChain Visit website
Job Details
Department AI, Product & Technology
Category DevOps & SRE
Posted 2 months ago