Responsibilities
- Create and manage batch and real-time pipelines for ingesting, transforming, and orchestrating data.
- Construct and optimize data Lakehouse systems using platforms like Azure Synapse or Delta Lake.
- Integrate diverse data sources including relational databases, NoSQL stores, files, and IoT streams.
- Implement and manage ETL and ELT workflows using tools such as Azure Data Factory, Databricks, or Apache Spark.
- Partner with data science teams to deliver clean, model-ready datasets for machine learning tasks.
- Enforce data quality standards and implement lineage and governance practices across data layers.
- Support business analytics by integrating with visualization tools like Power BI, Superset, or Tableau.
- Operationalize data APIs and machine learning models using Azure ML, Kubernetes, and CI/CD systems.
- Ensure data workflows are scalable, observable, and performant through monitoring and automation.
- Work closely with product, engineering, and business units to turn insights into actionable outcomes.
Requirements
- 4 to 8 years of professional experience in data engineering or related fields.
- Proficiency in Python, SQL, and at least one compiled language such as C#, Java, or Scala.
- Hands-on experience with relational databases including SQL Server, PostgreSQL, and MySQL, as well as NoSQL systems like MongoDB or Cosmos DB.
- Direct experience working with Azure Data Lake and Databricks platforms.
- Experience using ETL/ELT tools such as Azure Data Factory, Apache Airflow, or dbt.
- Knowledge of messaging and streaming technologies including Kafka, Event Hubs, or Service Bus.
- Exposure to machine learning frameworks like TensorFlow or PyTorch and understanding of MLOps principles.
- Experience with data visualization and analytics platforms such as Power BI, Apache Superset, or Tableau.
- Strong background in cloud platforms, containerization, and DevOps tools including Azure, Docker, Kubernetes, and GitHub Actions or Azure DevOps.
- Solid grasp of data modeling, version control, and CI/CD practices specific to data systems.
Nice to Have
- Experience building knowledge graphs or semantic search systems.
- Understanding of retrieval-augmented generation (RAG) patterns using large language models.
- Familiarity with emerging data architectures such as data mesh, data fabric, or domain-driven design.
- Experience with MLflow, Delta Live Tables, or Databricks Unity Catalog.
Work Arrangement
On-site — Chennai, Hyderabad, Bangalore
We’re on a Mission
Founded in 2005, the company pioneered the first digital validation lifecycle management system, transforming compliance processes in life sciences. Its flagship platform set the industry benchmark and continues to evolve as part of a broader digital transformation suite. Today, the organization advances innovation by combining purpose-built software with expert consulting to help regulated industries meet changing quality and compliance demands.
The Team You’ll Join
The team prioritizes customer success in every function, from product development to support. Believing in open communication, mutual support, and shared responsibility, members consistently go beyond their core roles. Innovation and ambition are central to the culture, driving both product evolution and personal growth. The organization is committed to becoming the leading intelligent validation platform, aiming for market leadership without compromise.
How We Work
The offices in Chennai, Hyderabad, and Bangalore operate on a full-time in-person schedule, five days a week. The company emphasizes face-to-face collaboration as essential for creativity, team cohesion, and long-term success. As an equal-opportunity employer, hiring and advancement are based on merit. All qualified candidates are considered without regard to race, religion, sex, sexual orientation, gender identity, national origin, disability, or other protected characteristics under local law.