Responsibilities
- Lead agile sprint execution in collaboration with stakeholders across business areas, collecting requirements and delivering practical data-driven solutions.
- Integrate patient-level data from sources such as clinical trials and electronic health records across organizational divisions.
- Work closely with project leads to ensure data systems, tools, and AI applications are scalable, effective, and deliver measurable impact.
- Support long-term data, AI, and tooling strategies by contributing practical insights and process improvements from hands-on experience.
- Promote adoption and change by engaging teams early in the research and development lifecycle.
- Develop AI pipelines that process raw clinical trial data—including adverse events, lab results, medical history, procedures, drug names, and subject exposure—and convert it into FDA-compliant SDTM formats using LLM-based term resolution, phonetic matching, and rule-based engines.
- Evaluate and verify clinical trial datasets across extensive study collections to confirm accurate transformation into CDISC SDTM standards.
- Design and manage a Retrieval-Augmented Generation (RAG) system that indexes clinical documentation from OHDSI resources into vector databases and supplies contextual data to LLMs for automated OMOP CDM field mapping.
- Create automated ETL processes that generate SQL code by analyzing source database schemas, producing Impala-compatible queries to populate OMOP CDM tables, replacing lengthy manual efforts with a single command.
- Develop systems that resolve clinical codes (e.g., LOINC, MedDRA, ICD-10, RxNorm) to OMOP concept identifiers using vocabulary databases, embedding models, and LLM reasoning for uncertain cases.
- Implement multi-tiered data quality validation using Impala SQL translations of R/JDBC checks, including constraint validation, descriptive analytics, and sanity checks, generating automated HTML reports post-ETL.
- Build neuro-symbolic AI workflows that process clinical term standardization through four stages—rule-based logic, phonetic algorithms, embedding similarity, and LLM inference—with full traceability and audit logs for each decision.
- Design and implement full-stack applications with React frontends, Python APIs, and Impala databases, including dashboards for reviewing mappings with confidence scores, RAG references, and approval workflows.
- Integrate interactive analytics platforms such as Qlik Sense and Power BI into applications to visualize data quality, mapping progress, and stakeholder feedback.
- Construct semantic classification models that automatically detect meaning, data type, clinical domain, sensitivity, and target OMOP field for any database column.
- Build LLM-driven data profiling tools that analyze source schemas, assess tables and columns, map relationships, and produce natural language documentation.
- Apply unsupervised machine learning methods like embedding-based clustering (e.g., K-Means) to standardize medical terms, including indications.
- Design feedback mechanisms where data quality errors and manual corrections generate new rules, expose vocabulary gaps, and enrich training datasets to improve system performance over time.
- Lead and mentor interdisciplinary teams and university student groups in developing clinical data tools, such as terminology and lab name mappers, delivering on defined timelines by translating technical needs into sprint tasks.
- Collaborate across organizational units to scale data extraction platforms through training sessions and joint initiatives.