About IQVIA

IQVIA provides scientific services spanning clinical trials, real world evidence, and consulting in all areas of the product lifecycle. Our Clinical Outcomes Assessments (COAs) organisation leads the industry in generating data to ensure that the patient voice is incorporated into the development and commercialisation of medication and other drug/non-drug interventions.

Role Summary

The COA Accelerator ecosystem is expanding its AI-enabled capabilities to help internal teams and external clients generate evidence-backed COA strategy recommendations. These capabilities depend on high-quality, traceable, well-governed, AI-ready data assets. The Data Engineer will play a critical role in transforming fragmented clinical, regulatory, scientific, and proprietary content into structured, searchable, and secure knowledge assets that power AI-enabled COA strategy workflows.

Responsibilities

  • Design, build, and maintain data infrastructure that supports IQVIA’s AI-enabled COA strategy and COA Accelerator capabilities.
  • Own the ingestion, transformation, normalisation, enrichment, indexing, versioning, and governance of public, proprietary, and client-specific data sources.
  • Build ingestion pipelines for structured and unstructured sources, including PDFs, Word documents, slide decks, spreadsheets, databases, APIs, clinical trial registries, regulatory documents, scientific publications, and internal repositories.
  • Transform raw source material into standardised, searchable, AI-ready formats that support evidence retrieval, source citation, recommendation generation, and expert review workflows.
  • Develop repeatable processes for document parsing, OCR, text extraction, metadata enrichment, chunking, deduplication, versioning, indexing, and quality control.
  • Provide technical support to the teams building and maintaining the platform’s core knowledge layer.
  • Support integration of public data sources such as clinical trial registries, FDA labels, EMA EPARs, HTA records, scientific literature, FDA guidance, public qualification documents, and other relevant evidence repositories.
  • Prepare data for retrieval-augmented generation workflows through high-quality chunking, embeddings, indexes, metadata filters, and source reference structures.
  • Collaborate with AI engineers to improve retrieval precision, recall, relevance, and citation accuracy.
  • Implement hybrid retrieval approaches combining semantic search, keyword search, structured database queries, and metadata filtering.
  • Maintain traceability between AI-generated outputs and source documents.
  • Implement data quality controls to identify incomplete, outdated, duplicated, poorly parsed, incorrectly tagged, or otherwise unreliable content.
  • Maintain audit trails for source ingestion, transformation, updates, deletions, access rights, and downstream use.
  • Work with legal, security, compliance, product, and domain stakeholders to ensure data use aligns with contractual, licensing, privacy, intellectual property, and governance requirements.

Individuals joining us are assured of a rewarding and progressive career in patient-focused research. You’ll have the opportunity to address challenging client issues, across multiple geographies, with a hands-on influence in developing and delivering innovative solutions. We operate in a truly multi-cultural, collegial and collaborative work environment that is rich in development and growth.

Core Qualifications

  • Degree in computer science, data engineering, data science, information systems, bioinformatics, computational biology, engineering, or a related technical field.
  • Experience designing, building, and maintaining data pipelines for structured and unstructured data.
  • Strong Python and SQL skills.
  • Experience with APIs, relational databases, document stores, search indexes, cloud data platforms, and ETL/ELT workflows.
  • Experience handling large volumes of text-heavy documents such as PDFs, slide decks, reports, publications, regulatory files, scientific literature, or knowledge repositories.
  • Strong understanding of data cleaning, normalisation, metadata management, document parsing, indexing, lineage, versioning, and auditability.
  • Familiarity with data modelling approaches for complex knowledge domains, including scientific, clinical, regulatory, or healthcare content.
  • Ability to translate domain expert requirements into practical data structures, metadata models, retrieval-ready content, and maintainable pipelines.
  • Strong attention to detail and ability to identify data quality issues.
  • Ability to collaborate effectively with AI engineers, product managers, COA scientists, software engineers, security stakeholders, legal teams, and commercial teams.
  • Strong documentation skills.

Additional Requirements

  • Experience in life sciences, clinical research, healthcare, regulatory data, scientific publishing, HEOR, clinical outcome assessments, patient-reported outcomes, or medical evidence management is strongly preferred.
  • Experience with clinical trial registries, regulatory labels, HTA reports, scientific literature databases, medical knowledge repositories, or similar evidence sources.
  • Experience with vector databases, embeddings, semantic search, Elasticsearch/OpenSearch, Azure AI Search, Pinecone, Weaviate, Milvus, Qdrant, or similar technologies.
  • Experience with cloud data platforms such as Azure, AWS, or GCP.
  • Experience with document AI, OCR, layout-aware parsing, table extraction, metadata enrichment, taxonomy development, or controlled vocabularies is desirable.
  • Familiarity with ontology development, biomedical terminologies, controlled vocabularies, evidence classification, and structured knowledge representation.
  • Understanding of GDPR, data privacy, intellectual property constraints, licensed content, confidential client data, access control, and secure data handling.
  • Ability to work independently in a remote or hybrid environment while collaborating across global teams.
  • Significant experience leveraging AI tools for work.
  • Fluency in English.