I build and evaluate large language models for healthcare. At M42 in Abu Dhabi I lead work on the Med42 family of open-weight clinical LLMs, covering the full pipeline — continuous pretraining, instruction tuning, preference alignment — and the evaluation and safety frameworks used to decide whether any of it is fit for clinical use.

Much of my recent work argues the same point from different angles: benchmark scores are a poor proxy for clinical reliability. That has led to research on evaluation-framework variability, dataset transparency and bias, sycophancy under authoritative pressure, and multilingual clinical care in Arabic.

Before healthcare I spent six years at EDF R&D, where I also completed a PhD on detecting novelty in textual data streams as early as possible.

Selected work

Med42 — open-weight clinical LLMs

A suite of clinical LLMs built on Llama 2 and Llama 3, adapted with specialised clinical data and multi-stage preference alignment. The v2 models outperform their Llama 3 base counterparts and GPT-4 across standard medical benchmarks.

Med42 paper Med42-v2 paper Llama3-Med42-70B Llama3-Med42-8B med42-70b

MEDIC — evaluating LLMs in clinical applications

A framework that assesses clinical LLMs across five dimensions — reasoning, ethics and bias, data and language understanding, in-context learning, and clinical safety — rather than collapsing everything into a single benchmark score.

Paper Leaderboard

BioToken & BioFM — genomic foundation models

A tokenization framework that encodes genetic variants and structural annotations — coding regions, transcript boundaries — directly into genomic representations, instead of treating DNA as plain linear text. The resulting 265M-parameter model stays competitive with genomic foundation models an order of magnitude larger.

Paper BioFM-265M

Selected publications

Full list on Google Scholar.

Experience

Education