Develops intelligent methods for discovering, completing, and querying complex tabular data at scale. This line of work addresses the full analytical pipeline over large-scale data lakes — from locating relevant tables and recovering missing values, to understanding semi-structured table layouts and executing accurate numerical reasoning. A central design principle is cost-effectiveness: leveraging budget-aware LLM selection to maximise analytical accuracy under real-world resource constraints.

Tabular Data Discovery Missing Value Imputation Numerical Table QA Semi-structured Tables Data Lakes LLM Routing

Develops AI-driven systems that transform heterogeneous clinical data into structured, actionable knowledge. This research spans the full spectrum of clinical data understanding — from learning robust patient representations and identifying clinically similar patients, to integrating medical knowledge for precise question answering and diagnostic reasoning. The overarching goal is to make clinical data more accessible, interpretable, and useful for both researchers and practitioners.

Medical Question Answering Clinical Representation Learning Patient Similarity Search Electronic Health Records