I completed my PhD in the department of Electrical Engineering and Computer Science at UC Berkeley, advised by Yun S. Song and Nir Yosef. Broadly speaking, I work on mathematical and computational biology. More specifically, I am interested in statistical phylogenetics and its intersection with single-cell biology. Two key areas I am interested in are protein evolution and cancer evolution. Methodologically, my work is themed around finding clever tradeoffs between computational efficiency and statistical efficiency (or accuracy) when analyzing large biological datasets. Indeed, many biological datasets have reached a point where they are too large to analyze with conventional methods, and more clever, scalable methods are needed. Here, by ‘clever’ we mean that the increased scalability comes at a small cost of accuracy. The best example of this is my work on CherryML for rate matrix estimation, where we developed a method that is thousands of times faster than the state-of-the-art, while being only two times less statistically efficient (rather than thousands, which would be trivially achievable by subsampling the dataset). This has enabled fitting deep models of protein evolution for the first time, see our recent preprint Deep models of protein evolution in time generate realistic evolutionary trajectories and functional proteins.
See my Google Scholar page for the most up-to-date list of publications and preprints.