CV

Senior Principal Researcher & Technical Lead at Huawei Canada, working on LLM post-training and agentic learning. Greater Toronto Area, Canada.

Summary

Research and technical leader with industrial-scale LLM post-training experience across thousands of Ascend NPUs. I lead a team spanning data, environments, SFT/RL and evaluation, and combine post-training research with distributed-systems engineering, with publications at EMNLP, AACL-IJCNLP, ICSE and ASE.

Experience

Huawei Canada, Centre for Software Excellence

Dec 2022 – present

Senior Principal Researcher & Technical Lead, May 2026 – present
Senior Researcher → Principal Researcher, Dec 2022 – Apr 2026

  • Industrial-scale post-training. Lead about 20 researchers and engineers across Pangu code-model SFT and RL for 8B–718B models, coordinating data curation, environments, NPU training and agent evaluation.
  • SFT data curation. Contributed to MindForge, an automated SFT data pipeline spanning source-free environments, teacher rollouts, build-validity filtering, infrastructure-failure recovery and reasoning repair. Fine-tuning Qwen3.6-27B on 973 curated trajectories raised the ProgramBench average test pass rate from 37.98% to 49.51%, with gains on seven more SE benchmarks.
  • End-to-end agent learning. Led RepoForge, integrating repository mining, 7,304 executable environments, teacher trajectories, SFT and RL. Its 8B agent reached 17.4% on SWE-bench Verified, leading the ≤8B non-thinking category at its August 2025 release.
  • Data-centric capability improvement. Co-authored an industrial study that increased usable teacher supervision 2.84× under the same teacher and attempt budget, improving held-out LiveCodeBench v6 pass@1 by 6.11 points and CodeForces by 2.59 points while keeping AIME/MATH regression suites within tolerance.
  • Cross-scaffold generalization. Built trajectory collection, filtering and distillation pipelines across Claude Code, OpenCode and OpenHands; co-authored DCAS on planning-aware fine-tuning that improves performance on scaffolds not seen in training.
  • Data quality. Designed SPICE for issue-clarity, test-coverage and effort labeling, at roughly 19,000× lower cost than estimated manual annotation in a 1,000-instance comparison.
  • Trustworthy evaluation. Designed SWE-agent evaluations across bug fixing, feature implementation, code editing and architecture; co-authored SWE-Effi and When Elo Lies on resource-bounded performance and Codeforces evaluation bias.
  • Distributed systems and open source. Led heterogeneous-computing research across Ascend NPUs and NVIDIA GPUs, including Ray on 10,000 NPUs. The team contributed 50+ upstream pull requests to Ray for cluster scalability, stability and performance.

Huawei Canada

Apr 2020 – Dec 2022

Research Intern → Senior Researcher

Researched reproducible deep learning, build systems and code clones; published in ICSE, TSE, TOSEM and ICSE-SEIP, including first-author work on training reproducible deep learning models.

Baidu, Cloud Testing Group

Aug – Dec 2017

Research Intern

Built a log-based code-coverage estimation prototype evaluated on five industrial projects; published at ASE 2018.

IBM Canada, Platform Symphony

Jan – Aug 2016

Research Partner

Built a Hadoop/Python framework to analyze distributed-system logs and detect logging anti-patterns and problematic message sequences.

Education

York University

Sep 2014 – Oct 2020

Ph.D. (2020) and M.A.Sc. (2017), Computer Engineering

Advised by Zhen Ming (Jack) Jiang. NSERC Canada Graduate Scholarship – Doctoral (CGS-D).

University of Science and Technology of China

Sep 2010 – Jun 2014

B.E., Computer Science

Talks and service

Publications

35 papers and preprints and 2 theses, plus 6 patents and patent applications, listed on the publications page and Google Scholar.

Technical focus

  • Training: SFT; RLVR; multi-turn, execution-feedback RL; teacher-trajectory distillation.
  • Data, evaluation and systems: data curation and labeling; agent evaluation; Ray; distributed training; heterogeneous NPU/GPU computing.
  • Research interests: domain adaptation, continual learning, and learning from software execution feedback.