Publications
Papers and preprints on LLM post-training for software engineering, AI systems, and empirical software engineering. The full, always-current list is on Google Scholar.
2026 12
TOSEM 2026Journal
Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap
A vision of AI-native, intent-first software engineering with AI teammates, and the technology roadmap to get there.
EMNLP 2026Industry Track, accepted
LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering
Treats industrial post-training as maintaining a deployed checkpoint through budgeted data-mixture patches, and distills what makes that hard.
ASE 2026Industry Showcase
DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds
An interception layer that decouples CLI coding-agent scaffolds from models; planning learned from Claude Code trajectories transfers to OpenCode and mini-swe-agent.
arXiv 2026Preprint
SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements
A benchmark of 188 tasks from merged pull requests for evaluating coding agents on behavior-preserving, non-functional improvements.
arXiv 2026Preprint
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch
Makes behavioral-specification elicitation an explicit first step before code synthesis when agents build programs from scratch.
arXiv 2026Preprint
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
Turns open-source CLI programs into source-free environments for whole-life-cycle program synthesis; fine-tuning Qwen3.6-27B lifts the ProgramBench pass rate from 37.98% to 49.51%.
AACL-IJCNLP 2026Main conference
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment
Agentic LLM judges curate data for architecture-aware fine-tuning; Qwen3 models trained on 3,360 instances reach up to 27.2% on SWE-bench Verified.
arXiv 2026Preprint
SynConfRoute: Syntax-Aware Routing for Efficient Code Completion with Small CodeLLMs
Training-free routing that combines token confidence with syntax checks to decide when a small local code model is good enough, cutting accelerator use by 58%.
ICSE 2026Technical briefing
Software Engineering for Foundation Models (SE4FM)
Companion paper for the ICSE 2026 technical briefing on software engineering for foundation models.
TOSEM 2026Journal
Software Performance Engineering for Foundation Model-Powered Software
Four performance-engineering challenges for software built on foundation models, from cognitive architecture design to deployment, with research directions.
ASE 2026Industry Showcase
When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models in Practice
Shows how submission order, contest selection and run-to-run variance bias Codeforces Elo ratings of LLMs; submission order alone shifts scores by 394 points.
EMNLP 2026Main conference
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
SemanticSpec verifies whole semantic steps instead of single tokens by probing hidden states, for up to 2.7× faster decoding on DeepSeek-R1-32B.
2025 7
ASE 2025Research track
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
Labels SWE-bench-style instances for issue clarity, test coverage and effort, cutting the cost for 1,000 instances from about $100,000 to $5.10.
ASE 2025Industry Showcase
Context-Aware CodeLLM Eviction for AI-assisted Coding
Context-aware model eviction for self-hosted, multi-model code-LLM serving under limited accelerator memory.
arXiv 2025Preprint
SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
Effectiveness metrics that weigh SWE-agent resolve rates against the tokens and time they consume, used to re-rank AI systems on SWE-bench.
arXiv 2025Technical report
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
Generates 7,304 executable environments from GitHub commits and trains with SFT and RL; RepoForge-8B-Agent reaches 17.4% on SWE-bench Verified.
KDD 2025Tutorial
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model Powered Software (FMware)
Tutorial on the challenges and technology roadmap for production-ready, trustworthy software built on foundation models.
arXiv 2025Preprint
Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees
SLA-aware dynamic batching for self-hosted code-LLM serving, with up to 26% higher goodput and 45% lower latency variability.
arXiv 2025Preprint
SLA-Awareness for AI-assisted coding
A runtime that serves mixed coding tasks with different latency targets while keeping cluster utilization high.
2024 2
FSE 2024Industry track
Rethinking Software Engineering in the Era of Foundation Models: A Curated Catalogue of Challenges in the Development of Trustworthy FMware
Ten challenges that make enterprise FMware development risky, and FMArts, Huawei Canada's platform for engineering trustworthy FMware.
arXiv 2024Preprint
Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers
Argues for goal-driven AI pair programmers instead of task-driven copilots.
2023 1
ICSE-SEIP 2023Software Engineering in Practice
An Empirical Comparison on the Results of Different Clone Detection Setups for C-based Projects
Compares how different clone-detection setups change the results on C-based projects.
2022 4
TSE 2022Journal
An Experience Report on Producing Verifiable Builds for Large-Scale Commercial Systems
A process and toolkit for producing verifiable builds, evaluated on three large commercial systems at Huawei.
ICSE 2022Technical track
Towards Training Reproducible Deep Learning Models
Reproducibility criteria plus record-and-replay and profile-and-patch techniques that make deep-learning training reproducible despite software and hardware non-determinism.
ICSE-SEIP 2022Software Engineering in Practice
Towards Build Verifiability for Java-based Systems
A systematic approach to verifiable builds for Java, applied to 46 Reproducible Central projects and 13 open-source projects used in Huawei products.
TOSEM 2022Journal
Towards a Consistent Interpretation of AIOps Models
Studies whether interpretations of AIOps models stay consistent across models, data and time.
2021 2
arXiv 2021Preprint
Can I use this publicly available dataset to build commercial AI software? – A Case Study on Publicly Available Image Datasets
Assesses potential license violations when commercial AI software is built from publicly available datasets.
CSUR 2021Journal
A Survey of Software Log Instrumentation
A survey of research on software log instrumentation: logging approaches, logging utilities and logging-code quality.
2020 2
Ph.D. thesisYork University
Improving the Logging Practices in DevOps
Automated approaches to improve logging practices on both the development and operations sides of DevOps.
ICSE 2020Technical track
Studying the Use of Java Logging Utilities in the Wild
A study of 3,856 logging utilities across 11,194 Java projects on GitHub: why projects use multiple and custom loggers.
2019 3
ASE 2019Industry experience report
An Industrial Experience Report on Performance-Aware Refactoring on a Database-Centric Web Application
Seventeen performance anti-patterns and refactorings applied to an industrial, database-centric web application.
EMSE 2019Journal
Extracting and studying the Logging-Code-Issue-Introducing changes in Java-based large-scale open source software systems
Extracts and studies logging-code-issue-introducing changes in six large Java systems; existing detectors catch only 3% of them.
ICSE 2019Doctoral symposium
Improving the Software Logging Practices in DevOps
Outlines doctoral research on improving logging practices across development and operations.
2018 1
ASE 2018Research track
An Automated Approach to Estimating Code Coverage Measures via Execution Logs
LogCoCo estimates code-coverage measures from execution logs, evaluated on industrial projects at Baidu and on open-source projects.
2017 3
M.A.Sc. thesisYork University
Characterizing and Improving Logging Practices in Java-based Open Source Software Projects - A Large-scale Case Study in Apache Software Foundation
Characterizes and improves logging practices in Java-based Apache projects; the basis of the EMSE 2017 paper.
ICSE 2017Technical track
Characterizing and Detecting Anti-Patterns in the Logging Code
Six logging-code anti-patterns derived from 352 changes in ActiveMQ, Hadoop and Maven, detected with the LCAnalyzer tool.
EMSE 2017Journal
Characterizing logging practices in Java-based open source software projects - a replication study in Apache Software Foundation
Replicates Yuan et al.'s logging-practice study on 21 Java projects from the Apache Software Foundation.
Patents 6
US 2026/0133843 A1Application, published 2026
Systems and Methods for Service Level Agreements for Foundation Model Applications (SLA-aware scheduler)
US 2026/0133848 A1Application, published 2026
Systems and Methods for Service Level Agreements for Foundation Model Applications (SLA-aware resource provisioner)
US 12,541,446 B2Granted 2026
Method and Apparatus to Trace and Visualize Data Movement
US 12,468,637 B2Granted 2025
Method and Apparatus for Providing Artificial Intelligence Model Swapping to Support Foundation Models
WO 2023/060525 A1PCT application, 2023
Methods and Systems for Generating Verifiable Software Releases
WO 2023/028996 A1PCT application, 2023
Methods and Devices for Ensuring the Reproducibility of Software Systems