Yifei Ding

Scientist II, Experimentation Data Science
Uber

Yifei Ding

Reliable experiments. Better decisions.

I build statistical methods and computation systems for large-scale experimentation. As part of Uber’s Experimentation Data Science team, I develop and advance the statistical methodology and computation behind our A/B and marketplace experimentation platforms. My work spans theoretical development, simulation-based validation, and production implementation.

My research combines machine learning and econometrics to study causal effects, individual heterogeneity, and endogeneity.

Background

I received my Ph.D. in Economics from the University of California, Riverside in September 2025. Before that, I earned an M.A. in Applied Statistics (2018), and a B.A. in Mathematical Finance and B.S. in Chemistry (2015), from Xiamen University.

My experience also includes an applied research scientist internship at Snap.

Experimentation at Uber

Scientist II · Experimentation Data Science · Sept. 2025–present

Full CV

Reliable inference · Citrus / ABPkg

Making results for rare-event metrics more reliable

~45%increased to~92%

Empirical confidence-interval coverage rate in validation tests.

Owned the inference correction end to end: methodology development and validation, code implementation, testing, and production launch.

Potential business applications

App crash rates, reservation metrics in email-campaign experiments, fraud-report rates, and delivery-defect rates—supporting more reliable product, marketing, and risk decisions across Uber.

Marketplace experimentation · MxPkg

Joint switchback analysis

Launched joint analysis that combines evidence from more cities to improve statistical power. Owned theoretical extensions, verification, implementation, testing, and the production release.

Theory → validation → implementation → launch

Statistical foundations

Rewrote and upgraded Citrus’s complete computation methodology reference (V2), providing company-wide guidance on A/B foundations, result computation, and interpretation.

Company-wide science guidance

Advise experimentation users through science on-call: experiment design, power analysis, configuration, debugging, result interpretation, and methodological questions.

2024

Deep Learning for Individual Heterogeneity with Generated Regressors by Adversarial Training

Yifei Ding and Ruoyao ShiWorking paper · Dissertation research

Abstract: Deep Learning for Individual Heterogeneity with Generated Regressors by Adversarial Training

We propose a semiparametric framework that combines machine learning with generated regressors via control function to capture individual heterogeneity while addressing endogeneity and sample selection bias in complex econometric models. This approach models individual heterogeneity through high-dimensional observable characteristics, with generated regressors supporting the control function to manage endogeneity or sample selection bias flexibly across various economic structures. Leveraging a tailored deep learning architecture, our framework integrates control functions and parameter functions seamlessly, enabling its adaptation to diverse econometric models. Using adversarial training, we achieve sup-norm convergence rates of parameter functions and control function at the optimal min-max rate, which enhances robustness and yields valid inferences for inferential structural parameters in high-dimensional settings. Extending the Double Machine Learning (DML) approach, we incorporate endogenous components and establish a new influence function that directly includes generated regressors, broadening the framework’s applicability across econometric models. With automatic differentiation in PyTorch, the influence function applies directly to data, streamlining inference and supporting various structural parameters without additional calculations. This integration makes the framework particularly useful in applied settings where individual heterogeneity and endogeneity are critical, such as personalized policy-making, targeted economic interventions, and customized optimizations in technology. Our simulations demonstrate superior performance, validating this framework’s practical use in econometric analysis where heterogeneity and endogeneity are key considerations.

2023

Comparing Methods for Continuous Treatment

Yifei Ding and Meng XuConference paper · CODE@MIT

Abstract: Comparing Methods for Continuous Treatment

This paper presents a comparative study of two advanced methodologies for estimating the effects of continuous treatments on outcome variables in large-scale tech applications. We focus on dose-response curves and marginal effects to address various business scenarios, such as the impact of ad frequency on user conversions, geolocation campaigns on local engagement, and latency on app performance. Our investigation centers around two promising approaches: entropy balancing for continuous treatment and double/debiased machine learning (DML). Using semi-synthetic data based on Snapchat user behavior, we evaluate these methods’ performance in terms of scalability, flexibility, and precision in handling high-dimensional, non-linear relationships between outcome variables, continuous treatments, and confounders. The study finds that tree-based machine learning models, particularly XGBOOST and Boostsmooth, outperform balancing approaches in estimating dose-response curves, while the balancing method performs best for marginal effect estimation. Notably, our findings also challenge the efficacy of the kernel-based selection model in the double machine learning process, prompting a reconsideration of its utility in real-world applications.

Does “Too-Connected” Network of Shareholders Exacerbate Crash Risk?

Haiqiang Chen, Yang Chen, Yifei Ding, and Muqing SongPublished · 2023(03), 1070–1087China Economic Quarterly (经济学季刊)

Abstract: Does “Too-Connected” Network of Shareholders Exacerbate Crash Risk?

Using quarterly data from the top 10 largest shareholders of A-share stock markets from 2003 to 2018, we construct a network of influential shareholders. Our findings reveal that firms with more interconnected shareholders face higher crash risk, especially when dominated by financial institutional shareholders or those with higher shareholding ratios. In contrast, state ownership and robust corporate governance significantly mitigate this risk. Mechanism analysis shows that firms with higher network centrality tend to have a higher goodwill-to-market value ratio, a greater proportion of related-party transactions to total assets, and larger M&A premiums, yet exhibit lower corporate governance transparency. These results suggest that overly connected shareholder networks may encourage tunneling behavior, exacerbating the crash risk for listed companies.

2022

Abstract: A Comparative Study of Machine Learning Models for Prediction: Insights from Tree-Based Models and Deep Neural Networks

The growing influence of machine learning (ML) and big data technologies has significantly reshaped many scientific disciplines, including econometrics. This paper conducts a detailed comparative analysis of various tree-based and deep learning models, focusing on their prediction capabilities. The models examined include neural networks (e.g., MLP, ResNet), and several advanced tree-based models (e.g., Boost-Smooth, SMARTboost and Random Forest). Additionally, we explore different prediction combination techniques to evaluate whether combining predictions from multiple models enhances predictive accuracy. Using simulations from the comprehensive data generating processes (DGP), we systematically compare the performance of these models under varying levels of noise and the presence of irrelevant features. Our findings reveal that tree-based models like SMARTboost and BooST demonstrate robust performance, particularly in low signal-to-noise scenarios, where they often outperform neural networks. Moreover, the inclusion of ensemble methods, such as median and simple average combinations, further improves prediction stability. Two real-world economic applications—Engel curve prediction and stock price crash risk prediction—highlight the practical implications of our analysis, showing the advantages of tree-based methods in capturing both linear and nonlinear data structures, while DNNs struggle in noisy and nonlinear environments. Our study emphasizes the need for careful model selection and the potential benefits of hybridizing prediction models for complex data tasks.

Abstract: Estimating Partial Effects Using Machine Learning

In this paper, we explore the use of machine learning techniques for estimating partial derivatives, which is a critical step towards understanding causal relationships in econometric analysis. By leveraging modern machine learning methods, such as tree-based models and deep neural networks, we assess their effectiveness in recovering regression functions and estimating partial derivatives. We introduce a novel tree-based model, Boosting Smooth Transition Regression Trees (BooST), and compare its performance with other models, including Boosting of Symmetric Smooth Additive Regression Trees (SMARTboost) and deep neural networks (DNNs). Simulations, based on the well-known Friedman data generating process (DGP), demonstrate the superiority of BooST in estimating partial effects across various signal-to-noise environments and in the presence of redundant variables. The empirical applications, including the study of Engel curves, further highlight the ability of BooST to outperform other machine learning models in accurately estimating partial derivatives. Our findings suggest that BooST provides a powerful tool for nonparametric regression and causal inference, especially in econometric contexts where accurate estimation of marginal effects is crucial.

Earlier experience

Teaching

Snap

June–Sept. 2023

Applied Research Scientist Intern

Continuous-treatment causal inference on 20M+ Snapchat observations, smooth tree-based models for marginal-effect estimation, and ensemble methods for stable prediction.

Related research

Xiamen University

June 2016–Aug. 2017

Research Assistant

Threshold models, financial-network modeling, and shareholder connections and stock-price crash risk, with web-scraped data and simulations in R and NetLogo.

Contact

For questions about my work or opportunities in experimentation and applied science, please get in touch.

yding067@ucr.edu