My name is Mike Nguyen. My research interests include
Download my resume in PDF and Word .
Postdoc in Marketing
University of Southern California
PhD in Marketing - minor Statistics
University of Missouri
MA in Economics - Econometrics and Quantitative Economics
University of Missouri
MBA in Marketing Analytics, Corporate Finance
University of Delaware
BBA in Marketing, International Business
Florida International University
R, SAS, STATA, SPSS
Python, NetLogo, Gephi
NEO4j, MongoDB
Write a data frame to CSV with base R, readr, or data.table, read it back with the same package, and each one destroys values in at least four of eleven columns, even when the reader is allowed to know the original schema. Leading zeros, the last digits of 18 digit IDs, the Namibian country code, the final bits of most doubles, milliseconds, and an hour on the night the clocks go back all go missing somewhere. Parquet written with arrow came back identical in every column, and pandas read the same file exactly. data.table was still the fastest, so the case for Parquet is correctness rather than speed.
On the CDNOW panel, with a real 39 week holdout as ground truth, a Pareto/NBD model flags 32 percent of the top decile by historical spend as probably inactive. The median flagged customer then buys nothing at all, against a median of 130 for the rest of the decile. Extrapolating repeat spending overstates the holdout by 35 percent and the model undershoots it by 16, so the level is not where the argument is strongest. Fitted in R with CLVTools, with a bootstrap confidence interval computed in Python.