Smishing scam-type classifier
Classifies 34k public smishing messages in 50+ languages into scam types, with an evaluation setup built for the question operations teams ask: how will this do on the next campaign?
Senior Data Scientist · ML Engineer
I take ML from problem framing to production: defining the right metric, building models that hold up on messy, imbalanced real-world data, and measuring whether they actually changed outcomes. Seven years across Trust & Safety, fraud, experimentation, and health AI.
chenghsiuwen.tw@gmail.com
Three threads run through my work.
Fine-grained scam and abuse classification under heavy class imbalance, deepfake detection, and turning model output into decisions that product and policy teams can act on.
A/B testing at scale, difference-in-differences for product changes, and self-service measurement so stakeholders stop waiting on analysts.
CGM glucose time-series prediction, and offline reinforcement learning to personalize when patients wear continuous glucose monitors.
Selected impact from industry and research roles.
Public code. Each repo has a README covering the problem, design choices, and results.
Classifies 34k public smishing messages in 50+ languages into scam types, with an evaluation setup built for the question operations teams ask: how will this do on the next campaign?
Four-person UCLA MSBA project extending a LangGraph demo into an audited multi-agent pipeline for specialty-medicine logistics.
Turns free-text requests into ranked menu items: an LLM extracts constraints into a strict schema, rules filter, and embeddings rank, so hard limits are never violated.
Systematic comparison of 14 CNN and Vision Transformer architectures, contrasting full fine-tuning with partial unfreezing.
Extends LipForensics with LLM-guided selection of lip-sync-friendly video segments. See also CNN deepfake detection.
Bronze → Silver → Gold pipeline turning raw listings into neighborhood market indicators across four US cities.
Health AI at UCLA.
Learning from observational data when a continuous glucose monitor adds the most value for each patient, using fitted Q-iteration and fitted Q-evaluation with bootstrap confidence intervals, plus reward design and wear-cost sensitivity analysis.
Forecasting glucose trajectories from CGM and wearable signals, and cleaning real-world CGM data so downstream models can be trusted.