Senior Data Scientist · ML Engineer

Coco Cheng

I take ML from problem framing to production: defining the right metric, building models that hold up on messy, imbalanced real-world data, and measuring whether they actually changed outcomes. Seven years across Trust & Safety, fraud, experimentation, and health AI.

What I work on

Three threads run through my work.

Trust & Safety / Fraud

Fine-grained scam and abuse classification under heavy class imbalance, deepfake detection, and turning model output into decisions that product and policy teams can act on.

Experimentation & causal inference

A/B testing at scale, difference-in-differences for product changes, and self-service measurement so stakeholders stop waiting on analysts.

Health AI

CGM glucose time-series prediction, and offline reinforcement learning to personalize when patients wear continuous glucose monitors.

Experience highlights

Selected impact from industry and research roles.

Trend Micro
Senior Data Scientist
  • Built a multi-layer transformer classifier (English and Japanese) sorting scam SMS into 19 sub-categories under severe class imbalance, deployed for daily inference.
  • Fixed deepfake-detection false positives, cutting online user complaints from thousands to dozens.
  • Reported scam trends to the product lead and worked across sales, product and engineering.
Far EasTone Telecommunications
Senior Data Scientist
  • Scaled experimentation and campaign measurement for 200+ product and marketing stakeholders with automated statistical testing and self-service dashboards, worth about $430K per year.
  • Used difference-in-differences to measure the effect of registration-funnel changes.
TikTok
Data Scientist Intern, Integrity & Safety · 2026
  • Worked cross-functionally with legal, engineering and product on integrity problems.
UCLA
Researcher · MSBA, Anderson (2026)
  • Research on offline reinforcement learning for personalized CGM wear scheduling: off-policy evaluation with bootstrap confidence intervals, reward design, and an EHR + CGM preprocessing pipeline.

Selected projects

Public code. Each repo has a README covering the problem, design choices, and results.

Trust & Safety · NLP

Smishing scam-type classifier

Classifies 34k public smishing messages in 50+ languages into scam types, with an evaluation setup built for the question operations teams ask: how will this do on the next campaign?

Template-aware splits reveal a 3-point leakage gap (0.829 → 0.800 macro-F1). Confidence routing auto-labels 60% of traffic at 96.8% accuracy.
MinHash LSHcluster bootstrapscikit-learnCI
View on GitHub →
Agents · Team project

Multi-agent dispatch QA

Four-person UCLA MSBA project extending a LangGraph demo into an audited multi-agent pipeline for specialty-medicine logistics.

I built the deterministic AuditAgent (7 rules; failed plans loop back to the planner), the what-if ScenarioAgent, audit routing, a Gemini backend with an LLM call budget, and tests.
LangGraphRAGGeminiStreamlit
View on GitHub →
LLM · Recommendation

LLM drink recommender

Turns free-text requests into ranked menu items: an LLM extracts constraints into a strict schema, rules filter, and embeddings rank, so hard limits are never violated.

Graded NDCG 0.989 on training queries, with caffeine thresholds tuned on a held-out split.
structured outputsembeddingsNDCG
View on GitHub →
Computer vision

Food-101 architecture benchmark

Systematic comparison of 14 CNN and Vision Transformer architectures, contrasting full fine-tuning with partial unfreezing.

ConvNeXt-Base reached 87.9% top-1 with full fine-tuning.
PyTorchConvNeXtViT
View on GitHub →
Trust & Safety · Video

Lip-sync deepfake detection

Extends LipForensics with LLM-guided selection of lip-sync-friendly video segments. See also CNN deepfake detection.

Documents each preprocessing iteration, lifting F1 from 23% to 58% on a small demo set of real videos.
video forensicsGeminiPyTorch
View on GitHub →
Data engineering

Airbnb market intelligence pipeline

Bronze → Silver → Gold pipeline turning raw listings into neighborhood market indicators across four US cities.

Geospatial joins with Apache Sedona, orchestrated in Airflow, served from Snowflake to a Tableau dashboard.
SparkAirflowSnowflake
View on GitHub →

Research

Health AI at UCLA.

Offline RL for personalized CGM wear scheduling

Learning from observational data when a continuous glucose monitor adds the most value for each patient, using fitted Q-iteration and fitted Q-evaluation with bootstrap confidence intervals, plus reward design and wear-cost sensitivity analysis.

CGM glucose time-series prediction

Forecasting glucose trajectories from CGM and wearable signals, and cleaning real-world CGM data so downstream models can be trusted.

Toolkit

ML
PyTorch · Hugging Face · scikit-learn · time-series forecasting (LSTM) · offline RL (FQI, FQE) · LLM structured outputs · LangGraph
Data & MLOps
Spark · Airflow · Snowflake · Databricks · FastAPI · Docker · GitHub Actions
Analysis
A/B testing · causal inference (DiD) · bootstrap inference · Tableau · Looker Studio
Languages
English · 繁體中文