Systematic Equity/ Quantitative Research/ Applied ML

Fan Zhu

Associate Portfolio Manager at Parametric Portfolio Associates (Morgan Stanley). I work on systematic equity SMAs — portfolio construction, rebalancing, and the Python analytics our team uses to run them day to day.

$9Bn+
Systematic equity SMAs
I help manage
400+
Benchmarks
monitored daily
70%
Credit memo drafting
time saved (prototype)

About

Between the model and the trade.

An optimizer result only becomes a portfolio once it clears tax lots, client restrictions, corporate actions and a transaction cost budget. Closing that gap is most of my job, and it is why I spend about as much time writing code as running analysis.

Portfolio Construction

Building and rebalancing custom core SMAs against client-specific benchmarks, balancing tracking error, tax impact and transaction costs against the constraints in each mandate. Scenario analysis on flows, rebalances and benchmark changes, written up so advisors can act on it.

Quantitative Research

Network analysis and machine learning applied to market structure. My undergraduate honors thesis forecast bull/bear regimes from the topology of industry correlation networks. The habit stuck: look for the structure first, then test carefully whether it predicts anything.

Systems & AI

Python applications used in daily PM workflows, plus the data-integrity checks behind 400+ benchmarks. Selected as an AI Ambassador at Parametric, where I help colleagues adopt AI tools and build their own.

Experience

From risk analytics to portfolio construction.

I started where risk numbers are produced, learned what breaks them, and moved to the desk that has to act on them. Each step has taken me closer to the investment decision itself — from measuring exposure, to building the portfolios that carry it, toward the quantitative research and systematic trading I want to work on next.

May 2025 — Present Atlanta, GA

Associate Portfolio Manager

Parametric Portfolio Associates · Morgan Stanley Investment Management

Construct and rebalance $9Bn+ of systematic equity custom core SMAs against client-specific benchmarks, trading tracking error against tax and transaction costs. Build the PM-facing Python analytics and the data-integrity layer under 400+ benchmarks.

Jun 2024 — May 2025 New York, NY

Associate, Risk COO

Morgan Stanley

Daily Rates and FX sensitivity analysis: risk-engine output aggregated under a delta–gamma approximation to estimate scenario PnL across market environments. Built the Python and SQL pipelines that monitored those sensitivities against limits.

Jun 2022 — May 2024 New York, NY

Analyst, Risk COO

Morgan Stanley

Portfolio analytics and reporting across a $200Bn+ lending portfolio in SQL, Python, R and VBA, including the error-checking routines behind senior management reporting.

Research

Market structure, and what actually predicts.

Work on market structure, machine learning and portfolio construction. The honors thesis below asked whether the topology of an industry correlation network carries a forecastable signal about the market regime that follows it — at short horizons it appears to, over longer ones the signal largely disappears. New work is added here as it is finished.

Honors Thesis · UNC Chapel Hill Economics · May 2020

Market Regime Forecasting using Correlation Networks and Machine Learning Classifiers

Dr. Michael Aguilar & Fan Zhu  ·  Faculty advisor: Dr. Jane Fruehwirth  ·  Graduated with Highest Honors

Historical U.S. equity returns are partitioned into bull and bear markets using a top-down turning-point algorithm. For each regime a correlation network is built over the 49 Fama–French industry portfolios, reduced to a minimum spanning tree, and summarised by its topological features. Those features — degree distribution, in-component degree, Kruskal stress and density, alongside the moments of S&P 500 returns — are fed to a bank of machine learning classifiers under 10-fold cross-validation. Across 16 combinations of estimation window and forecast lag, a Fine Tree on 1-month windows forecasting one week ahead reaches 91.7% accuracy; over longer horizons a Kernel Naive Bayes on 3-month windows at a 12-month lag holds 66.7%.

91.7%
Best accuracy
1 week ahead
0.92
AUC
Fine Tree model
53
Correlation networks
constructed
41 yrs
Daily data
1978–2019

Method

STEP 01

Partition the regimes

Hanna's top-down turning-point algorithm over daily S&P 500 prices, 1978–2019: 30 bull, 26 bear, 5 rallies, 7 corrections.

STEP 02

Correlate the industries

Within each regime window, the Pearson correlation matrix over daily log returns of the 49 Fama–French industry portfolios.

STEP 03

Build the tree

Correlations become Euclidean distances, dij = √(2(1−Cij)), reduced to a minimum spanning tree scored by Kruskal stress-1.

STEP 04

Extract topology

Degree distribution, in-component degree, stress and density per network — plus the standard deviation, kurtosis and range of market returns.

STEP 05

Classify & validate

Trees, KNN, SVM, ensembles and Naive Bayes across 16 window×lag combinations, each under 10-fold cross-validation.

Forecast accuracy by estimation window and lag

Sixteen model configurations, best classifier per cell. The signal is concentrated: only the shortest horizon breaks meaningfully away from chance.

52% 92% accuracy
Data table
Forecast accuracy and best classifier by estimation window and forecast lag
WindowLagAccuracyBest model
1 month1 week91.7%Fine Tree
1 month2 weeks77.1%Cosine KNN
1 month3 weeks62.5%Subspace Discriminant
1 month4 weeks62.5%Naive Bayes
3 months1 month62.5%Fine Tree
3 months3 months52.1%Boosted Tree
3 months6 months58.3%Linear Discriminant
3 months12 months66.7%Kernel Naive Bayes
6 months1 month64.2%Linear Discriminant
6 months3 months54.7%Cubic SVM
6 months6 months58.3%Fine Tree
6 months12 months52.1%Coarse KNN
12 months1 month54.7%Coarse Gaussian SVM
12 months3 months52.8%Coarse KNN
12 months6 months54.2%RUSBoosted Tree
12 months12 months52.1%Coarse KNN

Accuracy under 10-fold cross-validation. Rallies and corrections were excluded from training — too few observations for a valid sample — so the task is a bull/bear binary.

Where the winning model gets it right

Fine Tree, 1-month window, 1-week lag. Four misclassifications out of 48 regimes, split evenly between the two error types.

Data table
ActualPredicted bearPredicted bull
Bear212
Bull223

For context, Kole & Dijk (2010) forecast the same one-week-ahead bull/bear call with Markovian logit models at 89.3%. This model reaches 91.7% with an AUC of 0.92 — on 48 regime observations, so the difference sits well inside the noise.

Honest limitations. The MST minimisation criterion did not fully converge, so the trees may not be the most efficient representation of each network; and several algorithm-partitioned regimes are only days long — artefacts of black-swan events sitting too close to the following rally to be separately forecastable. Both are named in the paper as the first things to fix. With 48 regimes in the sample, the accuracy figures carry wide error bars either way.

Scenario summarisation for PM–advisor communication

Recurring trade and portfolio-analysis scenarios — security flows, rebalances, benchmark reconstitutions, tax-loss activity — summarised with an LLM copilot and clustered into a categorised template library. The work is in the categorisation: which scenarios recur often enough, and are distinct enough from one another, to earn their own template. Portfolio-manager correspondence to financial advisors then starts from a scenario-matched draft carrying the relevant risk and tax diagnostics rather than from a blank page.

LLM Summarisation Scenario Taxonomy PM ↔ Advisor Comms

LLM framework for fundamental credit risk

Co-authored at the Morgan Stanley firmwide GenAI Hackathon: generate text embeddings from 10-K filings, surface the fundamentals that matter to a credit view, and automate the first draft of the credit memo. Drafting time fell by about 70% in the hackathon evaluation.

LLM Embeddings 10-K Filings Credit Memos

Teaching & research assistantships

Teaching Assistant for data science at Duke's Fuqua School of Business and the Department of Statistical Science. Research Assistant at Duke's Center for Advanced Hindsight, a behavioural economics lab.

Duke Fuqua Statistical Science Behavioural Economics

Building

Tools that ended up in the daily workflow.

PM-facing Python applications I have built at Parametric. Small, specific tools that run inside a live investment process and are maintained in shared repositories.

Scenario extraction

Rules-based extraction and validation of trading scenarios from structured advisor notes — regex parsing with error handling, so an ambiguous instruction is flagged rather than silently misread.

PythonRegexValidation

Reconstitution analytics

Benchmark reconstitution impact analysis — adds and drops, resulting turnover, and which SMAs are affected — so rebalance work can be prioritised by impact rather than by list order.

Index EventsTurnoverPrioritisation

PM–advisor email generator

The shipped form of the scenario template library: portfolio-manager correspondence to financial advisors, drafted from copilot-summarised, categorised templates matched to the recurring trade and portfolio-analysis scenarios behind it.

LLM SummarisationTemplating

Restriction consistency

Consistency checks across securities and screening groups using hash-table comparison, catching restriction mismatches before they reach a portfolio that shouldn't hold the name.

Hash TablesComplianceScreening

Trade status reporting

A distributive trade status reporting system with automated scheduling and validation checks, streamlining the communication loop between portfolio managers and traders on what has actually filled.

SchedulingValidationPM ↔ Trader

Market data integrity

Dashboards tracking ingestion across 400+ benchmarks, reconciling corporate actions against Bloomberg, FactSet and Refinitiv. Not glamorous, but every rebalance downstream depends on it.

PythonSQLCorporate Actions

Toolkit

What I reach for.

Languages
PythonSQLR MATLABVBACSS
Portfolio & Risk
Portfolio OptimizationTracking Error Tax-Loss HarvestingTransaction Cost Analysis Delta–Gamma SensitivitiesScenario PnL Benchmark Reconstitution
ML & Statistics
ClassificationEnsemble Methods Network AnalysisCross-Validation LLM EmbeddingsTime Series
Data & Platforms
BloombergFactSetRefinitiv Git / GitHubExcel TableauPower BI

Education

Statistics, economics, and a psychology habit.

/ MAY 2022

MS, Statistical Science

Duke University

Teaching Assistant for data science courses at the Fuqua School of Business and the Department of Statistical Science. Research Assistant at the Center for Advanced Hindsight.

Off the desk

Endurance, mostly.

2025
NYC Marathon
finisher
2024
Spartan Race
Trifecta medalist
2023
Savage Race
Syndicate medalist

Contact

Let's talk.

Happy to talk about systematic equity, quantitative research, or the engineering behind either. Email is the fastest way to reach me.