Mathematical finance arrived late and then moved very fast. Seventy years at the start where the right idea existed and nobody used it. Then a fifteen-year window where most of the theory that matters got built. Then a long period of refinement that is still going, and my reading is that somewhere in it the field stopped receiving new theory and started fitting more parameters to the old kind.

I went through most of this while writing the econophysics review, and reading it in order changed how I think about which parts of the subject are load-bearing, which matters if you are deciding where to spend research time.

Bachelier, 1900

Louis Bachelier defended a thesis called Théorie de la spéculation at the Sorbonne in 1900, supervised by Henri Poincaré. He was trying to price options on French government bonds, and to do it he wrote down a mathematical model of a price that moves randomly in continuous time.

He derived what we now call Brownian motion. Five years before Einstein’s paper on the same process, for a completely different reason, and with a diffusion equation for the transition density. He then used it to compute option values.

Poincaré’s report on the thesis was positive and slightly puzzled. The work went nowhere for fifty years. Part of that is that mathematics had no rigorous foundation for the object Bachelier was manipulating until Wiener constructed the measure in the 1920s. Part of it is that nobody in finance was reading French doctoral theses in probability.

Bachelier’s model has a known defect: prices follow arithmetic Brownian motion, so they can go negative. That is wrong for equities, and the fix, which is to model the logarithm instead, took until the 1960s. The defect matters because it came back. In April 2020 crude oil futures settled below zero, and every desk running lognormal models had to switch to a Bachelier-style normal model in a hurry. A hundred and twenty years later the discarded model turned out to describe the situation better.

Ito, and then Samuelson finding Bachelier again

Kiyosi Ito built the stochastic calculus in the 1940s. The lemma that carries his name tells you how a function of a stochastic process evolves, and it is the reason any of the later work is possible. Without it, writing down the dynamics of a derivative whose underlying is random is not a well-posed operation.

Paul Samuelson came across Bachelier’s thesis in the 1950s, reportedly through a note from Leonard Savage, and did the obvious correction: model the logarithm of the price as Brownian motion, so the price itself is geometric Brownian motion and stays positive. This is the model that sits under everything that follows.

Samuelson could price an option under this model given an assumption about the investor’s risk preferences. That assumption was the obstacle. Different investors, different preferences, different prices, no unique answer.

1973

Fischer Black and Myron Scholes published “The Pricing of Options and Corporate Liabilities” in the Journal of Political Economy, and Robert Merton published “Theory of Rational Option Pricing” in the Bell Journal the same year.

The move that made it work is the hedging argument. Take a short position in the option and a continuously adjusted long position in the underlying, in the ratio that cancels the first-order sensitivity to price moves. The resulting portfolio has no exposure to the direction of the underlying. If it has no risk, it must earn the risk-free rate, or there is an arbitrage.

Set that up, apply Ito’s lemma, and you get a partial differential equation for the option value. Solve it and you get a formula. The investor’s risk preferences do not appear anywhere in it. Neither does the expected return of the underlying, which is the part that surprises people the first time they see it, and which is the whole reason the result is usable.

Cox, Ross and Rubinstein gave the binomial version in 1979, which is how most people meet the argument, because in discrete time the replication is visible rather than buried in a limit.

Harrison and Kreps in 1979 and Harrison and Pliska in 1981 then produced the general statement. Absence of arbitrage corresponds to the existence of an equivalent martingale measure. Completeness of the market corresponds to that measure being unique. Pricing becomes taking an expectation under a measure that is not the real-world one. That is the deepest layer of the theory and it is where the subject stops being about options and starts being about the structure of markets.

The smile, and the long era of patching

Black-Scholes assumes constant volatility. Invert the formula on real option prices to recover the volatility implied by each one and you should get the same number for every strike.

Before October 1987 you roughly did. After the crash you did not, and you have not since. Plot implied volatility against strike and you get a smile or a skew, which is the market saying that the lognormal model underprices large moves, and that it underprices downward moves more than upward ones.

Everything after this is a response to that observation.

Merton had already added jumps in 1976, with a Poisson process superimposed on the diffusion, which produces fat tails directly. Heston in 1993 made volatility itself a mean-reverting stochastic process correlated with the price, and found a semi-closed form, which is why it became the standard. Dupire in 1994 went the other way and asked what deterministic local volatility surface would exactly reproduce today’s observed option prices, which fits perfectly by construction and says nothing about tomorrow. Then SABR, then variance gamma, then a long list.

Each of these adds parameters and fits better. It is worth being honest that this is a different activity from what happened in 1973. Black, Scholes and Merton produced a result where risk preferences cancelled out. Most of what followed produced flexible families of models with enough freedom to match observed prices. Commercially that was enough, because a desk needs a surface it can quote from, not an explanation.

Microstructure, which is where the interesting theory went

The other branch stopped treating price as a given process and asked where it comes from.

Kyle’s 1985 model has an informed trader, noise traders and a market maker who cannot tell them apart, and derives how information gets into the price and how much the informed trader can extract before it does. Glosten and Milgrom in the same year derived the bid-ask spread as the market maker’s compensation for adverse selection, which explains why the spread exists at all without appealing to costs.

Almgren and Chriss around 2000 formalised optimal execution: you have a large order, trading fast costs you impact and trading slow exposes you to volatility, and there is a frontier. Avellaneda and Stoikov in 2008 did the market maker’s version, deriving optimal quotes as a function of inventory and time. Hawkes processes came in for order flow, because order arrivals cluster and self-excite in a way Poisson arrivals cannot capture. Bacry, Muzy and others built that out.

Gatheral, Jaisson and Rosenbaum in 2014 showed that volatility measured at high frequency behaves like a fractional Brownian motion with Hurst exponent around 0.1, which is far rougher than the standard models assume and rougher than anyone expected. That is a genuinely new empirical fact about markets and it is one of the few results in recent memory that forced a change in the modelling rather than an addition to it.

Where this leaves the subject

My reading is that the era of new foundational theory in derivatives pricing ended somewhere in the 1980s, and the era of new theory about market structure is still going, and that these two facts are related. Pricing theory had a clean question with a clean answer. Microstructure has messy questions and the data to attack them, which is a better place to be right now.

The part I work on sits in the second branch. Whether topological invariants of order book graphs carry information about regime transitions is a microstructure question, not a pricing one, and it exists because the data exists.

What I take from the history: the results that lasted are the ones where something cancelled. The expected return dropping out of Black-Scholes. Risk preferences dropping out of the martingale formulation. The spread appearing from adverse selection alone without assuming any costs. When a model needs more parameters to fit better, that is a signal about the model, and it is usually a signal that the structure has not been found yet.

Which is the bet. The edge is in structure nobody has found, not in parameters fitted to prices everybody can already see.

Related: