%20(5).png)
We are often asked how we infer the direction of a trade reported under the MiFID regime. This week we highlight our high level methodology for our initial model and how it can be used to gain even greater insight into the data.
Currently there is no reported direction or side within trade reports submitted under the MiFID regime, which contrasts the approach taken by FINRA in the US, which requires broker-dealers to report if they bought or sold securities on TRACE.
Therefore in order to calculate approximate buy and sell volumes for liquidity providers on a given day, we must infer the direction of each trade.
At a high level, the process involves considering each executed price and then determining whether that was closer to the bid or ask price at that moment in time, indicating that to be a dealer’s buy or sell trade.

Directionality over the short and long terms is a key indicator of momentum and investor activity in a market.
The above diagram provides a high-level, simplified view of the steps involved to gather and process the data in order to make inference.
Whilst in the absence of regulatory mandated reporting of the trade direction it is impossible to be 100% sure that an inferred direction is correct, by fine tuning this process we can significantly increase the probability of estimating correctly.
Our first example takes a liquid Spanish Government Bond (SPGB) and follows the process outlined on the previous page.
We obtain bid/offer quote data as well as trade report data and then merge the two datasets.
Once merged we calculate the mid-price at the time of each observable trade. We then create a rolling average for both the bid and offer, and from there we can infer which side the trade was closest to.

This description is an over-simplification of the process, and while correct at a high level, various data science techniques and multiple layers of testing go into the final output to improve its accuracy.
Chart 1 above highlights intraday activity using this methodology for a liquid government bond, which has a relatively frequent set of data points, but over the page we take a look at a less liquid bond and see if that presents us with any different challenges.
For this example we turn our attention to France and focus on an inflation-linked issue, which is far less liquid (and therefore has far fewer data points) than a conventional government bond.

Whilst data exists for less liquid bonds, the lack of available data points can make life difficult when considering standard data science techniques. For example machine learning often requires a vast number of rows in order to achieve reasonable predictions.
This does not make it an impossible problem to solve, it just makes the job significantly more difficult! As we refine this early methodology, we will be able to increase the accuracy of our inferences to ever greater degrees.
“It’s definitely a challenge to create a reliable method for inferring direction and is important to note no solution can ever be 100% perfect. At Propellant Digital, however, we are keen to extract maximum value from the MiFID data and we are very excited by the challenge ahead on this project. Please contact us to get involved in developing this service further.”