Academic
EMA/SMA Fusion: Evaluating a Crypto Trend-Following Strategy
An MSc algorithmic trading case study evaluating whether a moving-average trend strategy on Bitcoin, Ethereum and XRP beats buy-and-hold over 2020 to 2025. Using a corrected vectorbt backtest with next-bar execution and realistic fees, I found that the strategy's signals carry genuine trend information but that trading costs erase its edge on hourly data. On daily bars it matches buy-and-hold's risk-adjusted return while cutting the maximum drawdown from 80% to 61%, which makes it a risk-management overlay rather than a way to beat the market.
Want the full details? View the complete project on GitHub →
Academic
Predicting High-Volume Purchases with Statistical Machine Learning
MSc Statistical Machine Learning coursework in R predicting whether an online order line will be a bulk purchase of 10 or more units, using 540,000 transactions from a UK gift retailer. I engineered leakage-free product and customer purchase histories, grouped the train-test split and cross-validation by invoice, and compared ten models: logistic regression, a decision tree, a random forest, gradient boosting and a neural network, each with and without PCA. Gradient boosting performed best, reaching an AUC of 0.933 on held-out data, with the customer's buying history as the strongest predictor.
Want the full details? View the complete project on GitHub →
Data Projects
Evaluating a New Recommender System with an A/B Test
An online store's A/B test of a new recommendation system was left unanalysed when its analyst departed. I audited the test and evaluated it with user-level funnel analysis, z-tests with a Bonferroni correction and power analysis. The new system significantly reduced product page views by 14% and improved nothing, but the test was seriously flawed, with a severe sample ratio mismatch and contamination from a concurrent experiment, so I recommended fixing the group assignment and rerunning it.
Want the full details? View the complete project on GitHub →
Data Projects
Book Platform Market Analysis with SQL
A startup planning a reading app needed to understand the book market before designing its product. Using SQL on a PostgreSQL database of 1,000 books, together with their authors, publishers, ratings and reviews, I found that 82% of titles were published after 2000, that Penguin Books leads among full-length titles, that J.K. Rowling is the highest-rated author among widely rated books, and that the most active raters write about 24 reviews each.
Want the full details? View the complete project on GitHub →
Data Projects
E-commerce Product Range Analysis
Analysis of a year of transactions from an online household-goods store to guide decisions on its product range. After cleaning out postage, fees, stock write-offs and cancellations, I found that 22% of products generate 80% of net revenue while over half contribute just 5%, that sales double from spring to November across the whole range, and that the top 10% of customers bring in 61% of revenue while a third never return. The project combines ABC analysis, RFM segmentation and non-parametric hypothesis tests.
Want the full details? View the complete project on GitHub →
Data Projects
Wrangling and Analysing WeRateDogs Twitter Data
Data from three sources on the WeRateDogs Twitter account, a tweet archive, neural-network breed predictions and API engagement counts, gathered, assessed and cleaned into one master dataset. I documented and fixed twelve quality and three tidiness issues, correcting 30 misextracted ratings along the way, and found that the median rating inflated from 10/10 to 13/10 in under two years, that Samoyeds attract the most engagement, and that higher ratings predict more favourites even after controlling for the account's growth.
Want the full details? View the complete project on GitHub →