I am a Data Analyst and AI & Data Science professional focused on turning complex operational and healthcare datasets into clear, actionable insights for decision‑makers. I specialise in Python, SQL, Power BI and Excel, building end‑to‑end analytics workflows from data cleaning and modelling through to dashboards and narrative reporting. Recent projects include healthcare billing analytics, gender inequality index modelling, and biomedical text curation, all designed to support evidence‑based, public‑sector‑style decision making.
When I'm not working with data, I enjoy reading, playing piano, watching and playing football. I also enjoy playing chess and scrabble. I love the "aha!" moment when data reveals something new and useful.
Global Gender Inequality Index analysis with full data pipeline (cleaning, feature engineering, visualisation) and an interactive Streamlit app for scenario exploration and policy support.
Skills: Python, Pandas, Scikit‑learn, Streamlit, Data Visualisation, Statistical Analysis
Interactive Power BI report on 50K+ patient records, highlighting billing KPIs, chronic conditions and insurance patterns to inform resource allocation and service planning.
Skills: Power BI, Power Query, DAX, KPI Design, Healthcare Analytics
Ensemble machine learning models (Random Forest, XGBoost) for heart disease prediction, with thorough EDA, model evaluation (ROC‑AUC, F1), and a web GUI for non‑technical users.
Skills: Python, Scikit‑learn, XGBoost, Model Evaluation, Streamlit, Data Storytelling
Curated and documented a biomedical abstracts dataset (5,901 records) with deduplication, label taxonomy, and PubMed‑linked provenance to enable reproducible text‑classification research.
Skills: Data Cleaning, NLP, Corpus Design, Provenance Tracking, Reproducible Workflows
- 13+ public data & analytics projects across Python, Power BI and ML
- 3+ deployed web apps using Streamlit and AI chatbot frameworks
- Healthcare, inequality and business analytics domains with end‑to‑end pipelines (data cleaning → modelling → dashboards → narrative reports)
- Project A:
Gender Inequality Index (GII): Global Gender Inequality Index analysis using Python, EDA, and interactive visualisation to identify key inequality drivers.
- Project B:
Healthcare Machine Learning Research: Uncertainty-aware machine learning models for healthcare risk prediction and ethical clinical decision support.
- Advanced data analysis, machine learning, and model evaluation techniques using Python and Scikit-learn, with a focus on healthcare and social impact datasets.
- Applied data science workflows through the DataCamp Data Science Track, strengthening skills in EDA, SQL, and data storytelling.
- Python programming for data science via Programming with Mosh, deepening understanding of clean code, automation, and reproducible analysis.
- Practical AI skills and ethical AI deployment through a 19-week AI Skills Bootcamp with Breakthrough Enterprise, focused on inclusive technology and real-world applications.