Skip to content
View mazen-gebrel's full-sized avatar

Block or report mazen-gebrel

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please donโ€™t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this userโ€™s behavior. Learn more about reporting abuse.

Report abuse
mazen-gebrel/README.md

Hi there ๐Ÿ‘‹, I'm Mazen

Data Analyst & Data Engineer

I build end-to-end data solutionsโ€”from engineering automated ETL pipelines to designing business-driven analytics dashboards. I specialize in turning messy, unstructured data into scalable systems and actionable insights.

This is my new profile (old profile) where I share projects that interested me to develop locally.

๐Ÿ› ๏ธ Tech Stack

Languages & Core Data MySQL Python Pandas

Data Engineering & Cloud Dagster

BI & Data Visualization Tableau Power BI

Product Analytics & Tracking Firebase Tealium

Version Control & CI/CD GitHub GitHub Actions


๐Ÿš€ Featured Projects

An air-gapped Retrieval-Augmented Generation pipeline to chat with confidential PDFs locally.

  • The Build: Decoupled architecture utilizing LangChain, HuggingFace embeddings, and a persistent local ChromaDB for vector storage.
  • The Inference: Routes context to an open-source LLM (DeepSeek/Mistral) hosted locally via LM Studio, guaranteeing zero data leakage to external APIs. Paired with a Streamlit chat UI.

An end-to-end, automated data pipeline engineered to track technical skill demand.

  • The Build: Extracts live API data, cleans unstructured HTML using Regex/Pandas to flag specific skills, and loads it into a persistent SQLite database.
  • The Automation: Fully automated via GitHub Actions cron jobs to run weekly.

An end-to-end Machine Learning web application that predicts user cancellation risk.

  • The Build: Engineered a complete scikit-learn pipeline featuring automated scaling, one-hot encoding, and a tuned Random Forest Classifier, serialized for production.
  • The UI: Deployed an interactive Streamlit frontend that allows Customer Success teams to input user metrics and receive real-time churn probabilities and retention recommendations.

A deep-learning application that extracts structured text and timestamps from raw audio.

  • The Build: Implemented OpenAI's Whisper model locally, featuring dynamic compute allocation (swapping model weights based on hardware limits) and forced language mapping for complex dialects.
  • The UI: Engineered a Streamlit frontend that caches ML weights in RAM for performance and outputs structured timestamp DataFrames ready for downstream NLP analysis.

A full-stack, interactive analytics application that segments users based on purchasing behavior.

  • The Build: Processes synthetic transaction data to calculate complex Recency, Frequency, and Monetary (RFM) quantiles.
  • The UI: Features a custom-themed Streamlit interface with high-contrast light/dark modes and advanced Plotly visualizations (Treemaps, Scatter Plots).

๐Ÿ“ซ Let's Connect

GitHub Stats

Popular repositories Loading

  1. job-market-etl-pipeline job-market-etl-pipeline Public

    Python

  2. rfm-customer-segmentation rfm-customer-segmentation Public

    Python

  3. mazen-gebrel mazen-gebrel Public

  4. predictive-churn-model predictive-churn-model Public

    Python

  5. local-rag-document-chat local-rag-document-chat Public

    Python

  6. audio-transcription-pipeline audio-transcription-pipeline Public

    Python