Skip to content
View loueylahwel's full-sized avatar
💀
💀

Highlights

  • Pro

Block or report loueylahwel

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
loueylahwel/README.md
ascii skull animation

Louey Lahwel

Data Engineering Student · Sfax, Tunisia

LinkedIn GitHub Email

About

Engineering student specializing in Data Engineering at the Faculty of Sciences of Sfax, Tunisia. I build data infrastructure and distributed systems — pipelines and platforms that turn raw, unstructured data into actionable insight.

I enjoy solving architectural problems: orchestrating multi-node Hadoop/Spark clusters, designing medallion lakehouses, real-time processing, federated learning, and cloud-native warehousing.

location   ->  Sfax, Tunisia
education  ->  Engineering Cycle, Data Engineering — Faculty of Sciences of Sfax
focus      ->  Big Data | Distributed Systems | Data Pipelines | MLOps
currently  ->  Building scalable data infrastructure

Stack

Programming

Python SQL Java C C++

Big Data & Distributed Systems

Apache Spark Apache Hadoop Apache Iceberg Apache Airflow

Infrastructure & Orchestration

Docker Terraform Ansible Redis FastAPI Celery

Databases

ClickHouse PostgreSQL MySQL Azure SQL Azure Data Lake Azure Blob Storage

Visualization & BI

Grafana Apache Superset Prometheus

Cloud & DevOps

Azure Linux Git GitHub

Projects

Automated Multi-Node Distributed Cluster Orchestrator
A self-service platform for provisioning Apache Hadoop and Spark clusters via an async pipeline. Built with FastAPI, Celery, Redis, Terraform, and Ansible. Features JWT-based RBAC, real-time status streaming, and a fully containerized deployment stack with a monitoring layer.

FastAPI Celery Redis Terraform Ansible Docker Compose

GitHub Archive Trend & Virality Analytics Platform
End-to-end analytics platform ingesting GH Archive data through a Bronze-Silver-Gold Medallion architecture on Apache Iceberg and Spark. Includes a virality scoring engine, tech-stack trend analysis, and time-travel queries backed by LocalStack S3 and Iceberg REST Catalog.

Apache Spark Apache Iceberg LocalStack S3 Medallion Architecture

Intelligent Schema-Aware Web Scraping Framework
A Python-based framework for automated schema discovery and structural HTML analysis. Features modular parsing and extraction pipelines with reusable data models for downstream analytics and warehousing — designed for production-ready data ingestion.

Python HTML Parsing ETL Pipelines Data Modeling

Text-to-SQL Local Agent (ClickHouse + LLM)
A fully local Text-to-SQL system integrating FastAPI, ClickHouse, and an LLM runtime. Implements schema introspection, SQL validation, and Dockerized multi-service deployment for natural language querying over analytical databases.

FastAPI ClickHouse LLM Docker

Federated Learning Anomaly Detection System
LSTM autoencoder models for time-series anomaly detection within a federated learning architecture. Includes preprocessing pipelines, feature engineering, and integration into distributed model aggregation workflows for network security analytics.

LSTM Federated Learning Time-Series Feature Engineering

Real-Time Analytical Data Warehouse
Azure-hosted serverless data warehouse with automated ingestion via Azure Function Apps. Applied advanced data modeling for high-frequency financial data streams. Real-time dashboards via Grafana and Apache Superset, with OLAP queries and materialized views.

Azure Azure Function Apps Grafana Apache Superset OLAP

Certifications

Certification Issuer
Building Customized LLMs with OpenAI Columbia+
Learning AI Through Visualization Columbia+
CCNA: Introduction to Networks Cisco Networking Academy
Building Data Pipelines with Apache Airflow 365 Data Science
Advanced SQL for Data Engineering 365 Data Science

"The goal is to turn data into information, and information into insight."

Pinned Loading

  1. GitGud GitGud Public

    Turns your GitHub repositories into resume-worthy CV entries and inserts them into your LaTeX résumé automatically.

    Python 4

  2. Automated-Multi-Node-Distributed-Cluster-Orchestrator Automated-Multi-Node-Distributed-Cluster-Orchestrator Public

    A self-service platform that provisions, configures, monitors, and destroys Apache Hadoop + Apache Spark clusters using a fully automated, end-to-end infrastructure pipeline.

    TypeScript 2

  3. github-archive-analytics github-archive-analytics Public

    Viral Repo & Tech Trend Tracker | Ingests raw GH Archive data into an Iceberg data lake to surface trending repositories and ecosystem insights. Uses Spark for sliding-window aggregations and Icebe…

    Python 2

  4. CryptoPulse-DWH CryptoPulse-DWH Public

    Serverless Crypto Analytics & Volatility Engine

    Python 2

  5. sql-gen-agent-v2 sql-gen-agent-v2 Public

    A fully local, Dockerized Text-to-SQL application that translates natural language questions into ClickHouse queries using Ollama. It features a FastAPI backend with integrated SQL validation, sche…

    Python 2

  6. Collecty Collecty Public

    An Intelligent, Agentic Web Scraping Framework for Automated Schema Discovery and Structural Analysis.

    Python 5