Skip to content
View AhmadBilalDSA's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report AhmadBilalDSA

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AhmadBilalDSA/README.md

Typing SVG

A detail-driven Data Strategist and Systems Engineer with a rigorous foundation in corporate finance. I specialize in translating complex operational datasets into actionable intelligence, building high-performance local data engineering utilities, and securing software workflows.

LinkedIn Email

Profile Views


πŸ‘¨β€πŸ’» About Me & Focus

  • πŸ“Š What I Do: Data Analyst & Analytics Specialist specializing in Python, SQL, Power BI, and local-first data engineering tools.
  • ⚑ Currently Building: Automated data pipelines, issue-tracking scrapers, and SQL performance benchmarks (data-engine-benchmarks, duck-diff).
  • πŸ™ Open Source: Active contributor to data parsing and transformation ecosystems (sqlglot, ibis-project, scikit-learn).
  • 🌱 Currently Exploring: Advanced DuckDB optimizations and high-performance local data tooling.
  • 🀝 Open to: Roles in Data Analytics, Analytics Engineering, or Open-Source collaborations.

πŸ“Š GitHub Stats & Activity

GitHub Stats Top Languages


🧰 Tech Stack & Expertise

Skill Icons

Power BI Tableau Pandas NumPy DuckDB Polars Streamlit Pytest


🌟 Open Source Impact & Ecosystem Contributions

I actively contribute to the core infrastructure of major Python data ecosystems and machine learning frameworks:

🌐 Project 🎯 Domain πŸ’‘ Highlight & Contribution
scikit-learn/scikit-learn Machine Learning Contributed core algorithmic and numerical enhancements to Python's premier ML library (PR #34800).
tobymao/sqlglot SQL Transpiler Enhanced multi-dialect SQL parsing logic, AST validation, and query transformations.
ibis-project/ibis Portable Dataframes Resolved complex data manipulation bugs and optimized unified dataframe backend conversions.
scitex-ai/scitex-io AI Infrastructure Improved AI-driven analytical data pipelines and core developer tooling (PR #166).

πŸ”¬ Detailed Engineering Deep-Dives

πŸ›‘οΈ Adexa | Unit Testing & Security (PR #10)
Context: The AI engine lacked coverage to verify if its automated repair strategies were successfully mitigating SQL injection vulnerabilities.
What I Built: Engineered over 300 lines of comprehensive unit tests in test_repair_strategies.py. Validated the engine's behavior against standard SQL injection patterns to ensure generated code patches meet strict security standards. (Stack: Python, Pytest)

βš™οΈ py-simple-wrap | Core Implementation & Mocking (PR #199)
Context: The repository required tests for an easy_sql utility, but the base module itself was missing from the core directory.
What I Built: Developed the core easy_sql.py execution module to handle sqlite3 connections. Concurrently built the testing suite (test_easy_sql.py), utilizing unittest.mock.patch and MagicMock to isolate the database connection and validate query execution flows without requiring a live local database. (Stack: Python, SQLite3, Pytest, Unittest.mock)


πŸš€ Featured Data Science & Machine Learning Projects

πŸ“Š Project βš™οΈ Stack πŸ’‘ Core Impact & Scope
KSE-100 Financial Sentiment Analysis Python (NLTK), Pandas, APIs, Tableau Built an automated news-scraping pipeline and used NLP sentiment analysis to correlate public news trends with KSE-100 stock price movements.
Predictive Modeling of Employee Turnover Python, Scikit-learn, Random Forest, Tableau Analyzed HR metrics, engineered classification features, and deployed a tuned Random Forest model to flag employee attrition risk factors.
SpaceX Falcon 9 Landing Prediction Python, SQL, REST APIs, Plotly Dash Executed end-to-end data collection, wrangling, and multi-model classification (SVM, Logistic Regression) visualized via an interactive web dashboard.
Ames Housing Real Estate Valuation Python, Pandas, XGBoost, Feature Engineering Trained high-performance regression models handling 80+ features, utilizing log transformations and feature creation to minimize pricing error bounds.
Instacart Market Basket Analysis Pandas, Seaborn, EDA Processed over 1M records using heavy groupby aggregations to map multi-product associations and user reorder frequencies.

πŸ› οΈ Personal Engineering & Local Data Utilities

I design and build high-performance, air-gapped local-first tools for developers and data teams:

  • πŸ›‘οΈ repo-doctor β€” Automated codebase diagnostic and health-check runner auditing secrets, licenses, and AI-readiness.
  • πŸš€ duck-diff β€” High-speed data schema and table diffing utility leveraging embedded DuckDB for massive file reconciliation.
  • 🧹 sqlean-lint β€” Lightweight local-first SQL static analysis and rule validation engine targeting performance traps.
  • πŸ“ˆ dbt-optimizer β€” Compilation analyzer and cost auditor toolkit for dbt core models and DAG dependencies.
  • ⚑ data-engine-benchmarks β€” Benchmarking framework evaluating speed, memory efficiency, and throughput across local vs. distributed engines (DuckDB, Polars, Pandas).
  • πŸ” github-issue-hunter β€” Automated search engine and dashboard designed to query, tag, and track high-priority GitHub issues across open-source ecosystems.

πŸ‘¨β€πŸ’» Professional Experience

  • 🏫 Power BI Developer & Instructor | PNY Trainings (NAVTTC) (Feb 2026 – May 2026)
    • Engineered and delivered comprehensive technical curriculum in Power BI, SQL, and Python; established best practices for ETL pipelines and advanced data modeling.
  • πŸ• Junior Data Scientist | Timmy's Pizza (Nov 2023 – Dec 2025)
    • Optimized local delivery routes, staffing schedules, and supply chain visibility using Pandas and interactive Power BI dashboards.
  • πŸ’» Data Analyst Intern | PNY Trainings (Jun 2025 – Present)
    • Spearheaded a data-driven marketing analysis for a key e-commerce client, leveraging Power BI and Advanced Excel to project a 15% increase in customer engagement.
    • Automated a reporting pipeline using Python (Pandas) and SQL, reducing manual data processing for weekly sales reports by 10 hours per month and improving efficiency by 30%.
  • β›½ Procurement Intern | Sui Northern Gas Pipelines Limited (SNGPL)
    • Analyzed vendor performance metrics for a portfolio of 50+ suppliers, creating KPI dashboards in Excel that contributed to an estimated 5% reduction in procurement costs.
    • Streamlined the digital record-keeping process for Purchase Orders, designing a new workflow that reduced document retrieval times by over 50%.

πŸ“œ Education & Certifications

  • πŸŽ“ Bachelor of Science in Accounting and Finance β€” Hailey College of Commerce, University of the Punjab
  • πŸ† IBM Data Analyst Professional Certificate
  • πŸ† Google Advanced Data Analytics Professional Certificate
  • πŸ† IBM Data Science Professional Certificate
  • πŸ† IBM Data Engineering Professional Certificate
  • πŸ† Google Digital Marketing & E-commerce Professional Certificate

🐍 Contribution Activity Matrix

github contribution grid snake animation

[ Git Push / PR ] ──► [ GitHub Actions ] ──► [ Linting & PyTest Matrix ] ──► [ Deploy / Publish ]

Pinned Loading

  1. duck-diff duck-diff Public

    High-performance, constant-memory data diff engine powered by DuckDB SQL. Keyed/keyless reconciliation across Parquet, CSV, JSON, and SQLite with float epsilon tolerance.

    HTML 1

  2. github-issue-hunter github-issue-hunter Public

    Python 1

  3. fast-analytics-engine fast-analytics-engine Public

    High-performance in-memory SQL analytics engine powered by DuckDB, Polars, and Streamlit.

    Python 1

  4. data-engine-benchmarks data-engine-benchmarks Public

    Python 1

  5. financial-analytics-pipeline financial-analytics-pipeline Public

    Python 1

  6. sqlean-lint sqlean-lint Public

    Python 1