Building interpretable data products, reproducible analytics and privacy-aware AI systems.
Projects • Health AI R&D • Technical focus • Français
I am a Genomics Data Scientist intern at DataPathology and a final-year student in the MSc Data Analytics for Business at KEDGE Business School.
My work sits at the intersection of:
- interpretable machine learning;
- genomics and healthcare;
- data engineering and analytical quality;
- applied AI product development.
My MSc research focuses on interpretable machine-learning methods for complex genomic data, including pathway-level feature engineering, weak supervision and multilabel classification.
| Project | Main technical contribution | Academic context |
|---|---|---|
| 🌱 EcoMeal Bot · Live demo | Streamlit application combining 2,000+ recipes, environmental data, CO₂ estimation, user profiles and local/cloud LLM support. Automated GitHub-to-Hugging Face deployment. | Artificial Intelligence and SDG, MSc 2 S2 · Group project · 16.7/20 |
| 🏅 Olympic SQL Database | Seven-table SQLite model, deterministic synthetic data generation, 12 analytical queries, 11 integrity checks, ER diagram and data dictionary. | Data Management, MSc 2 S1 · Group project; responsible for the SQL component · 16/20 |
| ₿ Bitcoin Market Correlations | End-to-end financial-data ETL, rolling correlations, dynamic regressions and five interactive Plotly visualizations. | Python Bootcamp, MSc 2 S1 · Group project; responsible for all code and visualizations · 19.5/20 |
| 🌍 Global Economic & Environmental Analysis | Multi-source data preparation, international GDP/CO₂ analysis, statistical correlations, regression and geospatial visualization. | Fundamental of Data Analytics for Business, MSc 1 S2 · 17.2/20 |
Privacy-first, offline clinical voice workflow for pathology reporting. The project explores local speech processing, structured report generation, validation workflows and CPU-compatible deployment.
Offline-first medical document-intelligence engine for pathology and radiology workflows, with document ingestion, OCR orchestration, structured extraction, evidence traceability and mandatory human review.
These projects remain private because they involve proprietary product work and potentially sensitive clinical workflows. They are active prototypes, not clinically validated medical devices or production-ready systems.
Python · pandas · NumPy · scikit-learn · XGBoost · feature engineering · multilabel classification · weak supervision · NLP · time-series analysis · model evaluation
SQL · SQLite · PostgreSQL · ETL · data validation · data quality · Plotly · Tableau · Power BI
Streamlit · FastAPI · local LLMs · Hugging Face · GitHub Actions · pytest · Git · Linux · offline-first systems
- Developing interpretable ML workflows for genomic and healthcare data.
- Building privacy-aware, local-first clinical AI prototypes.
- Strengthening testing, data-quality controls and reproducibility.
- Turning academic analyses into clear, maintainable portfolio repositories.
🇫🇷 Version française
Je suis Genomics Data Scientist en stage chez DataPathology et étudiant en dernière année du MSc Data Analytics for Business à KEDGE Business School.
Je développe des pipelines de machine learning interprétables et des produits d’IA appliqués à la génomique, à la santé et à l’aide à la décision. Mon mémoire porte sur l’utilisation de méthodes de ML interprétables pour des données génomiques complexes, notamment le feature engineering par pathways, la supervision faible et la classification multilabel.
- EcoMeal Bot — application Streamlit combinant recettes, données environnementales, estimation CO₂ et modèles de langage locaux ou cloud.
- Olympic SQL Database — modélisation relationnelle SQLite, SQL analytique et contrôles d’intégrité.
- Bitcoin Market Correlations — pipeline ETL financier, corrélations glissantes, régressions et visualisations Plotly.
- Analyse économique et environnementale globale — analyse multi-sources du PIB, des émissions de CO₂, de la démographie et des dépenses militaires.
VoxPath et Scriptum sont deux prototypes privés consacrés respectivement aux workflows vocaux cliniques et à l’intelligence documentaire médicale locale. Ils restent volontairement privés afin de protéger les éléments propriétaires et les workflows potentiellement sensibles.