AI-powered marketing image generator. Turn a plain product photo into a finished, on-brand advertisement — while keeping the product itself pixel-for-pixel intact.
The web app — upload a product photo, describe the mood, and get a finished ad with an auto-written caption. Prompts work in English and Arabic.
Finished ads — the product in each ad is the user's original photo, composited back pixel-for-pixel; only the scene around it is generated.
Style transfer, base vs LoRA — same prompt and seed. The base model (left) drifts into a cluttered stock scene; the fine-tuned LoRA (right) produces a single clean hero product.
Diffusion models do not preserve objects. Ask any text-to-image model for "a Chanel N°5 bottle on marble" and it will paint a lookalike — with a distorted logo and hallucinated text. That is fine for art and useless for advertising, where the ad must show the real product.
adstyle inverts the usual flow. Instead of asking the model to draw the product, it asks the model to draw the world around the product:
product photo → cutout (U²-Net) → build canvas + mask
→ SDXL inpainting generates the scene around a protected region
→ re-paste the original cutout, pixel-for-pixel
→ contact shadow + light-wrap → design layer → caption
The product core is verified identical to the source (np.array_equal over ~96,652 pixels). The logo survives because it is never regenerated.
- Pixel-exact product fidelity — the real product is composited back over the generated scene, not synthesised.
- A custom LoRA fine-tuned on a 1,253-image advertising dataset built from scratch (rank 16, 2,000 steps, dual T4).
- An LLM "art director" that turns one short idea into several distinct creative concepts, with a deterministic 96-scene fallback so the system never depends on the model being online.
- Best-of-N selection using CLIP as a multimodal reward function.
- Physical grounding — contact-line detection, adaptive scale, and a computed shadow so the product sits on the surface instead of floating.
- A bilingual design layer (Arabic + English) rendered with Pillow + raqm for correct Arabic shaping.
- An input quality gate that inspects the uploaded photo and auto-repairs common problems before generation.
- A FastAPI + React web app, deployable from a single command.
adstyle/
├── app/ the generation pipeline
│ ├── schemas.py Pydantic data contracts
│ ├── pipeline.py orchestrator: routes a request through the stages
│ ├── main.py FastAPI entry point
│ ├── design.py text / price / logo layer (Pillow + raqm)
│ ├── prompts/ LLM system prompts
│ ├── assets/fonts/ bundled OFL fonts
│ └── stages/
│ ├── quality_gate.py Stage 0 — inspect the uploaded photo
│ ├── prompt_optimizer.py Stage 1 — idea → SDXL prompt (LLM)
│ ├── art_director.py Stage 1b — idea → N concepts (LLM)
│ ├── product_processor.py Stage 2 — segmentation → RGBA cutout
│ ├── composition.py Stage 4 — inpaint-around-product (core)
│ ├── staging.py staged podium / spotlight path
│ ├── integration.py contact line, support detection, shadow
│ ├── selector.py best-of-N CLIP scoring
│ ├── generation.py no-product text-to-image path
│ └── captions.py Stage 5 — caption + hashtags (LLM)
├── ui/
│ └── gradio_app.py lightweight Gradio front-end
├── web/ web-app add-ons
│ ├── preflight.py input gate: checks + auto-repairs (no GPU/LLM)
│ └── preflight_api.py /api/validate endpoint + in-page warning card
├── notebooks/ the offline model pipeline (Kaggle)
│ ├── 01_dataset_builder collect + filter + caption 1,253 images
│ ├── 02_lora_training fine-tune the SDXL LoRA
│ └── 03_evaluation CLIP + FID + ablation grid
├── tests/ 124 GPU-free tests
├── docs/ architecture diagram + write-ups
└── requirements.txt
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reloadThen open http://localhost:8000. Set OPENROUTER_API_KEY in the environment to enable the LLM stages; without it the deterministic fallbacks are used.
Run the tests:
pytest -q| Stage | File | What it does |
|---|---|---|
| 0 | quality_gate.py / web/preflight.py |
Inspect the uploaded photo — hands, blur, resolution, cut-off — and auto-repair the cutout |
| 1 | prompt_optimizer.py |
An LLM turns a short idea into a professional SDXL prompt |
| 1b | art_director.py |
An LLM invents N distinct concepts (scene, lighting, headline, font, colour) as validated JSON |
| 2 | product_processor.py |
U²-Net segments the product into an RGBA cutout |
| 4 | composition.py |
Builds the canvas, runs SDXL inpainting, then re-pastes the original cutout pixel-for-pixel |
| — | integration.py |
Detects the surface line and draws a contact shadow so the product is grounded |
| — | selector.py |
Scores N candidates with CLIP and keeps the strongest |
| — | design.py |
Renders the headline, price and logo, with correct Arabic shaping |
| 5 | captions.py |
An LLM writes the post caption, hashtags and call-to-action |
Every stage is optional: a null spec reproduces the previous behaviour exactly, which is what makes controlled A/B testing possible.
The LoRA was evaluated against the base SDXL model on 80 images (20 prompts × 2 seeds × 2 conditions):
| Metric | Base SDXL | + adstyle LoRA |
|---|---|---|
| CLIP Score (mean, n=40) | 32.59 | 30.57 |
| FID vs 400 real images | 227.48 | 230.57 |
The 2-point CLIP drop is a known style-vs-adherence trade-off; the FID difference is not statistically decisive at this sample size. The decisive evidence is visual — the ablation grid in docs/ shows the base model producing cluttered stock scenes while the LoRA produces a single clean hero product. See docs/Executive_Summary.md for the full account.
docs/Executive_Summary.md— five-page overview of the whole projectdocs/Code_Walkthrough.md— file-by-file technical walkthroughdocs/AI_Justification.md— the AI methods used and whydocs/ARCHITECTURE_AND_PLAN.md— original design notes
SDXL Inpainting · LoRA (PEFT) · U²-Net (rembg) · CLIP ViT-B/32 · MediaPipe · Pillow + raqm · FastAPI · React + Vite · Gradio · OpenRouter
MIT — see LICENSE.





