Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

adstyle

AI-powered marketing image generator. Turn a plain product photo into a finished, on-brand advertisement — while keeping the product itself pixel-for-pixel intact.

adstyle system architecture


Gallery

The web app — upload a product photo, describe the mood, and get a finished ad with an auto-written caption. Prompts work in English and Arabic.

adstyle web app — English

Finished ads — the product in each ad is the user's original photo, composited back pixel-for-pixel; only the scene around it is generated.

Red Diamond perfume ad Tobacco perfume ad Modelo beach ad

Style transfer, base vs LoRA — same prompt and seed. The base model (left) drifts into a cluttered stock scene; the fine-tuned LoRA (right) produces a single clean hero product.

Base SDXL vs adstyle LoRA


The problem

Diffusion models do not preserve objects. Ask any text-to-image model for "a Chanel N°5 bottle on marble" and it will paint a lookalike — with a distorted logo and hallucinated text. That is fine for art and useless for advertising, where the ad must show the real product.

The idea

adstyle inverts the usual flow. Instead of asking the model to draw the product, it asks the model to draw the world around the product:

product photo → cutout (U²-Net) → build canvas + mask
             → SDXL inpainting generates the scene around a protected region
             → re-paste the original cutout, pixel-for-pixel
             → contact shadow + light-wrap → design layer → caption

The product core is verified identical to the source (np.array_equal over ~96,652 pixels). The logo survives because it is never regenerated.


Highlights

  • Pixel-exact product fidelity — the real product is composited back over the generated scene, not synthesised.
  • A custom LoRA fine-tuned on a 1,253-image advertising dataset built from scratch (rank 16, 2,000 steps, dual T4).
  • An LLM "art director" that turns one short idea into several distinct creative concepts, with a deterministic 96-scene fallback so the system never depends on the model being online.
  • Best-of-N selection using CLIP as a multimodal reward function.
  • Physical grounding — contact-line detection, adaptive scale, and a computed shadow so the product sits on the surface instead of floating.
  • A bilingual design layer (Arabic + English) rendered with Pillow + raqm for correct Arabic shaping.
  • An input quality gate that inspects the uploaded photo and auto-repairs common problems before generation.
  • A FastAPI + React web app, deployable from a single command.

Repository layout

adstyle/
├── app/                    the generation pipeline
│   ├── schemas.py          Pydantic data contracts
│   ├── pipeline.py         orchestrator: routes a request through the stages
│   ├── main.py             FastAPI entry point
│   ├── design.py           text / price / logo layer (Pillow + raqm)
│   ├── prompts/            LLM system prompts
│   ├── assets/fonts/       bundled OFL fonts
│   └── stages/
│       ├── quality_gate.py       Stage 0 — inspect the uploaded photo
│       ├── prompt_optimizer.py   Stage 1 — idea → SDXL prompt (LLM)
│       ├── art_director.py       Stage 1b — idea → N concepts (LLM)
│       ├── product_processor.py  Stage 2 — segmentation → RGBA cutout
│       ├── composition.py        Stage 4 — inpaint-around-product (core)
│       ├── staging.py            staged podium / spotlight path
│       ├── integration.py        contact line, support detection, shadow
│       ├── selector.py           best-of-N CLIP scoring
│       ├── generation.py         no-product text-to-image path
│       └── captions.py           Stage 5 — caption + hashtags (LLM)
├── ui/
│   └── gradio_app.py       lightweight Gradio front-end
├── web/                    web-app add-ons
│   ├── preflight.py        input gate: checks + auto-repairs (no GPU/LLM)
│   └── preflight_api.py    /api/validate endpoint + in-page warning card
├── notebooks/              the offline model pipeline (Kaggle)
│   ├── 01_dataset_builder  collect + filter + caption 1,253 images
│   ├── 02_lora_training    fine-tune the SDXL LoRA
│   └── 03_evaluation       CLIP + FID + ablation grid
├── tests/                  124 GPU-free tests
├── docs/                   architecture diagram + write-ups
└── requirements.txt

Quick start

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --reload

Then open http://localhost:8000. Set OPENROUTER_API_KEY in the environment to enable the LLM stages; without it the deterministic fallbacks are used.

Run the tests:

pytest -q

How it works, stage by stage

Stage File What it does
0 quality_gate.py / web/preflight.py Inspect the uploaded photo — hands, blur, resolution, cut-off — and auto-repair the cutout
1 prompt_optimizer.py An LLM turns a short idea into a professional SDXL prompt
1b art_director.py An LLM invents N distinct concepts (scene, lighting, headline, font, colour) as validated JSON
2 product_processor.py U²-Net segments the product into an RGBA cutout
4 composition.py Builds the canvas, runs SDXL inpainting, then re-pastes the original cutout pixel-for-pixel
— integration.py Detects the surface line and draws a contact shadow so the product is grounded
— selector.py Scores N candidates with CLIP and keeps the strongest
— design.py Renders the headline, price and logo, with correct Arabic shaping
5 captions.py An LLM writes the post caption, hashtags and call-to-action

Every stage is optional: a null spec reproduces the previous behaviour exactly, which is what makes controlled A/B testing possible.


Model & evaluation

The LoRA was evaluated against the base SDXL model on 80 images (20 prompts × 2 seeds × 2 conditions):

Metric Base SDXL + adstyle LoRA
CLIP Score (mean, n=40) 32.59 30.57
FID vs 400 real images 227.48 230.57

The 2-point CLIP drop is a known style-vs-adherence trade-off; the FID difference is not statistically decisive at this sample size. The decisive evidence is visual — the ablation grid in docs/ shows the base model producing cluttered stock scenes while the LoRA produces a single clean hero product. See docs/Executive_Summary.md for the full account.


Documentation

  • docs/Executive_Summary.md — five-page overview of the whole project
  • docs/Code_Walkthrough.md — file-by-file technical walkthrough
  • docs/AI_Justification.md — the AI methods used and why
  • docs/ARCHITECTURE_AND_PLAN.md — original design notes

Tech stack

SDXL Inpainting · LoRA (PEFT) · U²-Net (rembg) · CLIP ViT-B/32 · MediaPipe · Pillow + raqm · FastAPI · React + Vite · Gradio · OpenRouter

License

MIT — see LICENSE.

About

AI-powered marketing image generator — turns a product photo into a finished ad while keeping the product pixel-for-pixel intact. SDXL inpainting + custom LoRA + LLM art director, with a bilingual (Arabic/English) web app.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages