diff --git a/.firebaserc b/.firebaserc
new file mode 100644
index 0000000..f83c837
--- /dev/null
+++ b/.firebaserc
@@ -0,0 +1,5 @@
+{
+ "projects": {
+ "default": "buildwithsparky"
+ }
+}
diff --git a/README.md b/README.md
index b556ddf..6fab522 100644
--- a/README.md
+++ b/README.md
@@ -4,8 +4,8 @@
Pixel runs a gaze model over an ad, finds what is pulling the eye away from the brand, and has a team of AI agents redesign it. Every edit is re-scored by the gaze model, and edits that don't raise the score are thrown out.
- 
- Try the demo
+ 
+ Try the demo
Built in one day at Multimodal Hacks (NY Tech Week, June 6 2026) with Rishi Shah.
@@ -14,7 +14,7 @@ Pixel runs a gaze model over an ad, finds what is pulling the eye away from the
1. **Look.** [DeepGaze IIE](https://github.com/matthias-k/DeepGaze), a neural saliency model trained on real eye-tracking data, predicts how likely each pixel of the ad is to be looked at.
2. **Score.** Pixel measures how much of that attention lands on the brand (the logo, the product, the call to action) compared with how big the brand is. 50 means the brand gets exactly its fair share for its size. Higher means it pulls more than its size.
-3. **Find the thieves.** The strongest spots outside the brand are the things stealing attention. Gemini looks at each one and names it ("Subway sign", "smiling woman's face").
+3. **Find the thieves.** The strongest spots outside the brand are the things stealing attention. Gemini looks at each one and names it ("Subway sign", "smartwatch band").
4. **Edit.** Agents propose changes, Gemini's image model ("Nano Banana") makes them, and DeepGaze re-scores every result. A brand-fit judge and a reward-hack guard can veto an edit. You watch each branch land and decide whether to grow another.

@@ -23,32 +23,33 @@ That is the app on the Nike sample. The billboard is the ad, but the first predi
## What the gaze model sees
-
+
A few things the model picks up that a person would also notice:
-- **Faces win.** On The Ordinary ad, the woman's face gets the first fixation and more than half the attention in the frame. The serum bottles she is selling come second.
- **Text is a magnet.** On McDonald's, the eye walks along the lit nameplate before it touches the arches.
- **Clutter competes.** On Red Bull, the phone and the mouse on the desk each pull a fixation away from the can.
- **Red on red disappears.** The Coca-Cola can scores well, but its heat barely shows against the red background. The model still finds the script logo.
## One edit that worked
-
+
-The search on the Nike billboard tried six edits over two rounds. Darkening only the slogan or only the store sign barely helped. Muting everything except the logo and product, then boosting their contrast and sharpness, did a lot: attention on target went from 59 to 81, and the absolute amount of attention on the brand almost tripled. Those numbers are recomputed from the saved images every time the figure is built, by [`scripts/make_figures.py`](scripts/make_figures.py).
+The Red Bull can shares the desk with a phone, a keyboard, a drawing tablet and a mouse, and the mouse pulls a fixation of its own. One live edit asked Nano Banana to clear the desk and change nothing else: no text, same light, same can. Attention on target went from 79 to 82, and the share of all attention that lands on the can went from 49% to 59%. Those numbers are recomputed from the saved images every time the figure is built, by [`scripts/make_figures.py`](scripts/make_figures.py).
+
+It is a small win, and it is the kind Pixel now accepts. An earlier version of this section showed a Nike edit that went from 59 to 81 by muting everything except the billboard. That is a cheat, not a better ad, and the guard below now rejects it.
## Most edits don't work, and Pixel says so
-
+
-This is a live run on The Ordinary ad from September 2026, with every edit saved. Each one looks more like a finished campaign than the original, and each one scored lower. A new headline gives the eye something else to read. Two of the edits moved the bottles, and the brand box does not move with them. So Pixel kept the original and reported a change of zero.
+This is a recorded optimizer run on the same Red Bull ad, with every edit saved. Each one looks more like a finished campaign than the original, and each one scored lower. Every one of them added a headline or a slogan, and text pulls the eye off the can. So Pixel kept the original and reported a change of zero. The edit that won above added no text at all.
The first version of Pixel could not do this. It measured raw attention inside the brand box, so "make the logo bigger" always won, and it floored every result at the baseline, so the number could only go up. The version here fixes both:
- **The score is size-invariant.** Attention share is divided by the box's share of the frame, so enlarging the target stops being a free win.
- **Losing edits stay visible.** The before and after show the real change, including zero or negative.
-- **A guard catches two cheats** ([`backend/eval_guard.py`](backend/eval_guard.py)). If the brand's share of attention went up but the absolute attention on it did not, the edit won by dimming everything else and is rejected. If the edit is too small to see, it is rejected as a likely adversarial trick.
+- **A guard catches three cheats** ([`backend/eval_guard.py`](backend/eval_guard.py)). If the brand's share of attention went up but the absolute attention on it did not, the edit is rejected. If the edit is too small to see, it is rejected as a likely adversarial trick. And if the rest of the frame came out darker, flatter, grayer or blurrier than before, the edit won by degrading the scene and is rejected. That last check compares the frame outside the brand box cell by cell, so removing one cluttering object passes and dimming the whole scene does not.
- **The badge tells the truth.** If DeepGaze fails to load and the cheap fallback runs instead, `/health` says so and the app shows DEMO instead of LIVE.
The full list of ways the score can still be gamed, and what is done about each, is in [`docs/EVAL_FLAWS.md`](docs/EVAL_FLAWS.md).
@@ -86,14 +87,14 @@ npm run dev # http://localhost:5173
Keys go in git-ignored files: `GEMINI_API_KEY`, `PINECONE_API_KEY` and `PINECONE_INDEX` in `backend/.env`, and `VITE_CLERK_PUBLISHABLE_KEY` in `frontend/.env.local`. Seed the competitor index once with `python backend/pinecone_seed.py`. Without a Gemini key the app still analyzes ads, but edits come back unchanged. DeepGaze runs on the GPU when there is one and on the CPU otherwise.
-The [live demo](https://sparkylab.web.app/pixel/) is the same frontend replaying runs recorded on the real backend, so it needs no server and no keys. `python scripts/record_demo.py` records them and `npm run build:demo` in `frontend` writes the static site to `web-dist/`.
+The [live demo](https://pixel-gaze.web.app/) is the same frontend replaying runs recorded on the real backend, so it needs no server and no keys. `python scripts/record_demo.py` records them and `npm run build:demo` in `frontend` writes the static site to `web-dist/`.
To rebuild the figures in this README:
```bash
python scripts/make_showcase.py # gallery, GIF, every-edit strip
-python scripts/make_figures.py # Nike before and after
-python scripts/capture_run.py the-ordinary "The Ordinary" # a new live run (needs a Gemini key)
+python scripts/make_figures.py # Red Bull before and after
+python scripts/try_edits.py red-bull "Red Bull" "" # a new live edit (needs a Gemini key)
python backend/test_scoring.py && python backend/test_eval_guard.py
```
@@ -102,7 +103,7 @@ python backend/test_scoring.py && python backend/test_eval_guard.py
| [`backend/`](backend/) | FastAPI app, DeepGaze runner and score, agents, branch search, reward-hack guard, tests |
| [`frontend/`](frontend/) | Vite, React and TypeScript app, the sample ads, the saved Nike run |
| [`scripts/`](scripts/) | Figure builders and the live-run capture |
-| [`results/`](results/) | The Ordinary run: every edit and its scores |
+| [`results/`](results/) | Live edits with their scores, Judge and guard verdicts, including the Red Bull win |
| [`docs/`](docs/) | How the score can be gamed and what stops it |
## Credits
diff --git a/assets/media/deepgaze.gif b/assets/media/deepgaze.gif
index 7228807..e2765a0 100644
Binary files a/assets/media/deepgaze.gif and b/assets/media/deepgaze.gif differ
diff --git a/assets/media/every-edit.png b/assets/media/every-edit.png
index 9f03c3b..dcbbd5f 100644
Binary files a/assets/media/every-edit.png and b/assets/media/every-edit.png differ
diff --git a/assets/media/gallery.png b/assets/media/gallery.png
index d7b2c33..eff6db9 100644
Binary files a/assets/media/gallery.png and b/assets/media/gallery.png differ
diff --git a/assets/media/hero-before-after.png b/assets/media/hero-before-after.png
index 4778e86..dd29777 100644
Binary files a/assets/media/hero-before-after.png and b/assets/media/hero-before-after.png differ
diff --git a/backend/agents.py b/backend/agents.py
index 7a40823..c3b04cd 100644
--- a/backend/agents.py
+++ b/backend/agents.py
@@ -182,7 +182,7 @@ def _step_finalize(state: dict) -> dict:
ratio_before=baseline, ratio_after=final,
target_sal_before=before.get("target_salience"),
target_sal_after=after.get("target_salience"),
- edit_is_semantic=True,
+ edit_is_semantic=True, target_box=target,
)
steps = [
diff --git a/backend/eval_guard.py b/backend/eval_guard.py
index 59b3d69..36eea1a 100644
--- a/backend/eval_guard.py
+++ b/backend/eval_guard.py
@@ -6,6 +6,11 @@
and (b) **imperceptible / global** tweaks the eye can't read as a design change.
This module gives the loop a cheap second opinion that catches those.
+A third cheat is the one the absolute-salience check misses: darken, desaturate, blur
+or black out everything except the brand. DeepGaze is a probability map, so pulling
+mass off the rest of the frame raises the brand's absolute share too. Pass
+`target_box` and the guard compares the rest of the frame before and after.
+
Wiring (Pixel): `deepgaze_runner.score_components(image, box)` returns both the
size-invariant prominence (the headline `ratio`) AND the absolute on-target salience
mass — feed both into `verdict(...)`:
@@ -17,7 +22,7 @@
before_img, after_img,
ratio_before=ratio_before, ratio_after=ratio_after,
target_sal_before=abs_before, target_sal_after=abs_after,
- edit_is_semantic=really_edited,
+ edit_is_semantic=really_edited, target_box=box,
)
accepted = bool(v["decision"] == "accept") # gate on this, surface v["reasons"]
@@ -42,6 +47,12 @@
GLOBAL_COVERAGE = 0.55 # fraction of pixels changed above which an edit looks "global"
GLOBAL_BBOX = 0.60 # change bounding-box covering more of the frame than this = "global"
EPS = 1e-4
+GRID = 8 # cells per side for the rest-of-frame comparison
+OUTSIDE_MARGIN = 0.04 # grow the target box by this before measuring "the rest of the frame"
+# Floors for the rest of the frame, as after/before ratios. Below any of these, the edit
+# raised the score by degrading everything that isn't the brand. Calibrated on the
+# recorded edits in frontend/public/replay and results/ (see test_eval_guard.py).
+OUTSIDE_FLOORS = {"lum": 0.80, "contrast": 0.75, "color": 0.70, "detail": 0.70}
def _gray(img: Image.Image, size=None) -> np.ndarray:
@@ -115,6 +126,71 @@ def perceptual_change(before: Image.Image, after: Image.Image) -> dict:
}
+def _rgb(img: Image.Image, size) -> np.ndarray:
+ return np.asarray(img.convert("RGB").resize(size, Image.BILINEAR), dtype=np.float64)
+
+
+def _outside_mask(h: int, w: int, box, margin: float = OUTSIDE_MARGIN) -> np.ndarray:
+ """True outside the target box, grown by `margin` (fraction of the frame) on every side
+ so a legit edit that spills slightly past the box isn't counted as 'the rest'."""
+ x, y, bw, bh = box
+ x0, y0 = int(max(0.0, x - margin) * w), int(max(0.0, y - margin) * h)
+ x1, y1 = int(min(1.0, x + bw + margin) * w), int(min(1.0, y + bh + margin) * h)
+ m = np.ones((h, w), bool)
+ m[y0:y1, x0:x1] = False
+ return m
+
+
+def _stats(rgb: np.ndarray, mask: np.ndarray) -> dict:
+ lum = rgb @ np.array([0.299, 0.587, 0.114])
+ rg = rgb[..., 0] - rgb[..., 1]
+ yb = 0.5 * (rgb[..., 0] + rgb[..., 1]) - rgb[..., 2]
+ gy, gx = np.gradient(lum)
+ grad = np.hypot(gx, gy)
+ l, r, b = lum[mask], rg[mask], yb[mask]
+ return {
+ "lum": l.mean(),
+ "contrast": l.std(),
+ # Hasler & Suesstrunk colorfulness
+ "color": np.hypot(r.std(), b.std()) + 0.3 * np.hypot(r.mean(), b.mean()),
+ "detail": grad[mask].mean(),
+ }
+
+
+def outside_change(before: Image.Image, after: Image.Image, box) -> dict:
+ """How the rest of the frame (outside the target) changed, as after/before ratios of
+ luminance, luminance contrast, colorfulness and edge detail. Each ratio is the MEDIAN
+ over a grid of cells, so removing one cluttering object (a few cells change) passes,
+ while dimming, flattening, desaturating or blurring the whole scene (most cells
+ change) does not. A ratio well below 1 means the edit degraded the rest of the frame."""
+ w, h = before.size
+ scale = min(1.0, MAX_SIDE / max(w, h))
+ size = (max(8, int(w * scale)), max(8, int(h * scale)))
+ a, b = _rgb(before, size), _rgb(after, size)
+ mask = _outside_mask(size[1], size[0], box)
+ ratios: dict[str, list[float]] = {"lum": [], "contrast": [], "color": [], "detail": []}
+ ys = np.linspace(0, size[1], GRID + 1).astype(int)
+ xs = np.linspace(0, size[0], GRID + 1).astype(int)
+ for i in range(GRID):
+ for j in range(GRID):
+ cell = (slice(ys[i], ys[i + 1]), slice(xs[j], xs[j + 1]))
+ m = mask[cell]
+ if m.mean() < 0.5: # mostly target; not "the rest of the frame"
+ continue
+ sa, sb = _stats(a[cell], m), _stats(b[cell], m)
+ for k in ratios:
+ ratios[k].append((sb[k] + 1.0) / (sa[k] + 1.0))
+ if not ratios["lum"]: # target fills the frame; nothing outside to judge
+ return {k: 1.0 for k in ratios}
+ return {k: round(float(np.median(v)), 3) for k, v in ratios.items()}
+
+
+def degradation(oc: dict) -> list[str]:
+ """Which of the rest-of-frame signals fell past its floor."""
+ words = {"lum": "darkened", "contrast": "flattened", "color": "desaturated", "detail": "blurred"}
+ return [f"{words[k]} ({oc[k]:.2f}x)" for k, floor in OUTSIDE_FLOORS.items() if oc[k] < floor]
+
+
def verdict(
before: Image.Image,
after: Image.Image,
@@ -126,10 +202,13 @@ def verdict(
sal2_before: float | None = None,
sal2_after: float | None = None,
edit_is_semantic: bool = True,
+ target_box=None,
) -> dict:
"""Decide accept / reject / review for one edit. Gate `accepted` on decision == 'accept'.
Priority of checks (reasons explain every outcome):
+ 0. rest of frame degraded -> reject (score rose, but the frame outside the target got
+ darker / flatter / grayer / blurrier; needs `target_box`)
1. imperceptible -> reject (invisible tweak / adversarial)
2. score didn't improve -> reject
3. suppression hack -> reject (share up but absolute target salience flat/down)
@@ -150,11 +229,17 @@ def verdict(
== ((sal2_after - sal2_before) > 0)
)
+ oc = outside_change(before, after, target_box) if target_box is not None else None
+
def out(decision: str) -> dict:
- return {"decision": decision, "reasons": reasons, **pc,
+ return {"decision": decision, "reasons": reasons, **pc, "outside": oc,
"ratio_gain": round(ratio_after - ratio_before, 4),
"abs_gain": round((target_sal_after - target_sal_before), 4) if have_abs else None}
+ # First, because SSIM runs on luminance and can't see a pure desaturation.
+ if ratio_gain and oc is not None and degradation(oc):
+ reasons.append("raised the score by degrading the rest of the frame: " + ", ".join(degradation(oc)))
+ return out("reject")
if not pc["perceptible"]:
reasons.append(f"change is imperceptible (ssim {pc['mean_ssim']}/{pc['p1_ssim']}) — likely reward-hack")
return out("reject")
diff --git a/backend/main.py b/backend/main.py
index e33bd79..ff065f7 100644
--- a/backend/main.py
+++ b/backend/main.py
@@ -136,12 +136,13 @@ async def optimize_step(image: UploadFile = File(...), brand: str = Form("the br
vetoed = quality < gemini.settings.judge_gate
# Reward-hack guard: a perceptible, localized edit whose ABSOLUTE on-target salience
- # rose — not a suppression cheat (share up, target salience flat) or an invisible tweak.
+ # rose — not a suppression cheat (share up, target salience flat), an invisible tweak,
+ # or a win bought by darkening, desaturating or blurring the rest of the frame.
guard = eval_guard.verdict(
img, variant,
ratio_before=current, ratio_after=new_score,
target_sal_before=abs_before, target_sal_after=abs_after,
- edit_is_semantic=really_edited,
+ edit_is_semantic=really_edited, target_box=tbox,
)
accepted = guard["decision"] == "accept"
return {
diff --git a/backend/test_eval_guard.py b/backend/test_eval_guard.py
index 44045b8..6abb74d 100644
--- a/backend/test_eval_guard.py
+++ b/backend/test_eval_guard.py
@@ -2,14 +2,20 @@
The guard is the second opinion that keeps the score honest: it rejects edits that
raise the *proxy* without a real on-target improvement. These pin the three cheats
-it must catch, plus the case it must allow. numpy + Pillow only.
+it must catch, plus the case it must allow, then the rest-of-frame cheat (darken,
+desaturate, blur or black out everything but the brand), calibrated on real recorded edits. numpy + Pillow only.
Run: python -m pytest test_eval_guard.py (or) python test_eval_guard.py
"""
from __future__ import annotations
+import base64
+import io
+import json
+from pathlib import Path
+
import numpy as np
-from PIL import Image
+from PIL import Image, ImageFilter
import eval_guard
@@ -83,6 +89,86 @@ def test_perceptual_change_flags_global_vs_local():
assert glob["is_global"]
+# --- the rest-of-frame cheat: win by degrading everything outside the brand ---------
+_BOX = [0.35, 0.35, 0.3, 0.3] # normalized x, y, w, h
+_PX = (70, 70, 130, 130) # the same box in pixels on the 200x200 base
+
+
+def _outside(fn) -> Image.Image:
+ """Apply `fn` to the whole frame, then paste the untouched target back in."""
+ out = fn(_BASE.copy())
+ out.paste(_BASE.crop(_PX), _PX[:2])
+ return out
+
+
+def _gain_verdict(after: Image.Image) -> dict:
+ # Every score signal says "win": share up AND absolute target salience up.
+ return eval_guard.verdict(_BASE, after, ratio_before=0.4, ratio_after=0.6,
+ target_sal_before=0.20, target_sal_after=0.30, target_box=_BOX)
+
+
+def test_darkening_the_rest_is_rejected():
+ dark = _outside(lambda im: Image.fromarray((np.asarray(im) * 0.35).astype("uint8")))
+ v = _gain_verdict(dark)
+ assert v["decision"] == "reject"
+ assert any("darkened" in r for r in v["reasons"])
+
+
+def test_desaturating_the_rest_is_rejected():
+ v = _gain_verdict(_outside(lambda im: im.convert("L").convert("RGB")))
+ assert v["decision"] == "reject"
+ assert any("desaturated" in r for r in v["reasons"])
+
+
+def test_blurring_the_rest_is_rejected():
+ v = _gain_verdict(_outside(lambda im: im.filter(ImageFilter.GaussianBlur(6))))
+ assert v["decision"] == "reject"
+ assert any("blurred" in r for r in v["reasons"])
+
+
+def test_blacking_out_the_rest_is_rejected():
+ v = _gain_verdict(_outside(lambda im: Image.new("RGB", im.size, (0, 0, 0))))
+ assert v["decision"] == "reject"
+
+
+def test_edit_inside_the_target_is_still_accepted():
+ after = _BASE.copy()
+ after.paste((255, 255, 255), (85, 85, 115, 115)) # a bright mark on the brand only
+ assert _gain_verdict(after)["decision"] == "accept"
+
+
+# --- calibration on real recorded edits (skipped if the files aren't there) ----------
+_ROOT = Path(__file__).resolve().parents[1]
+_PUB = _ROOT / "frontend" / "public"
+_NIKE_BOX = [0.3, 0.18, 0.45, 0.3]
+
+
+def test_june_nike_edit_is_rejected():
+ """The old 59 -> 81 Nike 'win' muted everything except the billboard."""
+ run = _PUB / "precomputed" / "nike.json"
+ if not run.exists():
+ return
+ d = json.loads(run.read_text(encoding="utf-8"))
+ img = lambda u: Image.open(io.BytesIO(base64.b64decode(u.split(",", 1)[-1]))).convert("RGB")
+ oc = eval_guard.outside_change(img(d["original_png"]), img(d["variant_png"]), _NIKE_BOX)
+ assert eval_guard.degradation(oc), oc
+
+
+def test_recorded_honest_edits_pass_the_rest_of_frame_check():
+ """Edits that changed the brand and left the scene alone must not trip the check."""
+ cases = [("replay/nike/step3.jpg", "nike", _NIKE_BOX),
+ ("replay/spotify/step1.jpg", "spotify", [0.32, 0.28, 0.4, 0.5]),
+ # clutter removed from the desk: a few cells change, the scene doesn't
+ ("../../results/red-bull/edit1.jpg", "red-bull", [0.34, 0.28, 0.26, 0.5])]
+ for rel, sample, box in cases:
+ after = _PUB / rel
+ if not after.exists():
+ continue
+ before = Image.open(_PUB / "samples" / f"{sample}.jpg")
+ oc = eval_guard.outside_change(before, Image.open(after), box)
+ assert not eval_guard.degradation(oc), (sample, oc)
+
+
if __name__ == "__main__":
fns = [v for k, v in sorted(globals().items()) if k.startswith("test_") and callable(v)]
for fn in fns:
diff --git a/firebase.json b/firebase.json
new file mode 100644
index 0000000..18f904c
--- /dev/null
+++ b/firebase.json
@@ -0,0 +1,13 @@
+{
+ "hosting": {
+ "site": "pixel-gaze",
+ "public": "web-dist",
+ "ignore": ["firebase.json", "**/.*"],
+ "headers": [
+ {
+ "source": "/assets/**",
+ "headers": [{ "key": "Cache-Control", "value": "public, max-age=31536000, immutable" }]
+ }
+ ]
+ }
+}
diff --git a/frontend/public/replay/apple/steps.json b/frontend/public/replay/apple/steps.json
index d36d415..18bb9ee 100644
--- a/frontend/public/replay/apple/steps.json
+++ b/frontend/public/replay/apple/steps.json
@@ -51,9 +51,9 @@
"judge": 0.3,
"judge_reason": "This creative attempts to capture Apple's focus on user experience, using a clean font and the iconic logo. However, several factors make it unlikely to run as a real Apple ad: \n1. **Outdated Product:** The ad features an iPhone X, which is an older model. Apple typically highlights its latest products in primary ad campaigns.\n2. **Quality & Polish:** The 'floating' effect and the glow around the phone appear somewhat amateurish and not seamlessly integrated. Apple ads are known for their extremely high production value, precision, and minimalist perfection, which this creative doesn't quite achieve.\n3. **Generic Message:** While 'EXPERIENCE' is a core tenet of Apple's brand, a single, broad word like this is often too generic for a standalone ad creative. Apple usually pairs its general ethos with specific features, innovations, or emotional benefits ('Privacy. That's iPhone,' 'Shot on iPhone,' 'The new way to interact with your phone').\n4. **Background:** The blurred background, while aiming for a lifestyle feel, lacks the distinctive or aspirational quality typical of Apple's curated environmental shots or their signature clean, often studio-like aesthetic.\n5. **Brand Fit (Aesthetics):** While the Apple logo and text font are on-brand, the overall execution lacks the sophisticated, premium feel Apple meticulously maintains in its advertising.",
"vetoed": true,
- "guard": "accept",
+ "guard": "reject",
"guard_reasons": [
- "perceptible, localized/justified, and target salience improved"
+ "raised the score by degrading the rest of the frame: darkened (0.68x), flattened (0.59x), desaturated (0.50x)"
],
"improved": false,
"n_directives": 8,
@@ -71,11 +71,11 @@
"judge": 0.9,
"judge_reason": "This creative demonstrates strong brand-fit for Apple. The minimalist aesthetic, featuring a clean product shot against a plain background, is a hallmark of Apple's advertising. The typography is simple and elegant, and the concise tagline 'CAPTURE THE MOMENT' directly highlights a key iPhone feature (the camera) and a core user benefit, aligning perfectly with Apple's benefit-driven messaging. While the phone is an older model (iPhone X/XS era), if evaluated within its historical context, the quality of the image and the clarity of the text are high. It's a straightforward, effective, and visually consistent ad that Apple would very likely run to promote the iPhone's camera capabilities.",
"vetoed": false,
- "guard": "accept",
+ "guard": "reject",
"guard_reasons": [
- "perceptible, localized/justified, and target salience improved"
+ "raised the score by degrading the rest of the frame: flattened (0.55x), desaturated (0.30x), blurred (0.49x)"
],
- "improved": true,
+ "improved": false,
"n_directives": 8,
"variant_heatmap": "replay/apple/step3-heat.png"
}
diff --git a/frontend/public/replay/manifest.json b/frontend/public/replay/manifest.json
index ee11f7c..98abae9 100644
--- a/frontend/public/replay/manifest.json
+++ b/frontend/public/replay/manifest.json
@@ -9,7 +9,6 @@
"nike",
"pepsi",
"red-bull",
- "spotify",
- "the-ordinary"
+ "spotify"
]
}
\ No newline at end of file
diff --git a/frontend/public/replay/the-ordinary/heat.png b/frontend/public/replay/the-ordinary/heat.png
deleted file mode 100644
index 94cf340..0000000
Binary files a/frontend/public/replay/the-ordinary/heat.png and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/image.jpg b/frontend/public/replay/the-ordinary/image.jpg
deleted file mode 100644
index e046ab4..0000000
Binary files a/frontend/public/replay/the-ordinary/image.jpg and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/predict.json b/frontend/public/replay/the-ordinary/predict.json
deleted file mode 100644
index 7cfcb19..0000000
--- a/frontend/public/replay/the-ordinary/predict.json
+++ /dev/null
@@ -1,51 +0,0 @@
-{
- "width": 1200,
- "height": 1800,
- "attention_score": 0.7595,
- "heatmap_png": "replay/the-ordinary/heat.png",
- "target_box": [
- 0.29,
- 0.44,
- 0.3,
- 0.26
- ],
- "scanpath": [
- {
- "x": 0.548,
- "y": 0.391,
- "order": 1
- },
- {
- "x": 0.317,
- "y": 0.479,
- "order": 2
- },
- {
- "x": 0.532,
- "y": 0.595,
- "order": 3
- },
- {
- "x": 0.635,
- "y": 0.999,
- "order": 4
- },
- {
- "x": 0.34,
- "y": 0.998,
- "order": 5
- }
- ],
- "distractors": [
- {
- "region": [
- 0.4484,
- 0.2906,
- 0.2,
- 0.2
- ],
- "share": 0.537,
- "desc": "Smiling woman's face"
- }
- ]
-}
\ No newline at end of file
diff --git a/frontend/public/replay/the-ordinary/step0-heat.png b/frontend/public/replay/the-ordinary/step0-heat.png
deleted file mode 100644
index 3623c27..0000000
Binary files a/frontend/public/replay/the-ordinary/step0-heat.png and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/step0.jpg b/frontend/public/replay/the-ordinary/step0.jpg
deleted file mode 100644
index 29f72f6..0000000
Binary files a/frontend/public/replay/the-ordinary/step0.jpg and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/step1-heat.png b/frontend/public/replay/the-ordinary/step1-heat.png
deleted file mode 100644
index 98d6825..0000000
Binary files a/frontend/public/replay/the-ordinary/step1-heat.png and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/step1.jpg b/frontend/public/replay/the-ordinary/step1.jpg
deleted file mode 100644
index 616499a..0000000
Binary files a/frontend/public/replay/the-ordinary/step1.jpg and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/step2-heat.png b/frontend/public/replay/the-ordinary/step2-heat.png
deleted file mode 100644
index 9793edd..0000000
Binary files a/frontend/public/replay/the-ordinary/step2-heat.png and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/step2.jpg b/frontend/public/replay/the-ordinary/step2.jpg
deleted file mode 100644
index 4dd9cc8..0000000
Binary files a/frontend/public/replay/the-ordinary/step2.jpg and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/step3-heat.png b/frontend/public/replay/the-ordinary/step3-heat.png
deleted file mode 100644
index ac37d6e..0000000
Binary files a/frontend/public/replay/the-ordinary/step3-heat.png and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/step3.jpg b/frontend/public/replay/the-ordinary/step3.jpg
deleted file mode 100644
index 80153c7..0000000
Binary files a/frontend/public/replay/the-ordinary/step3.jpg and /dev/null differ
diff --git a/frontend/public/replay/the-ordinary/steps.json b/frontend/public/replay/the-ordinary/steps.json
deleted file mode 100644
index 1ab9fbe..0000000
--- a/frontend/public/replay/the-ordinary/steps.json
+++ /dev/null
@@ -1,82 +0,0 @@
-[
- {
- "step": 0,
- "directive": "make SEVERAL coordinated changes at once to turn this into a polished campaign: remove the clutter and any objects, hands, props or stray text crowding or blocking the The Ordinary logo and product; reframe it head-on and enlarge it as the clear hero; clean and simplify the background; and add a bold, legible on-brand headline, call-to-action and wordmark on the The Ordinary logo and product",
- "variant_png": "replay/the-ordinary/step0.jpg",
- "current_score": 0.7595,
- "new_score": 0.7166,
- "delta": -0.0429,
- "target_salience_before": 0.2465,
- "target_salience_after": 0.1974,
- "judge": 0.2,
- "judge_reason": "This ad creative has some elements that align with The Ordinary's brand, such as the minimalist product display and the 'SIMPLIFY YOUR SKIN' message, which resonates with their straightforward approach. However, it fails significantly on two critical points for a 'real ad':\n\n1. **Placeholder Text on Labels**: The most glaring issue is the use of lorem ipsum-like text on the product labels. The Ordinary is renowned for its transparent and scientific labeling, clearly stating product names (e.g., 'Niacinamide 10% + Zinc 1%') and ingredients. Placeholder text immediately signals that this is not a genuine product advertisement.\n2. **Background Color**: While The Ordinary does use color in some campaigns, their primary aesthetic is often very clinical, utilizing white, grey, or black backgrounds. This particular shade of pink feels a bit off-brand for their typically gender-neutral and scientific image, leaning more towards a conventionally 'feminine' skincare brand.\n\nWhile the composition and idea are clean, the fundamental error of the placeholder text means this ad would absolutely not run as a real ad for The Ordinary, and therefore its quality as a production-ready creative is very low.",
- "vetoed": true,
- "guard": "reject",
- "guard_reasons": [
- "attention score did not improve"
- ],
- "improved": false,
- "n_directives": 7,
- "variant_heatmap": "replay/the-ordinary/step0-heat.png"
- },
- {
- "step": 1,
- "directive": "add a bold, legible, on-brand headline and call-to-action and strengthen the logo/wordmark right at the The Ordinary logo and product as a strong, high-contrast focal point",
- "variant_png": "replay/the-ordinary/step1.jpg",
- "current_score": 0.7595,
- "new_score": 0.5067,
- "delta": -0.2528,
- "target_salience_before": 0.2465,
- "target_salience_after": 0.0802,
- "judge": 0.2,
- "judge_reason": "The ad creative features good visual quality, a relatable model, and clearly displays The Ordinary's iconic products. The messaging around 'Science-Backed' and 'Discover Your Regimen' aligns well with the brand's identity. However, there is a critical and glaring typo: 'SKINKCARE' instead of 'SKINCARE' in the prominent headline. This significant error makes the ad unprofessional and fundamentally unacceptable for a brand that emphasizes scientific precision and attention to detail. This flaw alone would prevent it from running as a real ad, severely undermining its quality despite other strong elements.",
- "vetoed": true,
- "guard": "reject",
- "guard_reasons": [
- "attention score did not improve"
- ],
- "improved": false,
- "n_directives": 7,
- "variant_heatmap": "replay/the-ordinary/step1-heat.png"
- },
- {
- "step": 2,
- "directive": "clearly enlarge, brighten and sharpen the The Ordinary logo and product so it becomes the single biggest, boldest focal element while gently dimming and de-cluttering the surroundings \u2014 a polished, real ad",
- "variant_png": "replay/the-ordinary/step2.jpg",
- "current_score": 0.7595,
- "new_score": 0.537,
- "delta": -0.2225,
- "target_salience_before": 0.2465,
- "target_salience_after": 0.0906,
- "judge": 0.9,
- "judge_reason": "This ad creative demonstrates strong brand-fit for The Ordinary. The tagline 'SCIENCE. SIMPLIFIED. SKINCARE.' perfectly encapsulates the brand's core philosophy of accessible, ingredient-focused products. The amber dropper bottles are iconic and immediately recognizable as The Ordinary's packaging, even without a visible brand logo. The model appears natural and relatable, reinforcing the idea of real people using real products, which aligns with The Ordinary's unpretentious approach. The overall quality of the image is high, with good lighting, clear focus, and legible text. While The Ordinary often uses more minimalist, product-focused ads without models, this creative effectively blends a positive user experience with the brand's scientific messaging and product aesthetic. It would definitely run as a real ad, possibly targeting a slightly broader audience while staying true to its identity.",
- "vetoed": false,
- "guard": "reject",
- "guard_reasons": [
- "attention score did not improve"
- ],
- "improved": false,
- "n_directives": 7,
- "variant_heatmap": "replay/the-ordinary/step2-heat.png"
- },
- {
- "step": 3,
- "directive": "remove or clean away any objects, hands, props or background obstructions that crowd or block the The Ordinary logo and product so the product is fully visible, unobstructed, and the clear hero of the shot",
- "variant_png": "replay/the-ordinary/step3.jpg",
- "current_score": 0.7595,
- "new_score": 0.7398,
- "delta": -0.0197,
- "target_salience_before": 0.2465,
- "target_salience_after": 0.222,
- "judge": 0.95,
- "judge_reason": "This ad creative is excellent in terms of both quality and brand-fit. The minimalist design, clear typography, and focus on the product labels (which highlight the key ingredients) are perfectly aligned with The Ordinary's 'clinical authenticity' and ingredient-first approach. The use of a simple, soft pink background provides a clean, modern aesthetic that makes the amber bottles and white labels pop, without distracting from the products themselves. The 'THE ORDINARY.' branding is prominent at the top, and the 'SHOP NOW' call to action is clear and effective at the bottom. The slight angle of the bottles adds a touch of visual interest without cluttering the minimalist composition. This creative absolutely looks like it would run as a real ad for The Ordinary; it captures their essence perfectly and delivers a clear message.",
- "vetoed": false,
- "guard": "reject",
- "guard_reasons": [
- "attention score did not improve"
- ],
- "improved": false,
- "n_directives": 7,
- "variant_heatmap": "replay/the-ordinary/step3-heat.png"
- }
-]
\ No newline at end of file
diff --git a/frontend/public/replay/the-ordinary/thumb.jpg b/frontend/public/replay/the-ordinary/thumb.jpg
deleted file mode 100644
index 51af9ce..0000000
Binary files a/frontend/public/replay/the-ordinary/thumb.jpg and /dev/null differ
diff --git a/frontend/public/samples/the-ordinary.jpg b/frontend/public/samples/the-ordinary.jpg
deleted file mode 100644
index b22170f..0000000
Binary files a/frontend/public/samples/the-ordinary.jpg and /dev/null differ
diff --git a/frontend/src/samples.ts b/frontend/src/samples.ts
index 1fd9baa..6737de1 100644
--- a/frontend/src/samples.ts
+++ b/frontend/src/samples.ts
@@ -18,16 +18,6 @@ const img = (id: string) => import.meta.env.BASE_URL +
(import.meta.env.MODE === "demo" ? `replay/${id}/image.jpg` : `samples/${id}.jpg`);
export const SAMPLES: Sample[] = [
- {
- id: "the-ordinary",
- brand: "The Ordinary",
- campaign: "Serum droppers — model hero",
- img: img("the-ordinary"),
- target_box: [0.29, 0.44, 0.3, 0.26],
- target_desc: "The two amber serum dropper bottles she's holding, lower-center.",
- tint: "#7A5C3E",
- note: "stock model shot (Pexels) standing in for a serum brand; the face is the attention thief, the product the target — the textbook redirect case (vetted: product 28%→34%, face 64%→58% on one edit).",
- },
{
id: "coca-cola",
brand: "Coca-Cola",
diff --git a/frontend/vite.config.ts b/frontend/vite.config.ts
index 5feef79..698fa20 100644
--- a/frontend/vite.config.ts
+++ b/frontend/vite.config.ts
@@ -9,7 +9,7 @@ import react from "@vitejs/plugin-react";
const API = process.env.VITE_API_BASE || "http://127.0.0.1:8000";
const opt = { target: API, changeOrigin: true, timeout: 600000, proxyTimeout: 600000 };
-// `vite build --mode demo` builds the static replay demo served at sparkylab.web.app/pixel/.
+// `vite build --mode demo` builds the static replay demo served at pixel-gaze.web.app.
const DEMO_OUT = resolve(__dirname, "../web-dist");
// public/ also holds files the demo never loads: the 9 MB June Nike run, and the full-size
@@ -30,7 +30,6 @@ export default defineConfig(({ mode }) => {
const demo = mode === "demo";
return {
plugins: [react(), demo && dropUnusedPublicFiles()],
- base: demo ? "/pixel/" : "/",
build: demo ? { outDir: DEMO_OUT, emptyOutDir: true } : {},
server: {
proxy: {
diff --git a/results/apple/edit1.jpg b/results/apple/edit1.jpg
new file mode 100644
index 0000000..d2981b5
Binary files /dev/null and b/results/apple/edit1.jpg differ
diff --git a/results/apple/run.json b/results/apple/run.json
new file mode 100644
index 0000000..3092fb9
--- /dev/null
+++ b/results/apple/run.json
@@ -0,0 +1,35 @@
+{
+ "sample": "apple",
+ "brand": "Apple",
+ "box": [
+ 0.32,
+ 0.2,
+ 0.36,
+ 0.62
+ ],
+ "base": {
+ "prom": 0.7581,
+ "abs": 0.6983
+ },
+ "edits": [
+ {
+ "k": 1,
+ "directive": "remove the open hand and the wristwatch from the bottom of the photo, filling that area with the same blurred lakeside background. Do not add any text. Keep the iPhone, the lake, the colors and the lighting exactly as they are",
+ "prom": 0.7764,
+ "abs": 0.7737,
+ "judge": 0.7,
+ "judge_reason": "The creative exhibits high image quality and a visually appealing 'floating' effect, which aligns with Apple's aesthetic of sleekness, premium design, and advanced technology. The iPhone X is clearly the focus. While the composition is clean and product-centric, the background is somewhat muted and less aspirational or iconic than what Apple often features in its primary ad campaigns, which tend to use more vibrant or carefully curated environments to tell a story or evoke a specific emotion. Showing a generic home screen also means it doesn't highlight a specific feature or compelling content, which Apple frequently does. It could certainly be used as effective supplementary marketing material, social media content, or as part of a broader campaign, but is slightly less likely to be a standalone hero ad.",
+ "guard": "accept",
+ "guard_reasons": [
+ "perceptible, localized/justified, and target salience improved"
+ ],
+ "outside": {
+ "lum": 0.946,
+ "contrast": 0.923,
+ "color": 0.926,
+ "detail": 0.942
+ },
+ "kept": true
+ }
+ ]
+}
\ No newline at end of file
diff --git a/results/nike/edit1.jpg b/results/nike/edit1.jpg
new file mode 100644
index 0000000..0af1488
Binary files /dev/null and b/results/nike/edit1.jpg differ
diff --git a/results/nike/run.json b/results/nike/run.json
new file mode 100644
index 0000000..ce2dab2
--- /dev/null
+++ b/results/nike/run.json
@@ -0,0 +1,35 @@
+{
+ "sample": "nike",
+ "brand": "Nike",
+ "box": [
+ 0.3,
+ 0.18,
+ 0.45,
+ 0.3
+ ],
+ "base": {
+ "prom": 0.5922,
+ "abs": 0.196
+ },
+ "edits": [
+ {
+ "k": 1,
+ "directive": "make the Nike billboard look like a freshly lit, vivid display: brighter and crisper image of the player, cleaner and sharper. Do not add any text. Do not change anything outside the billboard",
+ "prom": 0.5263,
+ "abs": 0.1499,
+ "judge": 1.0,
+ "judge_reason": "This Nike ad creative is exceptionally high quality and an excellent fit for the brand. It features a recognizable elite athlete (Odell Beckham Jr.) in a dynamic, impactful pose, rendered in a classic black and white style that emphasizes raw athleticism and determination. The iconic 'Just do it.' tagline perfectly complements the image, reinforcing Nike's core message of inspiration and action. This is quintessential Nike advertising \u2013 powerful imagery, celebrity athlete endorsement, and strong brand messaging. It absolutely would run as a real ad, and likely has, given its classic and effective design.",
+ "guard": "reject",
+ "guard_reasons": [
+ "attention score did not improve"
+ ],
+ "outside": {
+ "lum": 1.076,
+ "contrast": 0.98,
+ "color": 0.856,
+ "detail": 1.006
+ },
+ "kept": false
+ }
+ ]
+}
\ No newline at end of file
diff --git a/results/red-bull/edit1.jpg b/results/red-bull/edit1.jpg
new file mode 100644
index 0000000..1c00f6a
Binary files /dev/null and b/results/red-bull/edit1.jpg differ
diff --git a/results/red-bull/edit2.jpg b/results/red-bull/edit2.jpg
new file mode 100644
index 0000000..5e61437
Binary files /dev/null and b/results/red-bull/edit2.jpg differ
diff --git a/results/red-bull/run.json b/results/red-bull/run.json
new file mode 100644
index 0000000..97cfabd
--- /dev/null
+++ b/results/red-bull/run.json
@@ -0,0 +1,54 @@
+{
+ "sample": "red-bull",
+ "brand": "Red Bull",
+ "box": [
+ 0.34,
+ 0.28,
+ 0.26,
+ 0.5
+ ],
+ "base": {
+ "prom": 0.79,
+ "abs": 0.4909
+ },
+ "edits": [
+ {
+ "k": 1,
+ "directive": "remove the black phone, the computer mouse, the keyboard and the drawing tablet from the desk so only the Red Bull can stands on the clean wooden desk. Do not add any text, headline or button. Keep the lighting, colors, camera angle and the can exactly the same",
+ "prom": 0.819,
+ "abs": 0.5906,
+ "judge": 0.6,
+ "judge_reason": "The ad creative is of good quality photographically \u2013 the lighting is pleasant, the product is in focus, and the setting (wooden desk, iMac) is relatable for a productivity scenario. However, its brand-fit for Red Bull is only moderate. While Red Bull is certainly consumed for focus and energy during work or study, this image lacks the dynamic, high-energy, or extreme elements that are central to Red Bull's core 'gives you wings' branding. It's a bit too static and generic, not capturing the aspirational, adventurous, or 'pushing limits' spirit usually associated with Red Bull campaigns. It *could* run as a highly targeted social media ad for productivity, but it wouldn't be a strong hero image for a major campaign, as it doesn't fully leverage Red Bull's distinctive brand identity.",
+ "guard": "accept",
+ "guard_reasons": [
+ "perceptible, localized/justified, and target salience improved"
+ ],
+ "outside": {
+ "lum": 0.995,
+ "contrast": 0.884,
+ "color": 0.978,
+ "detail": 0.906
+ },
+ "kept": true
+ },
+ {
+ "k": 2,
+ "directive": "make the Red Bull can slightly larger and its logo crisper and more vivid. Do not add any text, headline or button. Keep the desk, the lighting and everything else exactly as it is",
+ "prom": 0.7898,
+ "abs": 0.4905,
+ "judge": 0.9,
+ "judge_reason": "The ad creative is of high quality, featuring clear focus on the product, good lighting, and a clean composition. It has excellent brand-fit for Red Bull, portraying a common scenario where the drink is consumed for productivity or focus in a creative/work environment (e.g., with an iMac, graphics tablet). This aligns well with Red Bull's broader messaging beyond extreme sports, targeting students, creatives, and professionals who need an energy boost. It would absolutely run as a real ad, especially for digital or social media campaigns focused on daily energy and concentration.",
+ "guard": "reject",
+ "guard_reasons": [
+ "attention score did not improve"
+ ],
+ "outside": {
+ "lum": 0.995,
+ "contrast": 0.935,
+ "color": 0.97,
+ "detail": 0.929
+ },
+ "kept": false
+ }
+ ]
+}
\ No newline at end of file
diff --git a/results/spotify/edit1.jpg b/results/spotify/edit1.jpg
new file mode 100644
index 0000000..d2064da
Binary files /dev/null and b/results/spotify/edit1.jpg differ
diff --git a/results/spotify/run.json b/results/spotify/run.json
new file mode 100644
index 0000000..b2d8a89
--- /dev/null
+++ b/results/spotify/run.json
@@ -0,0 +1,35 @@
+{
+ "sample": "spotify",
+ "brand": "Spotify",
+ "box": [
+ 0.32,
+ 0.28,
+ 0.4,
+ 0.5
+ ],
+ "base": {
+ "prom": 0.7596,
+ "abs": 0.6324
+ },
+ "edits": [
+ {
+ "k": 1,
+ "directive": "make the green Spotify logo on the phone screen about twice as large and brighter, centered on the screen. Do not add any text. Keep the hand, the phone and the background exactly as they are",
+ "prom": 0.7587,
+ "abs": 0.6292,
+ "judge": 0.9,
+ "judge_reason": "This creative is high quality and has excellent brand-fit for Spotify. The image is clear, well-composed, and features the prominent Spotify logo on a smartphone, which is the primary consumption method for the service. The casual, natural setting (person holding phone outdoors) aligns perfectly with Spotify's brand image of being integrated into daily life. It's a straightforward and effective brand recognition ad that would absolutely run as a real ad, likely for app installs or general brand awareness campaigns. The only minor point preventing a perfect score is that it's a very standard creative approach, not particularly innovative, but highly effective nonetheless.",
+ "guard": "reject",
+ "guard_reasons": [
+ "attention score did not improve"
+ ],
+ "outside": {
+ "lum": 0.962,
+ "contrast": 0.964,
+ "color": 0.973,
+ "detail": 0.96
+ },
+ "kept": false
+ }
+ ]
+}
\ No newline at end of file
diff --git a/results/the-ordinary/edit1.jpg b/results/the-ordinary/edit1.jpg
deleted file mode 100644
index ac6d280..0000000
Binary files a/results/the-ordinary/edit1.jpg and /dev/null differ
diff --git a/results/the-ordinary/edit2.jpg b/results/the-ordinary/edit2.jpg
deleted file mode 100644
index dbc9282..0000000
Binary files a/results/the-ordinary/edit2.jpg and /dev/null differ
diff --git a/results/the-ordinary/edit3.jpg b/results/the-ordinary/edit3.jpg
deleted file mode 100644
index d0607c9..0000000
Binary files a/results/the-ordinary/edit3.jpg and /dev/null differ
diff --git a/results/the-ordinary/edit4.jpg b/results/the-ordinary/edit4.jpg
deleted file mode 100644
index a3f7b4b..0000000
Binary files a/results/the-ordinary/edit4.jpg and /dev/null differ
diff --git a/results/the-ordinary/edit5.jpg b/results/the-ordinary/edit5.jpg
deleted file mode 100644
index 0a30f39..0000000
Binary files a/results/the-ordinary/edit5.jpg and /dev/null differ
diff --git a/results/the-ordinary/run.json b/results/the-ordinary/run.json
deleted file mode 100644
index a55ead5..0000000
--- a/results/the-ordinary/run.json
+++ /dev/null
@@ -1,96 +0,0 @@
-{
- "sample": "the-ordinary",
- "box": [
- 0.29,
- 0.44,
- 0.3,
- 0.26
- ],
- "base": {
- "prom": 0.7595,
- "abs": 0.2465
- },
- "edits": [
- {
- "k": 1,
- "directive": "add a bold, legible, on-brand headline and call-to-action and strengthen the logo/wordmark right at the The Ordinary logo and product as a strong, high-contrast focal point",
- "prom": 0.4969,
- "abs": 0.0771
- },
- {
- "k": 2,
- "directive": "clearly enlarge, brighten and sharpen the The Ordinary logo and product so it becomes the single biggest, boldest focal element while gently dimming and de-cluttering the surroundings \u2014 a polished, real ad",
- "prom": 0.5647,
- "abs": 0.1013
- },
- {
- "k": 3,
- "directive": "reframe the The Ordinary logo and product to a clean, head-on, front-facing hero angle \u2014 square it to the camera so it faces the viewer directly and reads instantly \u2014 keep it in roughly the same spot",
- "prom": 0.7253,
- "abs": 0.2062
- },
- {
- "k": 4,
- "directive": "remove or clean away any objects, hands, props or background obstructions that crowd or block the The Ordinary logo and product so the product is fully visible, unobstructed, and the clear hero of the shot",
- "prom": 0.648,
- "abs": 0.1437
- },
- {
- "k": 5,
- "directive": "make SEVERAL coordinated changes at once to turn this into a polished campaign: remove the clutter and any objects, hands, props or stray text crowding or blocking the The Ordinary logo and product; reframe it head-on and enlarge it as the clear hero; clean and simplify the background; and add a bold, legible on-brand headline, call-to-action and wordmark on the The Ordinary logo and product",
- "prom": 0.6645,
- "abs": 0.1546
- }
- ],
- "tree": [
- {
- "id": 0,
- "parent": null,
- "depth": 0,
- "score": 0.7595,
- "status": "best",
- "directive": "original"
- },
- {
- "id": 1,
- "parent": 0,
- "depth": 1,
- "score": 0.6645,
- "status": "dead",
- "directive": "make SEVERAL coordinated changes at once to turn this into a polished campaign: "
- },
- {
- "id": 2,
- "parent": 0,
- "depth": 1,
- "score": 0.648,
- "status": "dead",
- "directive": "remove or clean away any objects, hands, props or background obstructions that c"
- },
- {
- "id": 3,
- "parent": 0,
- "depth": 1,
- "score": 0.7253,
- "status": "dead",
- "directive": "reframe the The Ordinary logo and product to a clean, head-on, front-facing hero"
- },
- {
- "id": 4,
- "parent": 0,
- "depth": 1,
- "score": 0.4969,
- "status": "dead",
- "directive": "add a bold, legible, on-brand headline and call-to-action and strengthen the log"
- },
- {
- "id": 5,
- "parent": 0,
- "depth": 1,
- "score": 0.5647,
- "status": "dead",
- "directive": "clearly enlarge, brighten and sharpen the The Ordinary logo and product so it be"
- }
- ],
- "final": 0.7595
-}
\ No newline at end of file
diff --git a/scripts/capture_run.py b/scripts/capture_run.py
index 4908c0c..441ea91 100644
--- a/scripts/capture_run.py
+++ b/scripts/capture_run.py
@@ -1,6 +1,6 @@
"""Run the real optimizer on one sample ad and keep every edit it tried.
- python scripts/capture_run.py the-ordinary "The Ordinary"
+ python scripts/capture_run.py red-bull "Red Bull"
Needs GEMINI_API_KEY and the backend deps. Writes results//edit.jpg for each
Nano Banana edit and results//run.json with each edit's directive, its
diff --git a/scripts/make_figures.py b/scripts/make_figures.py
index 3462381..4b3ca1f 100644
--- a/scripts/make_figures.py
+++ b/scripts/make_figures.py
@@ -2,8 +2,8 @@
python scripts/make_figures.py
-Reads the captured winning run in frontend/public/precomputed/nike.json, re-scores
-the before and after with the CURRENT scorer, and writes assets/media/hero-before-after.png.
+Reads the live Red Bull edit saved by scripts/try_edits.py in results/red-bull/, re-scores
+the before and after with the CURRENT scorer and guard, and writes assets/media/hero-before-after.png.
The shared drawing helpers here are also used by make_showcase.py. Every number printed on a figure is measured at build
time, so a figure can never drift away from the code that produced it.
@@ -24,14 +24,14 @@
ROOT = Path(__file__).resolve().parents[1]
OUT = ROOT / "assets" / "media"
-RUN = ROOT / "frontend" / "public" / "precomputed" / "nike.json"
SAMPLES = ROOT / "frontend" / "public" / "samples"
+WIN_SAMPLE, WIN_EDIT = "red-bull", ROOT / "results" / "red-bull" / "edit1.jpg"
# App palette (frontend/src/index.css :root)
BG, INK, MUTED, LINE = "#faf7f1", "#1b1813", "#8a8478", "#ebe4d6"
ACCENT, ACCENT_INK, GOOD, GOOD_WASH, PANEL = "#ee3d23", "#c22d16", "#0e9f6e", "#e6f6ef", "#ffffff"
-NIKE_BOX = [0.3, 0.18, 0.45, 0.3] # frontend/src/samples.ts
+WIN_BOX = [0.34, 0.28, 0.26, 0.5] # Red Bull target box, frontend/src/samples.ts
def font(name: str, size: int):
@@ -63,8 +63,9 @@ def SANSB(s):
return font("calibrib.ttf", s)
-def load_run() -> dict:
- return json.loads(RUN.read_text(encoding="utf-8"))
+def load_pair() -> tuple[Image.Image, Image.Image]:
+ before = Image.open(SAMPLES / (WIN_SAMPLE + ".jpg")).convert("RGB")
+ return before, Image.open(WIN_EDIT).convert("RGB").resize(before.size, Image.LANCZOS)
def data_url_image(url: str) -> Image.Image:
@@ -110,18 +111,17 @@ def overlay_heat(img: Image.Image, density: np.ndarray) -> Image.Image:
def fig_hero(scores: dict | None, dens: dict | None = None):
- """assets/media/hero-before-after.png - the captured winning run, re-scored today.
+ """assets/media/hero-before-after.png - the honest Red Bull win, re-scored today.
With densities, each pane carries its DeepGaze heat layer so the move is visible."""
- run = load_run()
- before, after = data_url_image(run["original_png"]), data_url_image(run["variant_png"])
+ before, after = load_pair()
PW, PH = 620, 930
W, H = 1440, 1330 if scores else 1190
canvas = Image.new("RGB", (W, H), BG)
d = ImageDraw.Draw(canvas)
- d.text((60, 52), "One campaign, one branch of the search", font=DISPLAY(38), fill=INK)
- d.text((60, 106), "Nike billboard with DeepGaze attention on top. Before, the strongest pull is the "
- "Subway sign. After the edit, it moves up onto the player.", font=SERIF(21), fill=MUTED)
+ d.text((60, 52), "One edit that worked", font=DISPLAY(38), fill=INK)
+ d.text((60, 106), "Red Bull with DeepGaze attention on top. The edit takes the phone, mouse, keyboard "
+ "and tablet off the desk and leaves the can and the light alone.", font=SERIF(21), fill=MUTED)
y0 = 172
for i, (label, img) in enumerate((("BEFORE", before), ("AFTER", after))):
@@ -169,17 +169,17 @@ def main():
sys.path.insert(0, str(ROOT / "backend"))
import deepgaze_runner as dg
import eval_guard
- run = load_run()
+ pair = dict(zip(("before", "after"), load_pair()))
scores, dens = {}, {}
- for k, key in (("before", "original_png"), ("after", "variant_png")):
- img = data_url_image(run[key])
- prom, abs_ = dg.score_components(img, NIKE_BOX)
+ for k, img in pair.items():
+ prom, abs_ = dg.score_components(img, WIN_BOX)
scores[k] = {"prom": prom, "abs": abs_}
dens[k] = dg._density(np.asarray(img))
print(" re-scored {}: prominence={} on-target salience={}".format(k, prom, abs_))
- g = eval_guard.verdict(data_url_image(run["original_png"]), data_url_image(run["variant_png"]),
+ g = eval_guard.verdict(pair["before"], pair["after"],
ratio_before=scores["before"]["prom"], ratio_after=scores["after"]["prom"],
- target_sal_before=scores["before"]["abs"], target_sal_after=scores["after"]["abs"])
+ target_sal_before=scores["before"]["abs"], target_sal_after=scores["after"]["abs"],
+ target_box=WIN_BOX)
scores["guard"] = g["decision"]
print(" guard:", g["decision"], g["reasons"])
fig_hero(scores, dens)
diff --git a/scripts/make_showcase.py b/scripts/make_showcase.py
index 45cb704..3ca103b 100644
--- a/scripts/make_showcase.py
+++ b/scripts/make_showcase.py
@@ -23,7 +23,7 @@
sys.path.insert(0, str(ROOT / "backend"))
import deepgaze_runner as dg # noqa: E402
-GIF_ORDER = ["nike", "the-ordinary", "red-bull", "coca-cola"]
+GIF_ORDER = ["nike", "spotify", "red-bull", "coca-cola"]
def load_samples() -> list[dict]:
@@ -179,29 +179,30 @@ def add(im, ms):
("enlarge", "enlarge product"), ("reframe", "head-on reframe"), ("remove", "remove clutter")]
-def fig_every_edit(sample: str = "the-ordinary"):
- """assets/media/every-edit.png - every Nano Banana edit from one captured run,
- each re-scored by DeepGaze against the same fixed brand box."""
- run_dir = ROOT / "results" / sample
- if not (run_dir / "run.json").exists():
- print("every-edit.png skipped: run scripts/capture_run.py first")
- return
+def fig_every_edit(sample: str = "red-bull"):
+ """assets/media/every-edit.png - every Nano Banana edit from one recorded optimizer run
+ (frontend/public/replay, made by scripts/record_demo.py), with the scores DeepGaze gave
+ each one against the same fixed brand box."""
import json
- run = json.loads((run_dir / "run.json").read_text(encoding="utf-8"))
- box = run["box"]
- base = run["base"]["prom"]
+ run_dir = ROOT / "frontend" / "public" / "replay" / sample
+ if not (run_dir / "steps.json").exists():
+ print("every-edit.png skipped: run scripts/record_demo.py first")
+ return
+ steps = json.loads((run_dir / "steps.json").read_text(encoding="utf-8"))
+ box = next(s["box"] for s in load_samples() if s["id"] == sample)
+ base = steps[0]["current_score"]
tiles = [("original", Image.open(SAMPLES / (sample + ".jpg")).convert("RGB"), base)]
- for e in run["edits"]:
- label = next((v for k, v in SHORT if k in e["directive"]), e["directive"][:24])
- tiles.append((label, Image.open(run_dir / "edit{}.jpg".format(e["k"])).convert("RGB"), e["prom"]))
+ for k, e in enumerate(steps):
+ label = next((v for key, v in SHORT if key in e["directive"]), e["directive"][:24])
+ tiles.append((label, Image.open(run_dir / "step{}.jpg".format(k)).convert("RGB"), e["new_score"]))
TW, TH, GAP, M, SS = 220, 330, 18, 50, 2
W = M * 2 + len(tiles) * TW + (len(tiles) - 1) * GAP
H = 150 + TH + 110
canvas = Image.new("RGB", (W, H), BG)
d = ImageDraw.Draw(canvas)
- d.text((M, 42), "Five real edits, all scored lower", font=DISPLAY(36), fill=INK)
- d.text((M, 94), "The Ordinary sample, one live run. Each edit is scored against the same "
+ d.text((M, 42), "{} real edits, all scored lower".format(len(tiles) - 1), font=DISPLAY(36), fill=INK)
+ d.text((M, 94), "Red Bull sample, one recorded run. Each edit is scored against the same "
"brand box. None beat the original, so Pixel kept it.", font=SERIF(19), fill=MUTED)
for i, (label, img, prom) in enumerate(tiles):
x0, y0 = M + i * (TW + GAP), 150
diff --git a/scripts/record_demo.py b/scripts/record_demo.py
index 9fefc42..604bb13 100644
--- a/scripts/record_demo.py
+++ b/scripts/record_demo.py
@@ -2,6 +2,7 @@
python scripts/record_demo.py # every sample that isn't recorded yet
python scripts/record_demo.py nike apple # just these (re-records them)
+ python scripts/record_demo.py --rejudge # re-run the current guard on saved runs
Calls the real FastAPI app in-process, the same way the frontend does: /predict on the
sample, then /optimize/step a few times, re-sending the current best creative after each
@@ -102,7 +103,37 @@ def record(client: TestClient, s: dict) -> None:
(out / "steps.json").write_text(json.dumps(steps, indent=1), encoding="utf-8")
+def rejudge(s: dict) -> None:
+ """Apply the current reward-hack guard to a saved run, using its saved scores and
+ images, so a guard change reaches the demo without paying for new edits."""
+ import eval_guard
+ out = OUT / s["id"]
+ steps = json.loads((out / "steps.json").read_text(encoding="utf-8"))
+ best = Image.open(PUBLIC / "samples" / (s["id"] + ".jpg")).convert("RGB")
+ for k, res in enumerate(steps):
+ after = Image.open(out / "step{}.jpg".format(k)).convert("RGB")
+ v = eval_guard.verdict(best, after, ratio_before=res["current_score"], ratio_after=res["new_score"],
+ target_sal_before=res["target_salience_before"],
+ target_sal_after=res["target_salience_after"], target_box=s["box"])
+ improved = bool(res["new_score"] > res["current_score"] and not res["vetoed"]
+ and v["decision"] == "accept")
+ if res["improved"] and not improved and k < len(steps) - 1:
+ sys.exit("{} step {} is no longer kept, so later steps edited the wrong image; "
+ "re-record it".format(s["id"], k))
+ if improved != res["improved"] or v["decision"] != res["guard"]:
+ print("{:14s} step {} {} -> {} {}".format(s["id"], k, res["guard"], v["decision"], v["reasons"][0]))
+ res.update(guard=v["decision"], guard_reasons=v["reasons"], improved=improved)
+ if improved:
+ best = after
+ (out / "steps.json").write_text(json.dumps(steps, indent=1), encoding="utf-8")
+
+
def main_() -> None:
+ if sys.argv[1:] == ["--rejudge"]:
+ for s in samples():
+ if (OUT / s["id"] / "steps.json").exists():
+ rejudge(s)
+ return
want = sys.argv[1:]
client = TestClient(main.app)
with client:
diff --git a/scripts/try_edits.py b/scripts/try_edits.py
new file mode 100644
index 0000000..196c21e
--- /dev/null
+++ b/scripts/try_edits.py
@@ -0,0 +1,61 @@
+"""Try hand-written edit prompts on one sample ad and keep every result.
+
+ python scripts/try_edits.py red-bull "Red Bull" "remove the phone ..." "..."
+
+Each prompt is one Nano Banana edit of the original. Every edit is re-scored by DeepGaze
+against the sample's fixed brand box, judged for brand fit, and run through the
+reward-hack guard, the same checks /optimize/step applies. Writes results//edit.jpg
+and appends to results//run.json. Needs GEMINI_API_KEY.
+"""
+from __future__ import annotations
+
+import json
+import sys
+from pathlib import Path
+
+from PIL import Image
+
+ROOT = Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT / "backend"))
+sys.path.insert(0, str(ROOT / "scripts"))
+import deepgaze_runner as dg # noqa: E402
+import eval_guard # noqa: E402
+import gemini # noqa: E402
+from capture_run import target_box # noqa: E402
+
+
+def main():
+ sample, brand, prompts = sys.argv[1], sys.argv[2], sys.argv[3:]
+ box = target_box(sample)
+ out = ROOT / "results" / sample
+ out.mkdir(parents=True, exist_ok=True)
+ path = out / "run.json"
+ img = Image.open(ROOT / "frontend" / "public" / "samples" / (sample + ".jpg")).convert("RGB")
+ p0, a0 = dg.score_components(img, box)
+ if dg.engine_name() != "deepgaze-iie":
+ sys.exit("DeepGaze did not load; refusing to score with the fallback engine")
+ run = json.loads(path.read_text(encoding="utf-8")) if path.exists() else {
+ "sample": sample, "brand": brand, "box": box, "base": {"prom": p0, "abs": a0}, "edits": []}
+ print("original prominence {:.1f}".format(p0 * 100))
+ for directive in prompts:
+ variant, desc = gemini.edit_image(img, directive)
+ if desc.startswith("["):
+ sys.exit("edit failed: " + desc[:80])
+ k = len(run["edits"]) + 1
+ variant.convert("RGB").save(out / "edit{}.jpg".format(k), quality=90)
+ prom, abs_ = dg.score_components(variant, box)
+ judge, reason = gemini.judge(variant, brand)
+ v = eval_guard.verdict(img, variant, ratio_before=p0, ratio_after=prom,
+ target_sal_before=a0, target_sal_after=abs_, target_box=box)
+ kept = prom > p0 and judge >= gemini.settings.judge_gate and v["decision"] == "accept"
+ run["edits"].append({"k": k, "directive": directive, "prom": prom, "abs": abs_,
+ "judge": judge, "judge_reason": reason, "guard": v["decision"],
+ "guard_reasons": v["reasons"], "outside": v["outside"], "kept": kept})
+ path.write_text(json.dumps(run, indent=1), encoding="utf-8")
+ print("edit {} prominence {:.1f} ({:+.1f}) salience {:.3f} -> {:.3f} judge {:.2f} guard {} {}".format(
+ k, prom * 100, (prom - p0) * 100, a0, abs_, judge, v["decision"], "KEPT" if kept else ""))
+ print(" ", v["reasons"][0])
+
+
+if __name__ == "__main__":
+ main()