Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
50b2431
Guard: reject edits that win by darkening, desaturating or blurring t…
EnesYilmazcode Oct 1, 2026
884ea8b
Test the rest-of-frame check, including the June Nike edit
EnesYilmazcode Oct 1, 2026
47b653d
record_demo: add --rejudge to re-run the guard on saved runs
EnesYilmazcode Oct 1, 2026
a4addc9
Re-judge the recorded runs: the Apple background swap is no longer kept
EnesYilmazcode Oct 1, 2026
5e8fcc8
Guard: compare the rest of the frame cell by cell so removing clutter…
EnesYilmazcode Oct 1, 2026
2b5b15e
Test that a real clutter removal passes the rest-of-frame check
EnesYilmazcode Oct 1, 2026
d49ce60
Re-judge the Apple run with the cell-based check
EnesYilmazcode Oct 1, 2026
2cb0697
Add a script to try hand-written edit prompts and keep every result
EnesYilmazcode Oct 1, 2026
9b8d8fb
Save live edits that don't add text or dim the scene; the Red Bull de…
EnesYilmazcode Oct 1, 2026
3706920
Drop The Ordinary from the sample ads and the demo
EnesYilmazcode Oct 1, 2026
249937f
Hero figure: use the Red Bull desk clean-up and the current guard
EnesYilmazcode Oct 1, 2026
25f23f1
Showcase: use the Red Bull run for the every-edit strip, drop The Ord…
EnesYilmazcode Oct 1, 2026
eb542d8
Rebuild the README figures
EnesYilmazcode Oct 1, 2026
80a3080
README: the Red Bull clean-up as the win, its losing run as the every…
EnesYilmazcode Oct 1, 2026
5548253
Serve the demo from the site root
EnesYilmazcode Oct 1, 2026
5a8e783
Add Firebase Hosting config for the pixel-gaze site
EnesYilmazcode Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .firebaserc
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
{
"projects": {
"default": "buildwithsparky"
}
}
29 changes: 15 additions & 14 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@
Pixel runs a gaze model over an ad, finds what is pulling the eye away from the brand, and has a team of AI agents redesign it. Every edit is re-scored by the gaze model, and edits that don't raise the score are thrown out.</p>

<p align="center">
<a href="https://sparkylab.web.app/pixel/"><img src="assets/media/deepgaze.gif" width="400" alt="DeepGaze attention heatmaps and predicted fixation order drawn over four sample ads"></a><br>
<a href="https://sparkylab.web.app/pixel/"><b>Try the demo</b></a>
<a href="https://pixel-gaze.web.app/"><img src="assets/media/deepgaze.gif" width="400" alt="DeepGaze attention heatmaps and predicted fixation order drawn over four sample ads"></a><br>
<a href="https://pixel-gaze.web.app/"><b>Try the demo</b></a>
</p>

<p align="center">Built in one day at Multimodal Hacks (NY Tech Week, June 6 2026) with <a href="https://github.com/rishis123">Rishi Shah</a>.</p>
Expand All @@ -14,7 +14,7 @@ Pixel runs a gaze model over an ad, finds what is pulling the eye away from the

1. **Look.** [DeepGaze IIE](https://github.com/matthias-k/DeepGaze), a neural saliency model trained on real eye-tracking data, predicts how likely each pixel of the ad is to be looked at.
2. **Score.** Pixel measures how much of that attention lands on the brand (the logo, the product, the call to action) compared with how big the brand is. 50 means the brand gets exactly its fair share for its size. Higher means it pulls more than its size.
3. **Find the thieves.** The strongest spots outside the brand are the things stealing attention. Gemini looks at each one and names it ("Subway sign", "smiling woman's face").
3. **Find the thieves.** The strongest spots outside the brand are the things stealing attention. Gemini looks at each one and names it ("Subway sign", "smartwatch band").
4. **Edit.** Agents propose changes, Gemini's image model ("Nano Banana") makes them, and DeepGaze re-scores every result. A brand-fit judge and a reward-hack guard can veto an edit. You watch each branch land and decide whether to grow another.

<p align="center"><img src="assets/media/app.png" width="820" alt="The Pixel app analyzing the Nike billboard: attention on target 59, with the Subway sign named as the main attention thief"></p>
Expand All @@ -23,32 +23,33 @@ That is the app on the Nike sample. The billboard is the ad, but the first predi

## What the gaze model sees

<p align="center"><img src="assets/media/gallery.png" width="820" alt="DeepGaze heatmaps and fixation order on all eight sample ads"></p>
<p align="center"><img src="assets/media/gallery.png" width="820" alt="DeepGaze heatmaps and fixation order on all seven sample ads"></p>

A few things the model picks up that a person would also notice:

- **Faces win.** On The Ordinary ad, the woman's face gets the first fixation and more than half the attention in the frame. The serum bottles she is selling come second.
- **Text is a magnet.** On McDonald's, the eye walks along the lit nameplate before it touches the arches.
- **Clutter competes.** On Red Bull, the phone and the mouse on the desk each pull a fixation away from the can.
- **Red on red disappears.** The Coca-Cola can scores well, but its heat barely shows against the red background. The model still finds the script logo.

## One edit that worked

<p align="center"><img src="assets/media/hero-before-after.png" width="820" alt="Nike billboard before and after an edit, with DeepGaze heat on both: attention moves from the Subway sign onto the player"></p>
<p align="center"><img src="assets/media/hero-before-after.png" width="820" alt="Red Bull can on a desk before and after an edit that removes the phone, keyboard, tablet and mouse, with DeepGaze heat on both"></p>

The search on the Nike billboard tried six edits over two rounds. Darkening only the slogan or only the store sign barely helped. Muting everything except the logo and product, then boosting their contrast and sharpness, did a lot: attention on target went from 59 to 81, and the absolute amount of attention on the brand almost tripled. Those numbers are recomputed from the saved images every time the figure is built, by [`scripts/make_figures.py`](scripts/make_figures.py).
The Red Bull can shares the desk with a phone, a keyboard, a drawing tablet and a mouse, and the mouse pulls a fixation of its own. One live edit asked Nano Banana to clear the desk and change nothing else: no text, same light, same can. Attention on target went from 79 to 82, and the share of all attention that lands on the can went from 49% to 59%. Those numbers are recomputed from the saved images every time the figure is built, by [`scripts/make_figures.py`](scripts/make_figures.py).

It is a small win, and it is the kind Pixel now accepts. An earlier version of this section showed a Nike edit that went from 59 to 81 by muting everything except the billboard. That is a cheat, not a better ad, and the guard below now rejects it.

## Most edits don't work, and Pixel says so

<p align="center"><img src="assets/media/every-edit.png" width="820" alt="The Ordinary ad and five Nano Banana edits of it, each scored lower than the original"></p>
<p align="center"><img src="assets/media/every-edit.png" width="820" alt="The Red Bull ad and four Nano Banana edits of it, each scored lower than the original"></p>

This is a live run on The Ordinary ad from September 2026, with every edit saved. Each one looks more like a finished campaign than the original, and each one scored lower. A new headline gives the eye something else to read. Two of the edits moved the bottles, and the brand box does not move with them. So Pixel kept the original and reported a change of zero.
This is a recorded optimizer run on the same Red Bull ad, with every edit saved. Each one looks more like a finished campaign than the original, and each one scored lower. Every one of them added a headline or a slogan, and text pulls the eye off the can. So Pixel kept the original and reported a change of zero. The edit that won above added no text at all.

The first version of Pixel could not do this. It measured raw attention inside the brand box, so "make the logo bigger" always won, and it floored every result at the baseline, so the number could only go up. The version here fixes both:

- **The score is size-invariant.** Attention share is divided by the box's share of the frame, so enlarging the target stops being a free win.
- **Losing edits stay visible.** The before and after show the real change, including zero or negative.
- **A guard catches two cheats** ([`backend/eval_guard.py`](backend/eval_guard.py)). If the brand's share of attention went up but the absolute attention on it did not, the edit won by dimming everything else and is rejected. If the edit is too small to see, it is rejected as a likely adversarial trick.
- **A guard catches three cheats** ([`backend/eval_guard.py`](backend/eval_guard.py)). If the brand's share of attention went up but the absolute attention on it did not, the edit is rejected. If the edit is too small to see, it is rejected as a likely adversarial trick. And if the rest of the frame came out darker, flatter, grayer or blurrier than before, the edit won by degrading the scene and is rejected. That last check compares the frame outside the brand box cell by cell, so removing one cluttering object passes and dimming the whole scene does not.
- **The badge tells the truth.** If DeepGaze fails to load and the cheap fallback runs instead, `/health` says so and the app shows DEMO instead of LIVE.

The full list of ways the score can still be gamed, and what is done about each, is in [`docs/EVAL_FLAWS.md`](docs/EVAL_FLAWS.md).
Expand Down Expand Up @@ -86,14 +87,14 @@ npm run dev # http://localhost:5173

Keys go in git-ignored files: `GEMINI_API_KEY`, `PINECONE_API_KEY` and `PINECONE_INDEX` in `backend/.env`, and `VITE_CLERK_PUBLISHABLE_KEY` in `frontend/.env.local`. Seed the competitor index once with `python backend/pinecone_seed.py`. Without a Gemini key the app still analyzes ads, but edits come back unchanged. DeepGaze runs on the GPU when there is one and on the CPU otherwise.

The [live demo](https://sparkylab.web.app/pixel/) is the same frontend replaying runs recorded on the real backend, so it needs no server and no keys. `python scripts/record_demo.py` records them and `npm run build:demo` in `frontend` writes the static site to `web-dist/`.
The [live demo](https://pixel-gaze.web.app/) is the same frontend replaying runs recorded on the real backend, so it needs no server and no keys. `python scripts/record_demo.py` records them and `npm run build:demo` in `frontend` writes the static site to `web-dist/`.

To rebuild the figures in this README:

```bash
python scripts/make_showcase.py # gallery, GIF, every-edit strip
python scripts/make_figures.py # Nike before and after
python scripts/capture_run.py the-ordinary "The Ordinary" # a new live run (needs a Gemini key)
python scripts/make_figures.py # Red Bull before and after
python scripts/try_edits.py red-bull "Red Bull" "<edit prompt>" # a new live edit (needs a Gemini key)
python backend/test_scoring.py && python backend/test_eval_guard.py
```

Expand All @@ -102,7 +103,7 @@ python backend/test_scoring.py && python backend/test_eval_guard.py
| [`backend/`](backend/) | FastAPI app, DeepGaze runner and score, agents, branch search, reward-hack guard, tests |
| [`frontend/`](frontend/) | Vite, React and TypeScript app, the sample ads, the saved Nike run |
| [`scripts/`](scripts/) | Figure builders and the live-run capture |
| [`results/`](results/) | The Ordinary run: every edit and its scores |
| [`results/`](results/) | Live edits with their scores, Judge and guard verdicts, including the Red Bull win |
| [`docs/`](docs/) | How the score can be gamed and what stops it |

## Credits
Expand Down
Binary file modified assets/media/deepgaze.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified assets/media/every-edit.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified assets/media/gallery.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified assets/media/hero-before-after.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
2 changes: 1 addition & 1 deletion backend/agents.py
Original file line number Diff line number Diff line change
Expand Up @@ -182,7 +182,7 @@ def _step_finalize(state: dict) -> dict:
ratio_before=baseline, ratio_after=final,
target_sal_before=before.get("target_salience"),
target_sal_after=after.get("target_salience"),
edit_is_semantic=True,
edit_is_semantic=True, target_box=target,
)

steps = [
Expand Down
89 changes: 87 additions & 2 deletions backend/eval_guard.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,11 @@
and (b) **imperceptible / global** tweaks the eye can't read as a design change.
This module gives the loop a cheap second opinion that catches those.

A third cheat is the one the absolute-salience check misses: darken, desaturate, blur
or black out everything except the brand. DeepGaze is a probability map, so pulling
mass off the rest of the frame raises the brand's absolute share too. Pass
`target_box` and the guard compares the rest of the frame before and after.

Wiring (Pixel): `deepgaze_runner.score_components(image, box)` returns both the
size-invariant prominence (the headline `ratio`) AND the absolute on-target salience
mass — feed both into `verdict(...)`:
Expand All @@ -17,7 +22,7 @@
before_img, after_img,
ratio_before=ratio_before, ratio_after=ratio_after,
target_sal_before=abs_before, target_sal_after=abs_after,
edit_is_semantic=really_edited,
edit_is_semantic=really_edited, target_box=box,
)
accepted = bool(v["decision"] == "accept") # gate on this, surface v["reasons"]

Expand All @@ -42,6 +47,12 @@
GLOBAL_COVERAGE = 0.55 # fraction of pixels changed above which an edit looks "global"
GLOBAL_BBOX = 0.60 # change bounding-box covering more of the frame than this = "global"
EPS = 1e-4
GRID = 8 # cells per side for the rest-of-frame comparison
OUTSIDE_MARGIN = 0.04 # grow the target box by this before measuring "the rest of the frame"
# Floors for the rest of the frame, as after/before ratios. Below any of these, the edit
# raised the score by degrading everything that isn't the brand. Calibrated on the
# recorded edits in frontend/public/replay and results/ (see test_eval_guard.py).
OUTSIDE_FLOORS = {"lum": 0.80, "contrast": 0.75, "color": 0.70, "detail": 0.70}


def _gray(img: Image.Image, size=None) -> np.ndarray:
Expand Down Expand Up @@ -115,6 +126,71 @@ def perceptual_change(before: Image.Image, after: Image.Image) -> dict:
}


def _rgb(img: Image.Image, size) -> np.ndarray:
return np.asarray(img.convert("RGB").resize(size, Image.BILINEAR), dtype=np.float64)


def _outside_mask(h: int, w: int, box, margin: float = OUTSIDE_MARGIN) -> np.ndarray:
"""True outside the target box, grown by `margin` (fraction of the frame) on every side
so a legit edit that spills slightly past the box isn't counted as 'the rest'."""
x, y, bw, bh = box
x0, y0 = int(max(0.0, x - margin) * w), int(max(0.0, y - margin) * h)
x1, y1 = int(min(1.0, x + bw + margin) * w), int(min(1.0, y + bh + margin) * h)
m = np.ones((h, w), bool)
m[y0:y1, x0:x1] = False
return m


def _stats(rgb: np.ndarray, mask: np.ndarray) -> dict:
lum = rgb @ np.array([0.299, 0.587, 0.114])
rg = rgb[..., 0] - rgb[..., 1]
yb = 0.5 * (rgb[..., 0] + rgb[..., 1]) - rgb[..., 2]
gy, gx = np.gradient(lum)
grad = np.hypot(gx, gy)
l, r, b = lum[mask], rg[mask], yb[mask]
return {
"lum": l.mean(),
"contrast": l.std(),
# Hasler & Suesstrunk colorfulness
"color": np.hypot(r.std(), b.std()) + 0.3 * np.hypot(r.mean(), b.mean()),
"detail": grad[mask].mean(),
}


def outside_change(before: Image.Image, after: Image.Image, box) -> dict:
"""How the rest of the frame (outside the target) changed, as after/before ratios of
luminance, luminance contrast, colorfulness and edge detail. Each ratio is the MEDIAN
over a grid of cells, so removing one cluttering object (a few cells change) passes,
while dimming, flattening, desaturating or blurring the whole scene (most cells
change) does not. A ratio well below 1 means the edit degraded the rest of the frame."""
w, h = before.size
scale = min(1.0, MAX_SIDE / max(w, h))
size = (max(8, int(w * scale)), max(8, int(h * scale)))
a, b = _rgb(before, size), _rgb(after, size)
mask = _outside_mask(size[1], size[0], box)
ratios: dict[str, list[float]] = {"lum": [], "contrast": [], "color": [], "detail": []}
ys = np.linspace(0, size[1], GRID + 1).astype(int)
xs = np.linspace(0, size[0], GRID + 1).astype(int)
for i in range(GRID):
for j in range(GRID):
cell = (slice(ys[i], ys[i + 1]), slice(xs[j], xs[j + 1]))
m = mask[cell]
if m.mean() < 0.5: # mostly target; not "the rest of the frame"
continue
sa, sb = _stats(a[cell], m), _stats(b[cell], m)
for k in ratios:
ratios[k].append((sb[k] + 1.0) / (sa[k] + 1.0))
if not ratios["lum"]: # target fills the frame; nothing outside to judge
return {k: 1.0 for k in ratios}
return {k: round(float(np.median(v)), 3) for k, v in ratios.items()}


def degradation(oc: dict) -> list[str]:
"""Which of the rest-of-frame signals fell past its floor."""
words = {"lum": "darkened", "contrast": "flattened", "color": "desaturated", "detail": "blurred"}
return [f"{words[k]} ({oc[k]:.2f}x)" for k, floor in OUTSIDE_FLOORS.items() if oc[k] < floor]


def verdict(
before: Image.Image,
after: Image.Image,
Expand All @@ -126,10 +202,13 @@ def verdict(
sal2_before: float | None = None,
sal2_after: float | None = None,
edit_is_semantic: bool = True,
target_box=None,
) -> dict:
"""Decide accept / reject / review for one edit. Gate `accepted` on decision == 'accept'.

Priority of checks (reasons explain every outcome):
0. rest of frame degraded -> reject (score rose, but the frame outside the target got
darker / flatter / grayer / blurrier; needs `target_box`)
1. imperceptible -> reject (invisible tweak / adversarial)
2. score didn't improve -> reject
3. suppression hack -> reject (share up but absolute target salience flat/down)
Expand All @@ -150,11 +229,17 @@ def verdict(
== ((sal2_after - sal2_before) > 0)
)

oc = outside_change(before, after, target_box) if target_box is not None else None

def out(decision: str) -> dict:
return {"decision": decision, "reasons": reasons, **pc,
return {"decision": decision, "reasons": reasons, **pc, "outside": oc,
"ratio_gain": round(ratio_after - ratio_before, 4),
"abs_gain": round((target_sal_after - target_sal_before), 4) if have_abs else None}

# First, because SSIM runs on luminance and can't see a pure desaturation.
if ratio_gain and oc is not None and degradation(oc):
reasons.append("raised the score by degrading the rest of the frame: " + ", ".join(degradation(oc)))
return out("reject")
if not pc["perceptible"]:
reasons.append(f"change is imperceptible (ssim {pc['mean_ssim']}/{pc['p1_ssim']}) — likely reward-hack")
return out("reject")
Expand Down
5 changes: 3 additions & 2 deletions backend/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -136,12 +136,13 @@ async def optimize_step(image: UploadFile = File(...), brand: str = Form("the br
vetoed = quality < gemini.settings.judge_gate

# Reward-hack guard: a perceptible, localized edit whose ABSOLUTE on-target salience
# rose — not a suppression cheat (share up, target salience flat) or an invisible tweak.
# rose — not a suppression cheat (share up, target salience flat), an invisible tweak,
# or a win bought by darkening, desaturating or blurring the rest of the frame.
guard = eval_guard.verdict(
img, variant,
ratio_before=current, ratio_after=new_score,
target_sal_before=abs_before, target_sal_after=abs_after,
edit_is_semantic=really_edited,
edit_is_semantic=really_edited, target_box=tbox,
)
accepted = guard["decision"] == "accept"
return {
Expand Down
Loading
Loading