Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Image as Score

English | 简体中文

A failed—but useful—experiment in translating personal photographs into editorial graphic-score posters.

image-as-score is a Codex skill that deconstructs a photograph, selectively preserves its identity, and reconstructs it as an experimental-music score poster. It tries to make something that feels collectible first and recognizably derived from the source photograph second.

Image as Score reference result

What it does

The skill treats the source image as a set of relationships rather than a collection of objects:

  • mass, spacing, direction, repetition, atmosphere, and material;
  • one or two recognizable identity-bearing structures;
  • a small number of interacting compositional systems;
  • exact notation changing into diffuse or granular states;
  • a real-photo banner, restrained title, and score-dominant page hierarchy.

Its working sequence is:

DECONSTRUCT → SELECTIVE PRESERVATION → ABSTRACT / DISTILL → RECONSTRUCT

The bundled reference is passed to image generation as a macro-composition and art-direction reference. The skill also includes a pre-generation Gate Card, a prompt compiler, and a post-generation visual review gate.

Installation

Copy or clone this repository into your personal Codex skills directory:

mkdir -p ~/.agents/skills
git clone https://github.com/ninggele/image-as-score ~/.agents/skills/image-as-score

Then ask Codex to use $image-as-score with an attached photograph.

Requirements

  • Codex with image-generation support
  • A source photograph supplied by the user
  • The bundled assets/style-reference.png
  • Human visual review of the actual generated bitmap

The skill is intentionally not a deterministic renderer. It does not generate score layouts with JSON, coordinates, or programmatic drawing.

Honest limitations

This project does not yet generalize reliably enough to be called a successful universal photo-to-score system.

1. It is too large

The skill accumulated a long Gate Card, prompt compiler, failure matrix, and maintenance contract while trying to prevent repeated visual failures. Much of that knowledge is useful, but the total instruction surface is heavy. It consumes substantial context and can make the workflow slow and difficult to maintain.

This is the clearest engineering failure of the project: each local correction made the system more capable, but also made the skill larger and more fragile.

2. It currently works best with wide, horizontally composed photographs

The strongest result was built from a landscape beach photograph with:

  • a long horizon;
  • distributed figures;
  • a distant architectural register;
  • broad empty atmosphere;
  • one scarce red event.

That source maps unusually well to the bundled reference. Wide landscapes, architectural panoramas, and other horizontal compositions are therefore the safest inputs.

Portrait-oriented photographs, close-up faces, vertically stacked scenes, compact interiors, and sources without a convincing horizontal duration may be forced into a page grammar that does not belong to them. The skill contains conditional rules intended to reduce this bias, but those rules have not been validated across enough real sources.

3. The reference is also an overfitting risk

The bundled style reference is powerful, but it can dominate the source image. Common failure modes include:

  • inventing skyline-like structures for unrelated subjects;
  • turning identity into generic ticks, dots, staff lines, or particles;
  • reusing the same curve, wave, grain, and red-event hierarchy;
  • preserving surface style while losing the photograph's real identity;
  • producing a visually competent poster that is still too close to the reference's composition.

4. Source identity is not guaranteed

Image-generation models may reconstruct rather than preserve the photographic banner, alter people or architecture, or weaken the source-specific identity voice. The skill includes a non-generative pixel-replacement safeguard, but the final result still requires active inspection.

Prompt compliance is not proof that the image works.

5. It is expensive in judgment and iteration

The workflow asks for thumbnail comparison, original-size inspection, reference comparison, and Tier 2 diagnosis. A first render can look polished while still failing structurally. The skill limits generation to two calls, so unresolved failures may remain.

6. The result is an artwork, not playable notation

The generated score is visual and editorial. It is not guaranteed to be consistently interpretable, performable, or useful as formal music notation.

Why publish a failed experiment?

Because several ideas survived the failure:

  • preserve source identity before abstraction;
  • translate relationships, not objects;
  • separate internal analysis from the model-facing prompt;
  • judge the real bitmap rather than the prompt;
  • treat a second generation as a targeted repair, not a free alternative;
  • stop after version one when it already works;
  • document failure modes as visible evidence rather than vague quality language.

These methods may be more reusable than the current graphic-score style itself.

Repository structure

image-as-score/
├── README.md
├── README.zh-CN.md
├── SKILL.md
├── agents/
│   └── openai.yaml
├── assets/
│   └── style-reference.png
└── references/
    ├── art-direction-prompt.md
    └── maintenance-contract.md

Feedback wanted

I am especially interested in results from:

  • portrait photographs;
  • people and pets;
  • still life and interiors;
  • plants and natural textures;
  • industrial details;
  • photographs without a strong horizontal axis or accent color.

If you try it, please share the source type, what remained recognizable, what became generic, and whether the final image felt like an artwork rather than a filter.

License status

The Skill instructions and documentation are released under the MIT License. assets/style-reference.png is licensed separately under CC BY-NC 4.0. See LICENSE and assets/LICENSE.md.

About

An experimental Codex skill that deconstructs personal photographs and reconstructs them as editorial graphic-score posters.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors