Skip to content

Latest commit

Β 

History

17 Commits

Folders and files

Repository files navigation

πŸš€ AFP-GIC: Controllable Generative Image Compression

✨ Good news! AFP-GIC pretrained weights are now available on Hugging Face. Download the model and use our inference code to compress and decompress your own images. See the code example below.

πŸ› οΈ Training code released! The three-stage training pipeline is now available, with stage-by-stage commands.

IEEE Access paper arXiv and supplementary material Hugging Face live demo Download pretrained checkpoint Code Ocean reproducible capsule

AFP-GIC is designed to make generative image compression content-adaptive and reduce hallucinations: invented details that do not match the original image. It adapts the visual knowledge transferred from a pretrained model to each image, guiding compression and reconstruction toward more realistic textures and better preservation of the original content, even at very low bitrates.

Our paper, Adaptive Fused Prior Transfer for Controllable Generative Image Compression, published in IEEE Access (2026), presents this content-adaptive design with five bitrate operating points in a single deployable pretrained model, without transmitting the fused prior or reloading model weights between operating points. System-level benchmarks on an NVIDIA RTX 4090 demonstrate 18.1% lower decoder latency and 31.1M fewer inference parameters than DC-VIC. Explore the released code and checkpoint, or upload your own images to experience compression and reconstruction in the live demo below.

AFP-GIC interactive demo: original and reconstructed image comparison, bitrate controls, and downloadable results.

Try it with your own images: choose an operating point, compress, and compare the reconstruction side by side. Download the actual compressed bitstream and decode it in the demo. The demo runs on CPU by default.

Try the live demo β†’

✨ Highlights

  • Multi-rate compression without a collection of models: one checkpoint covers all five reported operating points, simplifying model management.
  • Adaptive prior guidance without prior transmission: transfer image-adaptive knowledge from frozen AdaCode to guide encoding and predict the fused prior at the decoder.
  • 18.1% lower decoder latency: 80.47 ms versus 98.27 ms for DC-VIC.
  • 20.5% fewer inference parameters: 120.6M versus 151.7M, a reduction of 31.1M parameters.
  • From paper to hands-on evaluation: custom image uploads, downloadable bitstreams, standalone decompression, and per-image benchmark CSVs make the results accessible beyond the paper.

⚑ Efficiency

Method Inference parameters Encoder latency Decoder latency
DC-VIC 151.7M 61.61 ms 98.27 ms
AFP-GIC 120.6M 81.34 ms 80.47 ms

Benchmark: NVIDIA RTX 4090, 100 DIV2K patches of 256 x 256 pixels (paper, Table 5). Parameter counts include frozen components. These system-level measurements are separate from the CPU-hosted demo's response time.

🧩 Architecture

Continuing our research on learned image compression, AFP-GIC combines adaptive fused-prior transfer and single-model bitrate control in an asymmetric architecture. A frozen AdaCode model supplies image-adaptive guidance to the encoder; the decoder predicts the fused prior from the compressed representation instead of receiving it as side information.

AFP-GIC architecture: adaptive fused-prior guidance at the encoder and prior prediction at the decoder.

Figure 1. Overview of AFP-GIC. Blue and red indicate encoding and decoding, respectively; snowflakes and flames denote frozen and trainable modules.

πŸ–ΌοΈ Visual Comparisons

Original images and low-bitrate reconstructions from VVC Intra, MS-ILLM, CRDR, DC-VIC, and AFP-GIC on Kodak, including enlarged details.

Figure 5. Low-bitrate visual comparisons on Kodak. Baselines are shown at their closest available released bitrates; each image is labeled with its actual bpp. Images and annotations are reproduced from the paper.

πŸ“Š Paper Metrics

The metrics directory provides CSV results for all five AFP-GIC operating points on Kodak, CLIC2020, and DIV2K: dataset summaries, paper-rounded values, and 2,760 per-image records including PSNR, SSIM, MS-SSIM, SNR, LPIPS, DISTS, and NIQE. See the data description for the evaluation records and aggregation checks. FID is provided as a dataset-level metric.

πŸ“₯ Reconstructed Images and Metrics

To make research comparisons easier, we provide all 2,760 reconstructed images, per-image metrics, and dataset-average metrics in our GitHub Releases, covering 24 Kodak, 428 CLIC2020, and 100 DIV2K images at five bitrate operating points. Download the datasets and operating points you need to include AFP-GIC as a baseline under matched evaluation protocols, without rerunning the pretrained model.

Code Ocean Reproducible Capsule

Run the Kodak evaluation on Code Ocean with the pretrained model, input images, and configured environment. The published capsule evaluates all 24 Kodak images at five operating points and provides reconstructed images, metrics, and comparisons with the paper's reference values.

πŸ› οΈ Installation

This repository provides pretrained model inference, evaluation tools, an interactive demo, and the three-stage training pipeline. The instructions below set up inference and evaluation; see Training for the separate training environment.

The release was tested with Python 3.9, PyTorch 2.1.0, and torchvision 0.16.0. Create a separate environment and install the pinned dependencies:

git clone https://github.com/yifeipet/AFP_GIC.git
cd AFP_GIC
conda create -n afp-gic python=3.9 -y
conda activate afp-gic
python -m pip install -r public_release/requirements.txt

For GPU evaluation, use a PyTorch build compatible with your GPU and driver. Follow the official PyTorch installation instructions for the pinned version if a platform-specific build is needed. Do not replace the pinned versions with the latest releases when reproducing the paper.

πŸ“¦ Pretrained Model

Download the released checkpoint from Google Drive and place it at:

checkpoint/afp_gic_release/model/afp_gic_release.pth.tar

The checkpoint already includes the frozen prior component; no separate AdaCode weight download is required.

🧠 Hugging Face Model: Compress Your Own Images

After completing Installation, run the following from the repository root. The Hugging Face Hub provides the pretrained weights; the AFP-GIC code performs compression and decompression.

python -m pip install huggingface_hub
from pathlib import Path
import struct
import sys

import torch
from huggingface_hub import hf_hub_download

runtime = Path("public_release/runtime").resolve()
sys.path.insert(0, str(runtime))
import eval_public_release as afp

device = "cuda:0" if torch.cuda.is_available() else "cpu"
weights = hf_hub_download(
    repo_id="yifeipet/AFP-GIC",
    filename="afp_gic_release.pth.tar",
)
config = afp.load_infer_config(
    str(runtime / "config/afp_gic_release.yaml"), device
)
model = afp.build_comp_model(config).to(device)
model.load_learned_weight(ckpt_path=weights)
model.codec_setup()
model.eval()

# Encode your image at operating point 0 (choose 0 through 4).
with torch.no_grad():
    image = afp.read_real_tensor("input.png")
    parts = model.compress(image, quality_ind=0)["string_list"]
Path("compressed.afp").write_bytes(
    b"".join(struct.pack("<I", len(part)) + part for part in parts)
)

# Decode the saved file. The original image is not needed here.
data = Path("compressed.afp").read_bytes()
parts, offset = [], 0
for _ in range(3):
    if offset + 4 > len(data):
        raise ValueError("Truncated bitstream")
    size = struct.unpack_from("<I", data, offset)[0]
    offset += 4
    if size == 0 or offset + size > len(data):
        raise ValueError("Invalid payload length")
    parts.append(data[offset:offset + size])
    offset += size
if offset != len(data):
    raise ValueError("Unexpected trailing data")
with torch.no_grad():
    reconstruction, _, _ = model.decompress(parts)
afp.img_utils.imwrite("reconstruction.png", reconstruction)

Replace input.png with your image path. Outputs are compressed.afp and reconstruction.png. Encoding and decoding can run separately after the same model setup; decoding needs the bitstream and compatible weights, not the original image. This example reads files you created yourself, not untrusted uploads. Images are reconstructed lossily, and large inputs require more memory.

πŸ§ͺ Evaluation

πŸ“· Kodak

Obtain the Kodak dataset and place its 24 original PNG images directly in datasets/kodak/.

From the repository root, evaluate one operating point:

python public_release/test.py -d cuda:0 --dataset kodak --qualities 0

Evaluate all five operating points:

python public_release/test.py -d cuda:0 --dataset kodak --qualities 0 1 2 3 4
Quality index 0 1 2 3 4
Nominal target bpp 0.050 0.075 0.100 0.125 0.150

Actual bitrates vary with image content. All five indices use the same checkpoint.

πŸ—‚οΈ Other Datasets

The entry point also accepts clic2020_test and div2k_valid_hr. Place the original PNG images in the corresponding directories:

datasets/
|-- kodak/                              # Kodak PNG images
|-- CLIC/
|   `-- clic_test_images/               # CLIC2020 test PNG images
`-- DIV2K_valid_HR/
    `-- DIV2K_valid_HR/                 # DIV2K validation PNG images
python public_release/test.py -d cuda:0 --dataset clic2020_test --qualities 0 1 2 3 4
python public_release/test.py -d cuda:0 --dataset div2k_valid_hr --qualities 0 1 2 3 4

πŸ“ Outputs

The runner saves reconstructed images, actual bitrates, per-image metrics, and summary files. To choose an output directory:

python public_release/test.py -d cuda:0 --dataset kodak --qualities 0 --results-root results/kodak_demo

For this command, outputs include:

results/kodak_demo/
|-- summary_all.csv
|-- comparison_pivot.csv
`-- kodak/afp_gic_release/q0/
    |-- <image_name>.png
    |-- _bitrates.csv
    |-- per_image_metrics.csv
    |-- _metrics.json
    `-- summary.json

Metric protocols: per_image_metrics.csv records in-loop PSNR, MS-SSIM, and LPIPS; _metrics.json records metrics computed on saved reconstructions. These evaluation paths can produce different values. See the paper and Supplementary Material for the reporting protocols, and Paper Metrics for the released benchmark CSVs.

Run python public_release/test.py --help for the available options. Metric libraries may download their pretrained weights on first use.

πŸ‹οΈ Training

The training implementation provides three stages and seven steps, from initial training to a model fine-tuned on five selected control pairs.

Stage Steps Purpose
I 1–3 High-rate warmup, dual-control rate-distortion training, and adversarial training; 500K iterations per step
II 1–3 Prepare validation crops, search control settings, and select five control pairs; no gradient updates
III 1 Fine-tune on the five selected pairs for 500K iterations

Start with the training guide for environment setup, dataset preparation, and commands for all seven steps. Training uses a separate environment from inference.

  1. Install the dependencies in training/requirements.txt.
  2. Prepare the datasets and download AdaCode_S2_model_g.pth.
  3. Set the local paths in training/settings.yaml, then run the stages in order.

From the repository root, preview the first-stage command before starting training:

python training/run.py stage1-step1 --device cuda:0 --dry-run

Prior cosine loss is disabled from the start in all supplied training configurations; prior-consistency MSE remains enabled. AdaCode weights retain their upstream license and attribution.

πŸ“ Citation

If you find our work useful in your research, please cite our official IEEE Access paper:

@article{pei2026adaptive,
  title   = {Adaptive Fused Prior Transfer for Controllable Generative Image Compression},
  author  = {Pei, Yifei and Liu, Ying and Ling, Nam},
  journal = {IEEE Access},
  year    = {2026},
  doi     = {10.1109/ACCESS.2026.3737467},
  url     = {https://ieeexplore.ieee.org/document/11712133}
}

🀝 Acknowledgments and License

AFP-GIC builds on DC-VIC and AdaCode, with supporting components from BasicSR and CompressAI. We thank their authors for making these resources available.

Original AFP-GIC additions are provided for research and evaluation use. Third-party components retain their respective licenses; no single permissive license applies uniformly to this repository. Please consult LICENSE and THIRD_PARTY_NOTICES.md before reuse or redistribution.

For questions about this release, please open a GitHub issue.