Skip to content

About

Artifact of CONQuER: Compiler Optimisation for Network Quantisation with Evolutionary Refinement

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

CONQuER: Compiler Optimisation for Network Quantisation with Evolutionary Refinement

Overview

Deploying deep neural networks on resource-constrained hardware relies heavily on Mixed-Precision Quantisation (MPQ). However, traditional quantisation happens as a front-end preprocessing step, entirely disconnected from the downstream compilers that generate physical machine code. This separation leads to suboptimal configurations where assigned bit-widths map poorly to target hardware execution blocks.

CONQuER is a unified, compiler-integrated infrastructure for hardware-aware MPQ. It shifts quantisation directly into the compiler pipeline (at the MLIR TOSA level) to enable intelligent configuration handling based on true compiler support. To navigate the exponentially large combinatorial search space of layer-wise bit-width configurations, CONQuER couples an NSGA-II evolutionary algorithm with a surrogate pre-screening engine, ensuring optimal execution on the target device.


Repository Structure

The project bridges a Python-based search algorithm with a C++ MLIR compiler infrastructure:

  • conquer-opt/: The core C++ compiler infrastructure. Built on MLIR and integrated with the IREE compiler backend. Contains the custom MLIR passes for applying mixed-precision quantisation directly to intermediate representations.
  • conquer-search/: The Python-based search engine. Contains the NSGA-II Genetic Algorithm (ga.py), surrogate cost models (cost_model.py), and the orchestrator (conquer.py).
  • hardware/: JSON profiles representing the target execution environments (e.g., A100 GPUs, Snapdragon CPUs/Vulkan, EPYC CPUs) used by the search algorithm.
  • examples/: Example networks (ResNet, MobileNet, EfficientNet) and preprocessing scripts used for evaluation.
  • scripts/: Automation scripts for setting up the Docker environment, building the compiler, and running evaluations.

Setup & Building (Docker)

To ensure ease of use, a Docker setup is provided, for information on native installs please refer to the Dockerfile. Additionally, please note, the Dockerfile has only been tested on a machine with an Intel CPU, Nvidia and AMD deployments are untested with this release, however, the system has run on both previously, so this should work.

Prerequisites

  • Docker installed and running.
  • Git installed on the host machine.

Building the Environment

We provide a unified setup script that initialises all submodules (including the IREE compiler) and builds the isolated Docker container matching an Ubuntu 24.04 environment.

  1. Clone the repository and navigate to the root directory.
  2. Make the setup script executable:
   chmod +x scripts/00_setup.sh
  1. Run the setup script. By default, this builds a clean, lightweight CPU-only container:
./scripts/00_setup.sh

Build Arguments (Hardware & Testing)

The setup script accepts flags to compile the testing suite or inject specific hardware SDKs (like CUDA or Vulkan) into the Docker container. (only the base version of the Dockerfile without cuda, rocm, or vulkan has been tested).

  • --test: Compiles the C++ GoogleTests and executes them inside a temporary test container.
  • --vulkan: Installs Vulkan headers/tools and compiles the compiler with Vulkan backend support.
  • --cuda: Installs the NVIDIA CUDA Toolkit into the container and compiles with CUDA support (Note: Adds ~3GB to the image size).
  • --rocm: Installs AMD ROCm development packages and compiles with ROCm support.

Example: Build with tests and Vulkan support:

./scripts/00_setup.sh --test --vulkan

Accessing the Container

Once the initialisation is complete, you can drop into the container using:

docker run -it --rm conquer:latest

The compiled conquer-opt binary will be available inside the container at /app/build/conquer-opt.


Running the System

The repository includes a set of automated scripts to verify that the core components (Compiler, Genetic Algorithm, and Backend Evaluation) are functional.

1. Artifact Evaluation (GA Search)

This script performs an end-to-end dry run: downloading the dataset, exporting the MLIR model, and running a 10-generation GA search.

# Execute the search pipeline
./scripts/01_run_ga.sh

2. Compiling and Evaluating Policies

After the search, you can compile the best policies found by the GA and evaluate them for accuracy and latency:

# Compile the top Pareto solutions into IREE modules
./scripts/02_compile_policies.sh

# Benchmark latency and evaluate accuracy drop for compiled policies
./scripts/03_evaluate_policies.sh --gens 10

Artifact Evaluation Scope

Note to Reviewers: This artifact demonstrates the functionality of the CONQuER infrastructure. The provided scripts are configured for a lightweight, representative evaluation (ResNet-18, 10 GA generations) to ensure the artifact remains exercisable within a standard reviewer environment. Full-scale HIL validation results reported in the paper require significant computational resources, and a representative hardware configuration for the given platform, it is strongly recommended to run this GA on a host machine with a strong CPU for compilation however, that process is tricky and very sensitive to setup, so was ommitted here.


License

Apache 2.0

About

Artifact of CONQuER: Compiler Optimisation for Network Quantisation with Evolutionary Refinement

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages