Skip to content

Repository files navigation

ModelPulse

ModelPulse is a Python library and CLI tool that implements an end-to-end partial-weight transfer pipeline for running large language models (LLMs) across a two-device setup over a local network. While it is optimized for edge computing scenarios (such as a Raspberry Pi or low-powered device), it can be used across any two devices where a client (Device B) loads and runs GGUF-format language models that are hosted, managed, and pushed by a server (Device A).

ModelPulse Architecture

The core innovation is a zero-disk, delta-patching architecture: models are transferred as individual tensor shards, assembled in RAM, and loaded through llama.cpp, without ever writing the full model to persistent storage. When a model is updated, only the changed tensor shards are retransmitted, and the live model is hot-swapped in memory with no client restart required.

Table of Contents


Installation

Requirements: Python ≥ 3.10, Linux (POSIX)

pip install modelpulse

Or install from source:

git clone https://github.com/MdSufiyan005/ModelPulse
cd ModelPulse
pip install -e .

Quick Start

1. Convert a GGUF model to shards

modelpulse server convert ./my-model.gguf ./shards/my-model/

2. Start the server (Device A)

modelpulse server run --host 0.0.0.0 --port 8000 --shard-dir ./models-storage

3. Upload the model to the server

modelpulse server upload my-model ./shards/my-model/ --server http://192.168.1.10:8000

4. Start the client (Device B)

Run inference with a prompt:

modelpulse bridge run http://192.168.1.10:8000 --prompt "Explain edge computing."

Run the built-in benchmark suite:

modelpulse bridge run http://192.168.1.10:8000 --benchmark

Listen for future model updates (no prompt):

modelpulse bridge run http://192.168.1.10:8000

5. Push a delta update

After fine-tuning your model, convert the updated version and upload only the changed shards:

# Auto-diff: compare new shard directory against old one
modelpulse server upload my-model-v2 ./shards/my-model-v2/ \
  --base my-model \
  --base-dir ./shards/my-model/ \
  --server http://192.168.1.10:8000

Any connected clients will automatically receive and apply the delta no restart needed.


Dependencies

Package Role
fastapi + uvicorn REST + WebSocket server (Device A)
httpx Async HTTP client (Device B)
websockets WebSocket client (Device B)
llama-cpp-python Local LLM inference via llama.cpp
typer CLI framework
rich TUI output (panels, progress bars, tables)
psutil RAM, CPU temperature, and utilization monitoring
huggingface_hub (Available for model downloads)
numpy + scipy Numerical utilities

Detailed Documentation

For a deeper dive into ModelPulse's inner workings, refer to README-detailed.md. It covers:

  • Architecture & Project Structure
  • Key Concepts (GGUF Sharding, Delta Updates, Zero-Disk Loading, Control Plane)
  • CLI and REST API Reference
  • Shard and Manifest Formats
  • WebSocket Protocol
  • Metrics & Telemetry

License

MIT : see pyproject.toml for details.
Author: Mohammad Sufiyan (moahmmadsufiyan152@gmail.com)

About

End-to-end partial-weight transfer pipeline.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages