ModelPulse is a Python library and CLI tool that implements an end-to-end partial-weight transfer pipeline for running large language models (LLMs) across a two-device setup over a local network. While it is optimized for edge computing scenarios (such as a Raspberry Pi or low-powered device), it can be used across any two devices where a client (Device B) loads and runs GGUF-format language models that are hosted, managed, and pushed by a server (Device A).
The core innovation is a zero-disk, delta-patching architecture: models are transferred as individual tensor shards, assembled in RAM, and loaded through llama.cpp, without ever writing the full model to persistent storage. When a model is updated, only the changed tensor shards are retransmitted, and the live model is hot-swapped in memory with no client restart required.
Requirements: Python ≥ 3.10, Linux (POSIX)
pip install modelpulseOr install from source:
git clone https://github.com/MdSufiyan005/ModelPulse
cd ModelPulse
pip install -e .modelpulse server convert ./my-model.gguf ./shards/my-model/modelpulse server run --host 0.0.0.0 --port 8000 --shard-dir ./models-storagemodelpulse server upload my-model ./shards/my-model/ --server http://192.168.1.10:8000Run inference with a prompt:
modelpulse bridge run http://192.168.1.10:8000 --prompt "Explain edge computing."Run the built-in benchmark suite:
modelpulse bridge run http://192.168.1.10:8000 --benchmarkListen for future model updates (no prompt):
modelpulse bridge run http://192.168.1.10:8000After fine-tuning your model, convert the updated version and upload only the changed shards:
# Auto-diff: compare new shard directory against old one
modelpulse server upload my-model-v2 ./shards/my-model-v2/ \
--base my-model \
--base-dir ./shards/my-model/ \
--server http://192.168.1.10:8000Any connected clients will automatically receive and apply the delta no restart needed.
| Package | Role |
|---|---|
fastapi + uvicorn |
REST + WebSocket server (Device A) |
httpx |
Async HTTP client (Device B) |
websockets |
WebSocket client (Device B) |
llama-cpp-python |
Local LLM inference via llama.cpp |
typer |
CLI framework |
rich |
TUI output (panels, progress bars, tables) |
psutil |
RAM, CPU temperature, and utilization monitoring |
huggingface_hub |
(Available for model downloads) |
numpy + scipy |
Numerical utilities |
For a deeper dive into ModelPulse's inner workings, refer to README-detailed.md. It covers:
- Architecture & Project Structure
- Key Concepts (GGUF Sharding, Delta Updates, Zero-Disk Loading, Control Plane)
- CLI and REST API Reference
- Shard and Manifest Formats
- WebSocket Protocol
- Metrics & Telemetry
MIT : see pyproject.toml for details.
Author: Mohammad Sufiyan (moahmmadsufiyan152@gmail.com)
