This repository provides an enterprise grade, resource efficient observability stack designed for production environments. It implements full monitoring of the four golden signals (latency, traffic, errors, and saturation) and establishes OpenTelemetry OTLP ingestion standards for application logs, metrics, and distributed traces.
The stack is composed of the following core systems:
- Grafana: Visualizes metrics, logs, and traces through centralized dashboards.
- Prometheus: Collects time series metrics from hosts and containers.
- Loki: Provides log aggregation using TSDB indexing and local filesystem storage.
- Grafana Alloy: Exposes OTLP HTTP and gRPC endpoints VM-wide to receive application logs, metrics, and traces.
- Grafana Tempo: Aggregates and stores distributed traces from microservices and LLM applications.
- Node Exporter: Exports host level performance statistics (CPU, memory, disk, network).
- cAdvisor: Gathers container level metrics and resource utilization details.
All configurations are modular and decoupled:
observability/
alloy: Grafana Alloy collector pipeline configuration
backup: Backup and restore utilities
dashboards: Preconfigured Grafana dashboards
grafana: Configuration and provisioning rules
loki: Log aggregation database settings
prometheus: Metric scraping and alert rules
tempo: Grafana Tempo tracing backend configuration
- Docker Engine and Docker Compose V2.
- Target user must be a member of the docker group.
Create the environment file from the template:
cp .env.example .envOpen .env to configure version tags, port numbers, and database retention policies. Ensure you change the default Grafana admin password before deploying.
Deploy the stack in detached mode:
docker compose up -d- Grafana:
http://localhost:3030(Default credentials: admin / your password) - Grafana Alloy UI:
http://localhost:12345(Collector health and pipeline graph status) - OTLP gRPC Endpoint: Port
4317(Used by applications to send traces, metrics, and logs) - OTLP HTTP Endpoint: Port
4318(Alternative OTLP HTTP push endpoint)
This stack is production ready for Dokploy deployment using Traefik:
- Create a Docker Compose application in your Dokploy panel.
- Connect your Git repository.
- Configure domains for your services using the Dokploy UI. Map your Grafana domain to the
grafanaservice on container port3000. Map your OTLP endpoints toalloyon ports4317and4318. - Deploy the stack. Dokploy handles the Traefik routing, TLS certificate generation, and network isolation automatically.
The Loki pipeline integrates with Tempo through derived fields to automatically link logs to traces.
- Regex Extraction: The Loki data source uses the expression
trace_id[=:\s"]+(\w+)to parse trace identifiers from logs. - Direct Navigation: Clicking the View Trace button next to a log line opens the Tempo trace timeline split-screen instantly.
- Trace-to-Logs: When viewing a trace, you can search corresponding logs for that specific trace or span with one click.
An SRE test application is provided in test_ai.py to verify the end-to-end flow of logs, metrics, and traces. To run it locally on the host:
uv run --with python-dotenv --with openai --with opentelemetry-api --with opentelemetry-sdk --with opentelemetry-exporter-otlp --with opentelemetry-instrumentation-openai python test_ai.pyThis script performs the following actions:
- Connects to OpenAI using your API key from the local
.env. - Generates traces and logs structured under the OpenTelemetry GenAI Semantic Conventions.
- Flushes the telemetry directly to Alloy, which routes them to Loki and Tempo.
To support deployment on small hosts, CPU and memory boundaries are set on every container:
- Prometheus: Limited to 1200MB memory, 1.00 CPU. Retention set to 14 days or 10GB.
- Loki: Limited to 1000MB memory, 1.00 CPU. Retention set to 7 days.
- Grafana: Limited to 400MB memory, 0.50 CPU.
- Grafana Alloy: Limited to 250MB memory, 0.50 CPU.
- Grafana Tempo: Limited to 400MB memory, 0.50 CPU.
- cAdvisor: Limited to 150MB memory, 0.40 CPU.
- Node Exporter: Limited to 50MB memory, 0.20 CPU.
The total memory reservation is optimized for a 4GB RAM instance, leaving a safe buffer.
A shell script handles automated volume state archiving.
Run the utility:
./backup/backup.shThe backup workflow performs these steps:
- Compresses configuration directories and the
.envfile into a configuration archive. - Mounts active Docker volumes in read only mode to a temporary container to generate database archives.
- Stores all compressed files under the backup directory.
- Applies a 7 day retention window, deleting older files.
To restore the stack:
-
Unpack the configuration files:
tar -xzf backup/archive/config_timestamp.tar.gz -C /home/ubuntu/lemma/observability-stack
-
Provision the Docker volumes:
docker volume create lemma_prometheus_data docker volume create lemma_loki_data docker volume create lemma_grafana_data docker volume create lemma_tempo_data
-
Unpack database contents back into the volumes:
docker run --rm -v lemma_prometheus_data:/volume -v /home/ubuntu/lemma/observability-stack/backup/archive:/backup alpine sh -c "tar -xzf /backup/lemma_prometheus_data_timestamp.tar.gz -C /volume" docker run --rm -v lemma_loki_data:/volume -v /home/ubuntu/lemma/observability-stack/backup/archive:/backup alpine sh -c "tar -xzf /backup/lemma_loki_data_timestamp.tar.gz -C /volume" docker run --rm -v lemma_grafana_data:/volume -v /home/ubuntu/lemma/observability-stack/backup/archive:/backup alpine sh -c "tar -xzf /backup/lemma_grafana_data_timestamp.tar.gz -C /volume" docker run --rm -v lemma_tempo_data:/volume -v /home/ubuntu/lemma/observability-stack/backup/archive:/backup alpine sh -c "tar -xzf /backup/lemma_tempo_data_timestamp.tar.gz -C /volume"
-
Bring the stack up:
docker compose up -d