Skip to content
STiFLeR7Public

About

Gradia is a zero-latency monitoring dashboard that runs directly alongside your training script. No cloud uploads, no signup forms, just pure telemetry.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

11 Commits

Folders and files

Repository files navigation

G R A D I A

Next-Generation Local-First ML Training Visualization

PyPI Version Python Version License Build Status

From observing training to understanding learning.

Gradia Dashboard


πŸš€ What's New in v2.0.0

Gradia v2.0.0 introduces the Learning Timeline β€” a real-time, sample-centric view of how your models learn over time. This release transforms Gradia from a metrics dashboard into a learning behavior explorer.

✨ Flagship Feature: Learning Timeline

The Learning Timeline answers questions that aggregate metrics cannot:

  • 🎯 When did this sample become correctly classified?
  • πŸ”„ Which samples keep flipping predictions?
  • πŸ“ˆ Is the model memorizing or stabilizing?
  • ⚠️ Which data points drive learning instability?

Learning Timeline


πŸ“– Overview

Gradia is a high-performance, local-first monitoring solution for machine learning workflows. Unlike cloud-native platforms, Gradia focuses on zero-latency, privacy-first tracking that runs directly alongside your training loop.

Built on FastAPI and a Reactive UI, Gradia provides granular visibility into your model's training dynamics, system resources, and now β€” individual sample learning behavior.


⚑ Key Features

Feature Description
πŸ”¬ Learning Timeline Track how individual samples evolve during training with real-time visualization
πŸ“Š Real-Time Telemetry Nanosecond-precision tracking of Loss, Accuracy, and custom metrics
🧠 Intelligent Auto-Discovery Automatic task type inference (Classification vs Regression) and model suggestions
πŸ’» System Profiling CPU and RAM monitoring during training epochs
πŸ“ Artifact Management Automated checkpointing and structured logging (events.jsonl)
πŸ“‹ Comprehensive Reporting One-click PDF/JSON reports with full training history
πŸ”„ Backward Compatible Full support for v1.x runs with automatic migration

πŸ”¬ Learning Timeline Deep Dive

How It Works

The Learning Timeline tracks a bounded subset of samples (default: 100) throughout training, capturing:

  • Prediction β€” What the model predicts for each sample
  • Confidence β€” Model's certainty in its prediction
  • Correctness β€” Whether the prediction matches the true label
  • Flip Events β€” When predictions change between epochs

Sample Classification

Gradia automatically classifies tracked samples into categories:

Category Description Visual
Stable Correct Consistently correct predictions 🟒 Green
Late Learner Became correct after epoch N 🟑 Yellow
Unstable Predictions flip frequently 🟠 Orange
Persistent Error Never correctly classified πŸ”΄ Red

UI Blocks

The Timeline interface is organized into focused blocks:

  1. Block A: Timeline Overview β€” High-level view of learning stability across all tracked samples
  2. Block B: Sample Inspector β€” Deep-dive into individual sample trajectories with confidence curves
  3. Block C: Instability Panel β€” Top flipping samples, late learners, and persistent errors
  4. Block D: Training Context β€” Current epoch, status, and tracking metadata

πŸ› οΈ Architecture

Gradia employs a Producer-Consumer architecture:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Trainer Thread │───▢│   Event Queue   │───▢│   FastAPI UI    β”‚
β”‚   (Producer)    β”‚    β”‚ (Thread-Safe)   β”‚    β”‚   (Consumer)    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
        β”‚                                              β”‚
        β–Ό                                              β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Sample Tracker  β”‚                          β”‚ Timeline Logger β”‚
β”‚   (v2.0 New)    β”‚                          β”‚   (v2.0 New)    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

New v2.0 Components

  • gradia.events β€” Event model with LearningEvent, SampleState, EpochSummary
  • SampleTracker β€” Boundary-aware sample selection and tracking
  • TimelineLogger β€” Structured timeline event persistence
  • SchemaMigrator β€” Automatic v1.x to v2.0 config migration

πŸ“¦ Installation

pip install gradia --upgrade

From Source

git clone https://github.com/STiFLeR7/gradia.git
cd gradia
pip install -e ".[dev]"

πŸ’» Quick Start

Basic Usage

# Auto-detect datasets and start the dashboard
gradia run .

Advanced CLI

# Specify target column and port
gradia run . --target "label" --port 8080

Python API

from gradia.trainer.engine import Trainer
from gradia.core.scenario import ScenarioInferrer
from gradia.core.config import ConfigManager

# Infer scenario from dataset
inferrer = ScenarioInferrer()
scenario = inferrer.infer("data.csv", target_override="label")

# Configure training with timeline enabled
config_mgr = ConfigManager("./runs")
config = config_mgr.load_or_create()
config['model']['type'] = 'random_forest'
config['training']['epochs'] = 20
config['timeline']['enabled'] = True
config['timeline']['max_samples'] = 100

# Run training
trainer = Trainer(scenario, config, "./runs")
trainer.run()

# Get timeline insights
insights = trainer.get_timeline_insights()
print(f"Stable samples: {insights['stable_correct']}")
print(f"Flipping samples: {insights['top_flippers']}")

πŸ“Š Dashboard

Access the dashboard at http://localhost:8000 after running gradia run .

Pages

Page URL Description
Configure /configure Select model, hyperparameters, and start training
Metrics / Real-time training metrics and system resources
Timeline /timeline Learning Timeline visualization (v2.0)

Configuration Options

# Example gradia_config.yaml (auto-generated)
schema_version: "2.0"
project_name: "my-experiment"
save_model: true

model:
  type: "random_forest"
  params:
    n_estimators: 100
    max_depth: null

training:
  epochs: 20
  test_split: 0.2
  random_seed: 42

timeline:
  enabled: true
  max_samples: 100
  sampling_strategy: "boundary"

πŸ”„ Migration from v1.x

Gradia v2.0 is fully backward compatible. When you run gradia run . on a v1.x project:

  1. Existing configs are automatically migrated to v2.0 schema
  2. Old runs remain accessible
  3. Timeline features are enabled by default
# Migration happens automatically
gradia run .
# Output: Config migrated: Added timeline config, Set schema_version to 2.0

πŸ§ͺ Testing

# Run all tests
pytest tests/ -v

# Run with coverage
pytest tests/ --cov=gradia --cov-report=html

πŸ“ Project Structure

gradia/
β”œβ”€β”€ cli/              # Typer CLI application
β”œβ”€β”€ core/             # Configuration, inspection, migration
β”œβ”€β”€ events/           # v2.0 Event model and tracking
β”‚   β”œβ”€β”€ models.py     # LearningEvent, SampleState, EpochSummary
β”‚   β”œβ”€β”€ tracker.py    # SampleTracker with boundary sampling
β”‚   └── logger.py     # TimelineLogger for event persistence
β”œβ”€β”€ models/           # sklearn wrappers and model factory
β”œβ”€β”€ trainer/          # Training engine with timeline integration
└── viz/              # FastAPI server and UI templates
    β”œβ”€β”€ templates/    # Jinja2 HTML templates
    └── static/       # CSS and JavaScript

πŸ—ΊοΈ Roadmap

v2.1 (Planned)

  • WebSocket real-time updates
  • Dataset Intelligence Panel
  • Experiment Comparison (overlay 2-3 runs)
  • Export timeline to video/GIF

v2.2 (Future)

  • PyTorch integration
  • TensorFlow/Keras support
  • Remote monitoring mode

🀝 Contributing

We welcome contributions! Please see CONTRIBUTING.md for guidelines.

# Development setup
git clone https://github.com/STiFLeR7/gradia.git
cd gradia
pip install -e ".[dev]"

# Run tests
pytest tests/ -v

# Lint
flake8 gradia/

πŸ“„ License

Distributed under the MIT License. See LICENSE for more information.


πŸ”— Links


Built with ❀️ by STiFLeR for the ML Community.

Instagram β€’ Hugging Face β€’ PyPI

About

Gradia is a zero-latency monitoring dashboard that runs directly alongside your training script. No cloud uploads, no signup forms, just pure telemetry.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages