Skip to content

Latest commit

 

History

History
402 lines (260 loc) · 20.5 KB

File metadata and controls

402 lines (260 loc) · 20.5 KB

SENTRY: Core Code Reference

This document highlights the most important syntaxes, functions, and architectural patterns used in the SENTRY project. Every code block is followed by a plain-English walkthrough explaining exactly what is happening and why.


1. Python: Data Processing & Machine Learning

OS Input Hooking (pynput)

SENTRY uses pynput to silently monitor OS events without requiring window focus.

from pynput import mouse, keyboard

# Non-blocking background listeners
mouse_listener = mouse.Listener(on_move=on_move, on_click=on_click, on_scroll=on_scroll)
keyboard_listener = keyboard.Listener(on_press=on_press, on_release=on_release)

mouse_listener.start()
keyboard_listener.start()

What is happening here?

pynput is a library that talks directly to the operating system's input event system — the same pipeline that moves your cursor on screen. When you call .start(), it creates a new background thread that runs silently forever alongside the rest of the program.

The mouse.Listener takes three callback functions:

  • on_move — called every time the mouse moves even one pixel. It receives the (x, y) coordinates.
  • on_click — called when any mouse button is pressed or released.
  • on_scroll — called when the scroll wheel is used.

The keyboard.Listener takes two callbacks:

  • on_press — called the moment a key is pushed down.
  • on_release — called the moment the key is let go.

Why non-blocking? The start() method puts the listener in its own thread, so it doesn't stop the rest of the program from running. The main logic (feature extraction, ML scoring) can proceed while input events stream in continuously in the background.

Why does SENTRY need this and not just track window input? Standard web or app input only fires events when that window is focused. pynput hooks directly into the OS kernel event stream, so it works even when the user is in another application — which is essential for passive, continuous authentication.


Feature Mathematics (numpy, scipy)

Converting raw coordinates into behavioural features heavily relies on fast array math.

import numpy as np
from scipy.stats import entropy

# Calculate variance/spread of data points (e.g., flight times)
variance = np.std(flight_times)

# Calculate entropy (randomness) of directional bins
counts = np.bincount(directions, minlength=8)
probs = counts / np.sum(counts)
shannon_entropy = entropy(probs, base=2)

What is np.std(flight_times)?

std stands for standard deviation — a number that describes how spread out a list of values is.

Imagine you record the time gap between each keystroke during a 3-second session. If you always type at a very consistent speed, these gaps are all similar (e.g., [0.12, 0.11, 0.13, 0.12]). The standard deviation is tiny — nearly 0. If you're erratic or mashing keys, the gaps jump all over the place (e.g., [0.05, 0.40, 0.02, 0.90]). The standard deviation is large.

SENTRY uses np.std for KDV (key dwell variability) and DFT (digram flight time) — both are just measures of spread/consistency in your typing rhythm.

What is np.bincount + entropy?

This is how SENTRY computes Saccade Entropy (SE) — the randomness of your cursor's direction.

Step by step:

  1. Every cursor movement has a direction (angle in degrees, e.g., 45°, 270°).
  2. We divide the full 360° circle into 8 buckets (N, NE, E, SE, S, SW, W, NW — 45° each).
  3. np.bincount(directions, minlength=8) counts how many movements fell into each bucket. Result is an array like [5, 2, 0, 8, 1, 3, 0, 4].
  4. We divide by the total count to get probabilities — [0.22, 0.09, 0.00, ...]. These sum to 1.0.
  5. entropy(probs, base=2) calculates Shannon Entropy — a single number representing how random or uniform the distribution is.

If you only ever move the mouse in one direction (e.g., all rightward), one bucket has everything — entropy is 0 (totally predictable). If you move equally in all 8 directions, entropy is at its maximum value of 3.0 (totally random/uniform). Your personal value sits somewhere between these extremes and is unique to you.


Machine Learning Preprocessing (sklearn)

StandardScaler ensures all feature vectors have a mean of 0 and variance of 1.

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()

# Training: Learn the distribution AND transform the data
x_scaled_train = scaler.fit_transform(x_raw)

# Live Verification: Transform live data using the PREVIOUSLY LEARNED distribution
x_scaled_live = scaler.transform(x_new_data)

Why do we need a scaler at all?

The 8 features SENTRY computes exist on completely different numerical scales:

Feature Typical range
SE (Saccade Entropy) 0.0 → 3.0
KDV (Key Dwell) 0.01 → 0.15
MJA (Mouse Jerk) 500 → 8000

If you feed these raw values to the Isolation Forest, it will pay almost all its attention to MJA because that number is thousands of times bigger than KDV. The model essentially ignores the small-valued features. This makes it blind to typing patterns.

StandardScaler fixes this by transforming every feature so it has:

  • A mean of 0 (the average session value becomes zero).
  • A variance of 1 (the spread is normalised to the same scale).

After scaling, MJA and KDV look equally important to the model.

Why fit_transform during training, but only transform live?

  • fit_transform: During training, the scaler learns the mean and standard deviation from the training data AND applies the transformation. This is the "calibration" step.
  • transform (only): During live scoring, we use the same calibration values the scaler learned during training. We never re-learn from live data — that would defeat the purpose, because a hacker's behaviour would shift the scaler to make their values look normal.

The scaler is saved to disk (a .pkl file) alongside the model, and must be reloaded every time the app starts. If you use a different scaler for scoring than the one used for training, the numbers won't mean anything to the model.


Anomaly Detection (sklearn)

from sklearn.ensemble import IsolationForest

# Initialize the model with 5% expected garbage data
model = IsolationForest(contamination=0.05, random_state=42)

# Train the model (No labels required, unsupervised learning)
model.fit(x_scaled_train)

# Get the raw continuous anomaly score
# Negative = Anomaly, Positive = Normal
score = model.decision_function(x_scaled_live)[0]

What is Isolation Forest doing internally?

The model builds 100 decision trees (a "forest"). Each tree works like this:

  1. Pick a feature at random (e.g., MJA).
  2. Pick a random split point within the range of that feature's values.
  3. Divide the data. Repeat until a single data point is alone ("isolated").

Normal sessions sit in the dense cluster of your training data. To isolate one of them, the tree needs many, many random splits. The average path length (number of splits needed) is large.

Anomalous sessions sit far from your cluster. They get isolated in just a few splits. Short path length = anomaly.

decision_function() returns a score based on the average path length across all 100 trees:

  • Score > 0: Normal. The session looks like your training data.
  • Score < 0: Anomalous. The session is far from your cluster.
  • Score = 0: Right on the decision boundary.

contamination=0.05 tells the model to expect that approximately 5% of your training data itself was noisy or atypical. This sets the decision boundary such that 5% of training points are treated as outliers. Setting this too high (e.g., 0.30) means the model will be overly aggressive and flag your own normal behaviour.

random_state=42 is a fixed random seed, ensuring the model builds the same trees every time you train on the same data (reproducibility).


NDJSON Data Streaming

Python communicates with Electron by flushing JSON strings directly to the standard output buffer.

import sys, json

payload = {
    "type": "telemetry",
    "z_score": -0.15,
    "state_suggestion": "LOCK"
}

# Write a single line of JSON and force the OS to flush the buffer immediately
sys.stdout.write(json.dumps(payload) + '\n')
sys.stdout.flush()

What is stdout?

Every program has three standard "pipes" — channels of text:

  • stdin: input (what's typed/sent to the program).
  • stdout: output (what the program prints to the terminal).
  • stderr: error output.

When Python is launched by Electron as a child process, Electron holds direct references to these pipes. Electron's pythonProcess.stdout.on('data', ...) is literally listening to the bytes that Python writes here.

What is NDJSON?

NDJSON stands for Newline-Delimited JSON. Instead of one big JSON object, each message is a single-line JSON string, terminated by \n. This means Electron can split the incoming text stream by newline characters to get individual complete messages — even if multiple messages arrive in the same network packet or OS buffer flush.

Why sys.stdout.flush()?

By default, Python buffers stdout output (collects it and sends it in large chunks) for efficiency. But SENTRY needs real-time data every 3 seconds. Without flush(), the payload might sit in the buffer for seconds before being sent. flush() forces Python to immediately push everything in the buffer to Electron, ensuring the UI updates in real time. The -u flag when spawning Python (python -u main.py) also disables output buffering globally.


2. Electron: Process Management & IPC

Spawning the Python Daemon

Electron acts as the parent process, booting the Python script in the background.

const { spawn } = require('child_process');

const pythonProcess = spawn('python', ['main.py'], { 
    cwd: path.join(__dirname, '../') // Run from the root directory
});

// Catch errors and close events to prevent orphaned processes
pythonProcess.on('close', (code) => console.log(`Python exited with code ${code}`));

What is spawn?

spawn is a Node.js function that launches a new operating system process — completely separate from the Electron process itself. Think of it like Electron clicking "Run" on a program. The spawned Python process runs independently, has its own memory, and can be killed separately.

spawn returns a pythonProcess object which is a live reference to that running process. Through this object, Electron can:

  • Write to its stdin — to send commands to Python.
  • Read from its stdout — to receive data from Python.
  • Read from its stderr — to catch Python errors.
  • Call .kill() — to terminate Python when the app closes.

What does cwd do?

cwd (Current Working Directory) sets the folder from which the Python script runs. This is important because Python's file operations (loading .pkl model files, reading databases) use relative paths. If Python starts from the wrong directory, it won't find its files.

Why listen to pythonProcess.on('close', ...)?

This event fires when Python's process terminates for any reason — including crashes. SENTRY's main.js uses this to implement an auto-respawn system: if Python crashes unexpectedly, Electron waits 1.5 seconds and restarts it. This is capped at 3 attempts within 60 seconds to prevent an infinite crash loop from hammering the CPU.


Reading the NDJSON Stream

Electron listens to the standard output of the child process.

pythonProcess.stdout.on('data', (dataBuffer) => {
    // Data arrives as a buffer, convert to string
    const output = dataBuffer.toString();
    
    // Split by newline (NDJSON) and parse individually
    output.split('\n').forEach(line => {
        if (!line.trim()) return;
        try {
            const parsed = JSON.parse(line);
            // Send data to the React UI window via IPC
            mainWindow.webContents.send('telemetry-update', parsed);
        } catch (e) {
            console.error("JSON Parse Error");
        }
    });
});

What is dataBuffer?

The OS delivers data to Electron in raw Buffers — fixed-size chunks of bytes. These chunks are not guaranteed to align with your JSON message boundaries. In one data event, Electron might receive half a JSON message; in the next, the second half plus two more complete messages. This is why SENTRY maintains a stdoutBuffer string — partial lines are accumulated until a \n is detected, confirming the message is complete.

Why if (!line.trim()) return;?

After splitting by \n, the last element in the array is always an empty string (because the string ends with \n). This guard skips those empty lines to avoid a JSON parse error on an empty string.

Why try/catch around JSON.parse?

Python sometimes prints non-JSON text to stdout (e.g., Python warnings, print statements from libraries). If JSON.parse receives a non-JSON string like "UserWarning: ...", it throws an error. The try/catch handles this gracefully — non-JSON lines are just logged and discarded.

What is mainWindow.webContents.send(...)?

webContents is Electron's interface to the Chromium rendering engine running inside the app window. Calling .send('channel-name', data) is like shouting a message into the window — any React component that is "listening" on that channel name will receive the data. This is called IPC (Inter-Process Communication). It is the mechanism that makes the live-updating dashboard possible.


The Security Bridge (preload.js)

Exposes specific Electron IPC channels to the React frontend safely.

const { contextBridge, ipcRenderer } = require('electron');

contextBridge.exposeInMainWorld('electronAPI', {
    // Allows React to send commands to Electron
    sendCommand: (cmdStr) => ipcRenderer.send('python-command', cmdStr),
    
    // Allows React to listen to telemetry updates
    onTelemetry: (callback) => {
        const listener = (event, data) => callback(data);
        ipcRenderer.on('telemetry-update', listener);
        // Return a cleanup function so React can unmount cleanly
        return () => ipcRenderer.removeListener('telemetry-update', listener);
    }
});

Why does this file even exist? What problem does it solve?

In Electron, the UI runs inside a Chromium (browser) window. For security, this UI environment is treated just like a browser tab — it is not allowed to access the filesystem, spawn processes, or use Node.js APIs directly. If it could, malicious JavaScript embedded in a webpage could delete your files or run viruses.

However, SENTRY's React UI does need to talk to the Electron backend (to receive telemetry, send commands). preload.js runs in a special "middle ground" — it has access to both Node.js (specifically ipcRenderer) and the browser window (window). It acts as a controlled gateway.

What does contextBridge.exposeInMainWorld do?

It creates a new property on the browser's global window object — in this case, window.electronAPI. After preload runs, every React component in the app can call window.electronAPI.sendCommand(...) or window.electronAPI.onTelemetry(...). Crucially, these are the only things React can do with Electron. It can't access ipcRenderer directly, it can't read files — only the functions listed inside exposeInMainWorld are available. This is the principle of least privilege.

What is ipcRenderer?

ipcRenderer is the client-side half of Electron's IPC system. ipcRenderer.send('channel', data) sends a message to main.js. ipcRenderer.on('channel', callback) listens for messages sent from main.js. It's like a two-way radio — one side sends on a named channel, the other side listens on that same channel name.

Why does onTelemetry return a cleanup function?

ipcRenderer.on(...) registers a persistent listener. If a React component mounts, registers a listener, then unmounts (e.g., the user navigates away), the listener continues existing in memory — even though the component is gone. This is called a memory leak. The returned cleanup function (() => ipcRenderer.removeListener(...)) is called by React's useEffect cleanup to deregister the listener when the component unmounts.


3. React: The Frontend UI

Subscribing to IPC Events (useEffect)

React components hook into the Electron IPC channels when they mount, and clean up when they unmount to prevent memory leaks.

import { useEffect, useState } from 'react';

const Dashboard = () => {
    const [anomaly, setAnomaly] = useState(false);

    useEffect(() => {
        // Subscribe to telemetry
        const unsubscribe = window.electronAPI.onTelemetry((data) => {
            if (data.state_suggestion === 'LOCK') {
                setAnomaly(true);
            }
        });

        // Cleanup on unmount
        return () => unsubscribe();
    }, []); // Empty dependency array = run once on mount

    return <div>{anomaly ? "Compromised" : "Secure"}</div>;
};

What is useState?

useState is React's way of storing a value that, when changed, automatically triggers a re-render of the component. Think of it like a variable that React watches. When you call setAnomaly(true), React sees the value changed, re-runs the component function, and updates what's shown on screen.

useState(false) sets the initial value to false. anomaly is the current value; setAnomaly is the function to change it.

What is useEffect?

useEffect runs a block of code after the component has been rendered to the screen. It's how you hook into external systems (like Electron IPC) from within a React component.

The [] (empty array) at the end is the dependency array. It tells React: "Only run this effect once, when the component first appears on screen." Without the [], the effect would run every time the component re-renders — potentially registering dozens of duplicate IPC listeners.

What does the return () => unsubscribe() do?

The function returned from a useEffect is the cleanup function. React calls it when the component is about to be removed from the screen (unmounted). unsubscribe() calls ipcRenderer.removeListener(...) under the hood (via preload.js), ensuring the IPC listener is removed and no memory leak occurs.


Solving Stale Closures (useRef)

Because useEffect callbacks are closures, they capture state variables at the time they are created. To read the live value of a state variable inside a persistent IPC listener, you must mirror it in a useRef.

const [lockPaused, setLockPaused] = useState(false);
const lockPausedRef = useRef(lockPaused);

// Keep the Ref perfectly in sync with the State
useEffect(() => {
    lockPausedRef.current = lockPaused;
}, [lockPaused]);

useEffect(() => {
    window.electronAPI.onTelemetry((data) => {
        // We read from the Ref, NOT the state variable.
        // The state variable 'lockPaused' would always evaluate to its initial value (false).
        if (data.state_suggestion === 'LOCK' && !lockPausedRef.current) {
            triggerChallenge();
        }
    });
}, []);

What is a "stale closure"?

This is one of the trickiest React concepts. When useEffect(() => { ... }, []) runs with an empty dependency array, it captures a snapshot of all the variables it uses at that moment — the moment the component first mounts. This snapshot is "frozen" inside the callback function.

Imagine the component mounts with lockPaused = false. The IPC listener is registered and the callback captures lockPaused = false forever. Even if the user later clicks a button that calls setLockPaused(true), the IPC callback still reads the frozen old value: false. The button appears to do nothing.

How does useRef solve this?

A useRef creates a special object: { current: <value> }. The key difference from useState is that changing ref.current does NOT trigger a re-render. More importantly, the ref object itself (the container {}) never changes — only its .current property does. This means the frozen closure always holds a reference to the same container, and reading .current at any time gives the latest value.

The second useEffect with [lockPaused] in its dependency array runs every time lockPaused changes. It simply writes the new value into lockPausedRef.current, keeping the ref perfectly up to date.

In summary: Use useState to trigger re-renders. Use useRef to read the latest value inside persistent callbacks that were created once and never re-register.