ML systems engineer · Founder at ScholarLM
I make ML and developer infrastructure fast, and I prove it with numbers: Metal GPU kernels, Rust and Go runtimes,
and code-intelligence tools for AI agents, built as modules that every new project reuses.
Open to opportunities · bharath@vbcr.dev · Dallas–Fort Worth, TX
- Now: founder and sole engineer of ScholarLM, an AI research platform (React, Rust, Go and Python) that writes fully-cited manuscripts. In parallel I build the performance stack below.
- Strongest results: a Qwen3.5-2B engine that runs 2.1–2.2× faster than PyTorch on Apple GPUs, a code-graph indexer that builds its index 1.6–21× faster than five other tools on every corpus tested, and a Go↔Rust call path made 24× cheaper.
- Before: Associate Researcher at the Adidas Center for Engagement Science (ASU, 2023–2025), and stem-cell research at Texas Tech University Health Sciences Center.
- Education & papers: M.S. Biomedical Engineering, Arizona State University (2024) · 2 peer-reviewed papers (2025).
Each figure comes from the project's own benchmark record, with its conditions. Where a figure didn't survive a re-check, it isn't here.
| Project | Result | Conditions |
|---|---|---|
| tessl | 2.14× / 2.12× / 2.22× faster than PyTorch MPS at 200 / 2,048 / 8,192 tokens | Identical Qwen3.5-2B work (real-weight prefill, prefix state kept, 17 answer rows scored), one GPU hold, lengths interleaved, M5 Pro, 2026-10-04. PyTorch's gated-delta layer runs its pure-torch fallback, the only path on a Mac. Separately, bf16 GEMM is 2.55× MLX. |
| DevMap | 1.6–21× faster cold index and 2.2–107× faster refresh than five other code-graph tools, on all four corpora · 9.7 ms definition lookups | v0.2.2, M5 Pro, every corpus pinned to a commit (2026-09-14). Lost single-file re-index to CodeGraph. |
| Gusset | Go→Rust serial call 90.3 → 3.73 µs (24×) | Spin-then-park plus a shared-memory completion ring, linux-amd64 VM. Backed by a chaos hammer, fuzz targets and Miri. |
| ojas | Training-block step 41.5 ms vs PyTorch's 43.1 ms on CPU | Apple M5 Pro CPU, forward + backward, 124 of 127 outputs within tolerance (2026-10-02). Some single ops are still slower. |
| GitPulse | −48% median process spawn-and-wait (3.64 → 1.89 ms) | One controlled run, 200 samples. |
| BINN | 0.8320 on Spiking Heidelberg Digits, 12/12 seeds ≥ 0.80 | Both pre-registered crux gates for backprop-free learning failed, and are reported alongside it. |
| Sequence mixers | ~6.9× Mamba-2 throughput from a chunk-parallel SSD scan | nanolab, seed-paired intervals. |
Each project is a module the next one is built on. Every arrow below is a real dependency in the source: a crate, a go.mod require, a vendored engine or a spawned sidecar. Arrows point from a project to what it builds on.
flowchart TB
subgraph K["Kernels"]
tessl["tessl<br/>Metal 4 GEMM + Qwen3.5 engine"]
sparsl["sparsl<br/>sparse + scan kernels"]
end
subgraph E["Engines & boundaries"]
ojas["ojas<br/>deep learning engine"]
gusset["Gusset<br/>Rust-in-Go contract"]
end
subgraph A["Code intelligence & agents"]
devmap["DevCouncil · DevMap<br/>code graph + gate"]
manvi["MANVI<br/>agent harness"]
jarvis["Jarvis<br/>model-free replay"]
end
subgraph P["Products"]
gitpulse["GitPulse"]
scholarlm["ScholarLM"]
devtype["DevType"]
end
subgraph R["Research"]
binn["BINN"]
lappi["Lappi"]
gemma["gemma-metal"]
end
ojas -->|kernels| tessl
ojas -->|cgo| gusset
lappi -->|Mac backend| tessl
lappi -->|Mac trainer| ojas
binn -->|crates.io| sparsl
binn -.->|optional| tessl
gemma -->|GEMMs| tessl
devmap -->|engine + gate| gusset
manvi --> devmap
manvi --> gusset
jarvis -->|replay| manvi
gitpulse -->|vendored| devmap
gitpulse -->|sidecar| manvi
gitpulse --> gusset
scholarlm -.->|dev tooling| devmap
devtype -.->|dev tooling| devmap
devtype -.->|coverage| gitpulse
sparsl was lifted out of BINN's numeric core and published on its own. GitPulse also vendors MarkDev's renderer crates, and DevPrism embeds MANVI as its tool gate. Explore the same graph interactively on bharath.vbcr.dev.
| Project | Builds on · Used by | |
|---|---|---|
| Kernels | ||
| tessl · tessl.vbcr.dev Makes the matrix math inside LLMs fast on Apple GPUs. A Metal 4 GEMM and tensor runtime in Rust: MPP TensorOps matmul2d, cooperative register accumulators, fused epilogues. Now a Qwen3.5-2B engine with a full forward and backward training step, checked against transformers under pre-written tolerances. |
Used by ojas, Lappi, gemma-metal, BINN | |
| sparsl · Fast, reproducible kernels for spiking-network simulation: CSR SpMV, LIF membrane updates and a chunked prefix scan, all deterministic. A Device exists only for a backend that can actually execute, so results reproduce bit for bit and never misreport where they ran. |
Lifted out of BINN · used by BINN | |
| Engines & boundaries | ||
| ojas · ojas.vbcr.dev Trains and runs neural networks inside Go services, with no Python runtime. A Rust deep learning engine that reads its machine (cores, caches, unified memory, cgroup limits) before it plans work, with PyTorch kept only as the reference oracle. |
Builds on tessl, Gusset · used by Lappi | |
| Gusset · gusset.vbcr.dev Lets a Go service call a Rust engine safely and cheaply. The runtime contract: panic firewall, bounded concurrency, deadlines enforced inside Rust, poisoned handles, per-field ABI checks, allocator accounting. MIT / Apache-2.0. |
Used by DevCouncil, MANVI, ojas, GitPulse | |
| Code intelligence & agents | ||
| DevCouncil · DevMap · devcouncil.vbcr.dev Gives AI coding agents a fast, accurate map of a codebase. Native Go and Rust code-intelligence and verification components. DevMap is the code graph: devmap ask with an evidence pack of related code and tests, blast radius with owners, and commit regression suspects. The write gate runs fail-closed on Gusset. |
Builds on Gusset · used by MANVI, GitPulse, ScholarLM tooling | |
| MANVI · manvi.vbcr.dev Runs AI coding agents under explicit policy. A coding-agent harness in Go and Rust: dual-plane execution across a process boundary, and a six-step policy ladder whose outcomes stay distinct, so a check that could not run never reads as a pass. 1,031 cross-language parity cases hold the two planes to one behaviour. |
Builds on DevCouncil, Gusset · used by GitPulse, Jarvis, DevPrism | |
| Jarvis · jarvis.vbcr.dev Desktop capabilities discovered once with Gemini, frozen into typed artifacts, then replayed through MANVI with no model decisions: 35/40 macOS replays and 40/40 saved-state checks, not yet a clean stability pass. Human approval gates every account change. |
Builds on MANVI | |
| Products | ||
| GitPulse · gitpulse.vbcr.dev Native workspace for Git, review, tasks and AI agent sessions, in one Tauri 2 / Rust / Svelte 5 process. Links DevMap in-process for code-graph regression suspects and supervises Claude Code and Codex in a managed lane through MANVI. Zero telemetry. |
Builds on DevCouncil, MANVI, Gusset | |
| ScholarLM · showcase & architecture AI research platform that searches the literature and writes fully-cited, grounded manuscripts: plan → write → verify → review, with every claim traced to a retrieved source. React, a Rust edge, a Go orchestration core and a Python ML worker; also served as an MCP tool server. Its agent layer is open as WisDev. |
Its coding agents navigate it with DevMap | |
| DevType · devtype.vbcr.dev Native macOS text expander and on-device writing assistant in Swift/AppKit. Typed triggers expand in ordinary text fields; proofread, rewrite, translate and code actions run on Apple Foundation Models without leaving the Mac. Imports TextExpander and Espanso libraries, and keeps passwords apart from snippets behind Touch ID. |
DevCouncil verifies its changes; its coverage export feeds GitPulse (development tooling) | |
| Research | ||
| 🧠 | BINN · binn.vbcr.dev A from-scratch Rust instrument built to falsify one question: can a sparse, locally learned, event-driven network learn without backpropagation? Under pre-registered kill-gates the answer was no, and that stays on the record. The same instrument then earned a positive: temporal spike order is the mechanism behind its SHD result. |
Builds on sparsl, tessl |
| Lappi · lappi.vbcr.dev An open, calibrated typed-decision model: schema in, typed slots out (choice, score, span, abstain), with split-conformal calibration and line-level grounding. Promotion needs every gate to have run and passed; a gate that did not run is never counted as a pass. The 606K-parameter byte model reached 80.97% against the 88.5% control it must beat; the 2B campaign continues. |
Builds on tessl, ojas | |
| nanolab · attention.vbcr.dev Instrumented small-LM training lab: attention, Mamba-2, Gated DeltaNet and minGRU behind CLI flags, chunk-parallel scan kernels, and multi-seed ablations reported as intervals, with the experiment record and the replication manuscript. |
Research companion to the kernels | |
| gemma-metal Gemma inference runtime for Apple silicon with split sliding/global KV ring caches. It takes its general and INT4 GEMM kernels from tessl, and it is still below its own decode-speed gate. |
Builds on tessl | |
These are directions, each grounded in what the repositories themselves record as open:
- Train and serve Lappi on my own stack. Lappi's Mac trainer already runs on ojas and its backend on tessl's Qwen3.5 kernels, while the 2B campaign runs on cloud GPUs. The goal is a calibrated decision model that passes its own gates and is served locally.
- A device-aware engine for Go services. ojas reaches Go through Gusset, and its resource plan reads the machine. There is no device router yet, so the plan is only advice. Routing work between CPU and Metal from that plan is next.
- Spiking read-outs on the current kernels. BINN's tessl interop is pinned to 0.1.4, while tessl has moved to 0.2.0 with the Qwen3.5 engine. Bringing the attention read-out that earned the SHD result onto the current kernels comes next.
- The same tools in every repository. DevMap already sits under GitPulse and under the coding agents working on ScholarLM and DevType, and MANVI under GitPulse, Jarvis and DevPrism, so each improvement to the graph or the harness lands in all of them at once.
| Paper | Journal | Year |
|---|---|---|
| Investigation on the heating effects of intra-tumoral injectable magnetic hydrogels (IT-MG) for cancer hyperthermia | Biomedical Physics & Engineering Express | 2025 |
| The Therapeutic Scope of Orofacial Mesenchymal Stem Cells | Bioengineering | 2025 |
Biomedical computing: GenoThermal_Targeting, a patient-specific magnetic-nanoparticle therapy pipeline from genomic discovery through physics simulation. Research briefs are at research.vbcr.dev.
Side projects: finished or maintained, but outside the main stack
- Chronicle: local-first second brain across Mac and Android, with on-device embeddings and RAG.
- MarkDev: native macOS Markdown editor on a Swift + Rust core. Its renderer crates are vendored into GitPulse.
- DevPrism: local-first LaTeX and research workspace, forked from claude-prism, with MANVI as its tool gate.
- M5Blade: Apple-silicon fan controller that writes to the SMC behind a race-free control gate.
- Strait: macOS bulk transfer with BLAKE3 hash-on-write and resumable staging.
- Curio, ChronosFlow, Meridian: on-device AI mobile apps.
- SalEdge: multi-firm ERP for battery retailers, with GST e-invoicing and a local AI layer.
- AcademiaTrack, Void, Whimsical-Love: web apps and experiences.
Everything has a page at apps.vbcr.dev.

